AMN: Breaking the Parsing Bottleneck via Associate Memory Networks
Memory Network for Linguistic Structure Parsing
This paper introduces the Associate Memory Network (AMN), a novel component designed to augment linguistic structure parsers for syntactic dependency parsing and semantic role labeling (SRL). By leveraging an external memory bank of training instances, the model achieves new state-of-the-art heights on the Penn Treebank (PTB) and CoNLL-2009 benchmarks.
TL;DR
In the world of linguistic parsing, we often rely on models to "memorize" grammar through hidden weights. This paper proposes a different path: Associate Memory Network (AMN). By allowing a parser to "look up" similar sentences in a textbook of training data during inference, the authors achieved SOTA results in both Dependency Parsing and Semantic Role Labeling (SRL), effectively solving long-distance dependency errors that haunt standard models.
Problem & Motivation: The Dense Compression Trap
Modern parsers (like the Biaffine parser) have hit a wall. While BERT and ELMo provided massive boosts, architectural innovations have stagnated. The core issue is that standard neural networks are "eager learners"—they compress vast amounts of training data into fixed-size dense vectors.
During this compression, fine-grained structural nuances are often lost. When a model encounters a complex, long-distance dependency, it lacks the "buffer" to recall how similar structures were handled in the training set. The authors argue for a "lazy learning" approach: keep the training data (the treebank) as an explicit external memory and query it during the parsing process.
Methodology: The "Textbook" Lookup Mechanism
The proposed Associate Memory Network (AMN) functions like a neural controller paired with a RAM of annotated structures.
1. Memory Addressing (The Filter)
It is computationally impossible to compare a new sentence to every sentence in a treebank. The authors use a Filter to select the top most relevant sentences. Among methods like Word Moving Distance (WMD) and Smooth Inverse Frequency (SD), Edit Distance (ED) on POS tag sequences proved most effective, as it captures the structural skeleton of the sentence rather than just its vocabulary.
2. Feature Extraction & Alignment
Once the "memory sentences" are retrieved, the model doesn't just copy them. It uses a Transformer-based attention mechanism to align the input sentence with the memory sentences.

The model calculates a similarity matrix between the input and the memory. For dependency parsing, it retrieves relative distances and relation labels from the memory to form a "memory feature" that guides the final Biaffine scorer.
3. Read-Only vs. Read-Write Memory
Unique to this work is the dual-memory design:
- Read-Only (Gold): Fixed training data.
- Read-Write (Silver): During inference/training, the model writes its own predictions back into a buffer, allowing it to adapt to specific domains or maintain consistency in longer documents.
Experiments: New Heights in Performance
The AMN was tested on the Penn Treebank (PTB), Chinese Treebank (CTB), and Universal Dependencies (UD).
- Syntactic Performance: On PTB, the BIAF+AMN version reached 96.12% UAS, a statistically significant improvement over the 95.74% baseline.
- Multilingual Robustness: In the UD v2.3 test across 12 languages, AMN improved results across the board, proving that structural retrieval is a language-agnostic benefit.
- SRL Excellence: In Semantic Role Labeling, the system reached a 90.2% F1 score, proving competitive even against complex ensemble models.

Why it works: Handling the "Long Tail"
Error analysis (Fig 3 & 4 in the paper) shows that the AMN-enhanced model drastically reduces error rates in long-distance dependencies. While standard models struggle as sentences grow longer, the AMN provides a structural anchor by finding similar patterns in its memory bank.
Critical Analysis & Conclusion
The AMN framework represents a shift toward Retrieval-Augmented Parsing. Its strength lies in its ability to utilize "golden" examples directly, bypassing the lossy nature of weight-based parameters.
Limitations: The primary drawback is the computational overhead. Maintaining and querying a large memory bank increases inference time and GPU memory demands. The authors acknowledge this, noting that while it sets new SOTA records, real-time application requires further optimization of the memory addressing stage.
Future Outlook: This work paves the way for "Online Learning" parsers that get better as they process more data (by writing to their silver memory) without requiring full retraining. It suggests that for high-precision tasks like linguistic parsing, a "textbook" is often better than a "memory."
