AMN: Breaking the Parsing Bottleneck via Associate Memory Networks

Memory Network for Linguistic Structure Parsing

2020-01-01
Zuchao Li, Chaoyu Guan, Hai Zhao, Rui Wang, Kevin Parnow, Zhuosheng Zhang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces the Associate Memory Network (AMN), a novel component designed to augment linguistic structure parsers for syntactic dependency parsing and semantic role labeling (SRL). By leveraging an external memory bank of training instances, the model achieves new state-of-the-art heights on the Penn Treebank (PTB) and CoNLL-2009 benchmarks.

TL;DR

In the world of linguistic parsing, we often rely on models to "memorize" grammar through hidden weights. This paper proposes a different path: Associate Memory Network (AMN). By allowing a parser to "look up" similar sentences in a textbook of training data during inference, the authors achieved SOTA results in both Dependency Parsing and Semantic Role Labeling (SRL), effectively solving long-distance dependency errors that haunt standard models.

Problem & Motivation: The Dense Compression Trap

Modern parsers (like the Biaffine parser) have hit a wall. While BERT and ELMo provided massive boosts, architectural innovations have stagnated. The core issue is that standard neural networks are "eager learners"—they compress vast amounts of training data into fixed-size dense vectors.

During this compression, fine-grained structural nuances are often lost. When a model encounters a complex, long-distance dependency, it lacks the "buffer" to recall how similar structures were handled in the training set. The authors argue for a "lazy learning" approach: keep the training data (the treebank) as an explicit external memory and query it during the parsing process.

Methodology: The "Textbook" Lookup Mechanism

The proposed Associate Memory Network (AMN) functions like a neural controller paired with a RAM of annotated structures.

1. Memory Addressing (The Filter)

It is computationally impossible to compare a new sentence to every sentence in a treebank. The authors use a Filter to select the top most relevant sentences. Among methods like Word Moving Distance (WMD) and Smooth Inverse Frequency (SD), Edit Distance (ED) on POS tag sequences proved most effective, as it captures the structural skeleton of the sentence rather than just its vocabulary.

2. Feature Extraction & Alignment

Once the "memory sentences" are retrieved, the model doesn't just copy them. It uses a Transformer-based attention mechanism to align the input sentence with the memory sentences.

Model Architecture

The model calculates a similarity matrix between the input and the memory. For dependency parsing, it retrieves relative distances and relation labels from the memory to form a "memory feature" that guides the final Biaffine scorer.

3. Read-Only vs. Read-Write Memory

Unique to this work is the dual-memory design:

  • Read-Only (Gold): Fixed training data.
  • Read-Write (Silver): During inference/training, the model writes its own predictions back into a buffer, allowing it to adapt to specific domains or maintain consistency in longer documents.

Experiments: New Heights in Performance

The AMN was tested on the Penn Treebank (PTB), Chinese Treebank (CTB), and Universal Dependencies (UD).

  • Syntactic Performance: On PTB, the BIAF+AMN version reached 96.12% UAS, a statistically significant improvement over the 95.74% baseline.
  • Multilingual Robustness: In the UD v2.3 test across 12 languages, AMN improved results across the board, proving that structural retrieval is a language-agnostic benefit.
  • SRL Excellence: In Semantic Role Labeling, the system reached a 90.2% F1 score, proving competitive even against complex ensemble models.

Experimental Results

Why it works: Handling the "Long Tail"

Error analysis (Fig 3 & 4 in the paper) shows that the AMN-enhanced model drastically reduces error rates in long-distance dependencies. While standard models struggle as sentences grow longer, the AMN provides a structural anchor by finding similar patterns in its memory bank.

Critical Analysis & Conclusion

The AMN framework represents a shift toward Retrieval-Augmented Parsing. Its strength lies in its ability to utilize "golden" examples directly, bypassing the lossy nature of weight-based parameters.

Limitations: The primary drawback is the computational overhead. Maintaining and querying a large memory bank increases inference time and GPU memory demands. The authors acknowledge this, noting that while it sets new SOTA records, real-time application requires further optimization of the memory addressing stage.

Future Outlook: This work paves the way for "Online Learning" parsers that get better as they process more data (by writing to their silver memory) without requiring full retraining. It suggests that for high-precision tasks like linguistic parsing, a "textbook" is often better than a "memory."

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize non-parametric memory or retrieval-augmented generation (RAG) specifically for syntactic and semantic parsing tasks.
  • Which paper first proposed the "Biaffine Attention" mechanism for dependency parsing, and how has its implementation evolved in recent SOTA parsers?
  • Explore research that applies Associate Memory Networks or similar "textbook" paradigms to low-resource cross-lingual transfer learning in NLP.
Contents
AMN: Breaking the Parsing Bottleneck via Associate Memory Networks
1. TL;DR
2. Problem & Motivation: The Dense Compression Trap
3. Methodology: The "Textbook" Lookup Mechanism
3.1. 1. Memory Addressing (The Filter)
3.2. 2. Feature Extraction & Alignment
3.3. 3. Read-Only vs. Read-Write Memory
4. Experiments: New Heights in Performance
4.1. Why it works: Handling the "Long Tail"
5. Critical Analysis & Conclusion