Beyond Embeddings: Boosting QA Passage Ranking with Shallow Linguistic Features
Answer Passage Ranking Enhancement Using Shallow Linguistic Features
The paper proposes an enhanced answer passage ranking method for fact-seeking Question Answering (QA) by integrating shallow linguistic features with deep sentence embeddings. By augmenting a feed-forward neural network with features like noun counts and named entity counts, the authors achieve significant performance gains, notably outperforming strong baselines like R3 and pure InferSent rankers at R@5 on the QUASAR-T dataset.
TL;DR
While Large Language Models and sentence embeddings dominate modern NLP, they often struggle with the "specificity" required for fact-seeking Question Answering. This paper demonstrates that by injecting simple linguistic features—like how many nouns or named entities a passage contains—directly into a neural network's hidden layers, we can significantly improve answer recall and ranking accuracy, outperforming pure semantic-based models by up to 7%.
The "Semantic Gap" in Question Answering
In open-domain QA, the goal is to find the specific needle in the haystack: the passage that contains the actual answer. Current SOTA methods use distributed representations (embeddings) to measure how "related" a passage is to a question.
However, there is a catch. Consider the question: "When did Google start?"
- Passage A: "Google was launched by Larry Page and Sergey Brin at Stanford." (High semantic similarity, but no answer).
- Passage B: "Google began operations in 1998." (Lower semantic overlap, but contains the answer).
Traditional embeddings often rank Passage A higher because they prioritize general semantic relatedness over the presence of specific answer-type indicators (like timestamps or nouns).
Methodology: Fusing Syntax with Semantics
The researchers identified a major limitation: distributed embeddings "cram" information into a single vector, often losing track of shallow linguistic counts that indicate a passage's "answer-bearing" potential.
1. Feature Engineering
They identified 7 key shallow features:
- Noun & Named Entity Counts: Fact-seeking answers are almost always nouns or entities.
- Query Coverage: Proximity to the original question's terms.
- Pronoun Counts: High pronoun counts often "mask" the actual answer (e.g., using "She" instead of "The Queen").
2. The Architecture: Middle-Layer Augmentation
The core innovation isn't just using these features, but where they are used. The authors found that simply adding these features to a final classifier (Late Fusion) didn't work. Instead, they used a Deep Learning-Based Augmentation:

As shown in the diagram, the linguistic features are concatenated to the output of the first dense layer. This forces the model to learn a representation that balances "latent semantics" (from InferSent) with "explicit structure" (the counts) before making the final ranking decision.
Experiments & Results
The researchers tested their approach on the QUASAR-T dataset. The results were clear: pure deep learning (BL-NN) is significantly enhanced when shallow features are added.

- Mean Rank (MR): Improved from 9.58 to 8.79 (lower is better).
- Recall@5: Reached 0.63, significantly beating the baseline (0.58) and the previous InferSent state-of-the-art (0.56).
Interestingly, their statistical analysis showed that verbs count for very little in ranking fact-seeking answers, whereas nouns and named entities have a correlation of 0.57 and 0.55 with the actual answer probability.
Critical Analysis & Conclusion
This work serves as a vital reminder for AI practitioners: Neural networks are not magic. While embeddings are powerful, they are essentially "blurry" summaries of text. For tasks requiring high precision—like finding a specific date or name—the "shallow" signals we used a decade ago (POS tagging, entity counting) still provide a necessary inductive bias.
Limitations: The study focuses on "short passages." In a long-document retrieval scenario, these count-based features might become noisier. Future work should look at how these features interact with Transformer-based attention maps directly.
Takeaway for Researchers
If your retrieval model is struggling to distinguish between "topic-related" and "answer-containing" content, don't just reach for a larger LLM. Try looking at the linguistic structure—it might be the missing piece of the puzzle.
