Ontology-Enhanced ABSA: Achieving SOTA Performance with 80% Less Data
Ontology-Enhanced Aspect-Based Sentiment Analysis
This paper introduces an Ontology-Enhanced Aspect-Based Sentiment Analysis (ABSA) framework that integrates domain-specific knowledge into a machine learning pipeline. By leveraging a structured ontology for aspect detection and sentiment classification, the method achieves SOTA-level performance while significantly reducing the dependency on large labeled datasets.
TL;DR
In an era dominated by data-hungry deep learning, this paper revisits Knowledge Engineering to solve Aspect-Based Sentiment Analysis (ABSA). By injecting a domain-specific ontology into an SVM classifier, the authors created a system that matches SOTA performance on the SemEval-2015 benchmark while requiring only a fraction of the training data.
The "Data Hunger" Problem in Sentiment Analysis
Most current ABSA systems rely on deep neural networks (CNNs, LSTMs, Transformers). While powerful, these models are notorious for:
- Data Dependency: They require thousands of manually labeled examples to learn basic domain concepts.
- Contextual Blindness: They often struggle with word-sense disambiguation—knowing that "cold" is a compliment for a drink but an insult for an entree.
The authors argue that we shouldn't force a model to "re-learn" that a steak is a type of food in every new dataset. Instead, we should provide this Common Domain Knowledge via a structured ontology.
Methodology: Fusing Logic with Statistical Learning
The team developed a hybrid pipeline that combines a standard NLP preprocessing stack (Lemmatization, POS tagging, Synset extraction) with a specialized Restaurant Ontology.
The Ontology Design
The ontology is split into three core pillars:
- Target (Aspects): A hierarchy where
SteakMeatFood. - Sentiment Expressions: A collection of evaluative words (e.g., "delicious," "slow").
- Relational Logic: A specific mapping that connects sentiments to targets. This allows for reasoning: if the text says "cold," the model checks the target. If the target is
Beer, it assigns a positive weight; ifPizza, a negative one.

Feature Engineering
The model uses a Linear SVM. The "secret sauce" is the feature vector:
- Binary Features: Presence of lemmas, WordNet synsets, and ontology concepts.
- Sentiment Averaging: A weighted average combining scores from the Stanford Sentiment Tool, NRC dictionaries, and the ontology.
- Distance Correction: Grammatical proximity is used to ensure sentiment words are correctly linked to their respective aspects.
Experiments: Performance & Data Efficiency
The researchers tested their approach against the SemEval-2015 restaurant review dataset.
1. The Power of "Prior Knowledge"
The results showed a clear hierarchy in performance:
- Base SVM: F1 0.574
- SVM + WordNet (+S): F1 0.631
- SVM + Ontology (+O): F1 0.687
- Full Hybrid (+SO): F1 0.698
2. Sensitivity to Data Size (The "Killer" Result)
The most striking finding was in Aspect Detection. As shown in the graph below, while the "Base" algorithm's performance plummeted as training data was removed, the Ontology-enhanced (+SO) version remained remarkably stable. At just 20% of the training data, the +SO model performed as well as the Base model did with 100% of the data.

Deep Insight: Peeking into the "Black Box"
By analyzing the SVM weights, the authors proved that the ontology concepts were indeed the most influential features. For category-specific classifiers like DRINKS#PRICES, the top features were not just raw words, but the semantic concepts of Price and Drink.

Critical Analysis & Conclusion
Takeaway
This paper serves as a vital reminder for the AI community: Knowledge is a shortcut to Intelligence. In domains where data is expensive to label—such as legal, medical, or specialized industrial sectors—building a focused ontology is a much more efficient path to SOTA performance than collecting massive datasets.
Limitations
- Manual Effort: The ontology was manually curated. While effective, it creates a bottleneck for scaling across hundreds of different domains.
- Implicit Targets: The model still struggles with identifying "implicit" aspects (where the sub-aspect isn't explicitly named) compared to explicit ones.
Future Work
The next frontier is the automated construction of ontologies from web-scale data (Linked Open Data), combining the scalability of big data with the precision of formal logic.
