Ontology-Enhanced ABSA: Achieving SOTA Performance with 80% Less Data

Ontology-Enhanced Aspect-Based Sentiment Analysis

2017-01-01
Kim Schouten, Flavius Frasincar, Franciska de Jong
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces an Ontology-Enhanced Aspect-Based Sentiment Analysis (ABSA) framework that integrates domain-specific knowledge into a machine learning pipeline. By leveraging a structured ontology for aspect detection and sentiment classification, the method achieves SOTA-level performance while significantly reducing the dependency on large labeled datasets.

TL;DR

In an era dominated by data-hungry deep learning, this paper revisits Knowledge Engineering to solve Aspect-Based Sentiment Analysis (ABSA). By injecting a domain-specific ontology into an SVM classifier, the authors created a system that matches SOTA performance on the SemEval-2015 benchmark while requiring only a fraction of the training data.

The "Data Hunger" Problem in Sentiment Analysis

Most current ABSA systems rely on deep neural networks (CNNs, LSTMs, Transformers). While powerful, these models are notorious for:

  1. Data Dependency: They require thousands of manually labeled examples to learn basic domain concepts.
  2. Contextual Blindness: They often struggle with word-sense disambiguation—knowing that "cold" is a compliment for a drink but an insult for an entree.

The authors argue that we shouldn't force a model to "re-learn" that a steak is a type of food in every new dataset. Instead, we should provide this Common Domain Knowledge via a structured ontology.

Methodology: Fusing Logic with Statistical Learning

The team developed a hybrid pipeline that combines a standard NLP preprocessing stack (Lemmatization, POS tagging, Synset extraction) with a specialized Restaurant Ontology.

The Ontology Design

The ontology is split into three core pillars:

  • Target (Aspects): A hierarchy where Steak Meat Food.
  • Sentiment Expressions: A collection of evaluative words (e.g., "delicious," "slow").
  • Relational Logic: A specific mapping that connects sentiments to targets. This allows for reasoning: if the text says "cold," the model checks the target. If the target is Beer, it assigns a positive weight; if Pizza, a negative one.

Model Architecture and NLP Pipeline

Feature Engineering

The model uses a Linear SVM. The "secret sauce" is the feature vector:

  • Binary Features: Presence of lemmas, WordNet synsets, and ontology concepts.
  • Sentiment Averaging: A weighted average combining scores from the Stanford Sentiment Tool, NRC dictionaries, and the ontology.
  • Distance Correction: Grammatical proximity is used to ensure sentiment words are correctly linked to their respective aspects.

Experiments: Performance & Data Efficiency

The researchers tested their approach against the SemEval-2015 restaurant review dataset.

1. The Power of "Prior Knowledge"

The results showed a clear hierarchy in performance:

  • Base SVM: F1 0.574
  • SVM + WordNet (+S): F1 0.631
  • SVM + Ontology (+O): F1 0.687
  • Full Hybrid (+SO): F1 0.698

2. Sensitivity to Data Size (The "Killer" Result)

The most striking finding was in Aspect Detection. As shown in the graph below, while the "Base" algorithm's performance plummeted as training data was removed, the Ontology-enhanced (+SO) version remained remarkably stable. At just 20% of the training data, the +SO model performed as well as the Base model did with 100% of the data.

Data Size Sensitivity Comparison

Deep Insight: Peeking into the "Black Box"

By analyzing the SVM weights, the authors proved that the ontology concepts were indeed the most influential features. For category-specific classifiers like DRINKS#PRICES, the top features were not just raw words, but the semantic concepts of Price and Drink.

Feature Weights Table

Critical Analysis & Conclusion

Takeaway

This paper serves as a vital reminder for the AI community: Knowledge is a shortcut to Intelligence. In domains where data is expensive to label—such as legal, medical, or specialized industrial sectors—building a focused ontology is a much more efficient path to SOTA performance than collecting massive datasets.

Limitations

  • Manual Effort: The ontology was manually curated. While effective, it creates a bottleneck for scaling across hundreds of different domains.
  • Implicit Targets: The model still struggles with identifying "implicit" aspects (where the sub-aspect isn't explicitly named) compared to explicit ones.

Future Work

The next frontier is the automated construction of ontologies from web-scale data (Linked Open Data), combining the scalability of big data with the precision of formal logic.

Find Similar Papers

Try Our Examples

  • Find recent papers on Hybrid ABSA models that combine Knowledge Graphs with Transformer-based architectures like BERT or RoBERTa.
  • Who originally proposed the OntoClean methodology for evaluating ontological decisions, and how has it been automated in recent years?
  • Search for studies applying Domain Ontologies to Aspect-Based Sentiment Analysis in medical or financial document mining where labeled data is scarce.
Contents
Ontology-Enhanced ABSA: Achieving SOTA Performance with 80% Less Data
1. TL;DR
2. The "Data Hunger" Problem in Sentiment Analysis
3. Methodology: Fusing Logic with Statistical Learning
3.1. The Ontology Design
3.2. Feature Engineering
4. Experiments: Performance & Data Efficiency
4.1. 1. The Power of "Prior Knowledge"
4.2. 2. Sensitivity to Data Size (The "Killer" Result)
5. Deep Insight: Peeking into the "Black Box"
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Work