WEB-SOBA: Streamlining Domain Ontologies with Word Embeddings for Hybrid ABSA

WEB-SOBA: Word Embeddings-Based Semi-automatic Ontology Building for Aspect-Based Sentiment Classification

2021-01-01
Fenna ten Haaf, Christopher Claassen, Ruben Eschauzier, Joanne Tjan, Daniël Buijs, Flavius Frasincar, Kim Schouten
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces WEB-SOBA, a semi-automatic methodology for constructing domain-specific sentiment ontologies for Aspect-Based Sentiment Analysis (ABSA) using word embeddings. By integrating word2vec and sentiment-based vector refinement, it builds a structure that achieves state-of-the-art accuracy when used within the HAABSA hybrid framework.

Executive Summary

TL;DR: WEB-SOBA is a semi-automatic framework that builds domain sentiment ontologies by leveraging the semantic power of word embeddings (word2vec). It slashes the human labor required for ontology construction by over 50% while boosting the accuracy of Aspect-Based Sentiment Analysis (ABSA) to 87.16% when integrated into a hybrid neural-symbolic framework.

In the landscape of NLP, this work bridges the gap between traditional symbolic AI (ontologies) and modern connectionist AI (neural networks), proving that these two paradigms are most potent when used in tandem.

The "Knowledge Bottleneck" in ABSA

Aspect-Based Sentiment Analysis (ABSA) is crucial for businesses to understand specific feedback (e.g., "The food was great, but the service was slow"). Hybrid models—which use ontologies for logical reasoning and machine learning as a fallback—currently represent the SOTA.

However, the bottleneck has always been the ontology itself.

  1. Manual Construction: Too slow (hours/days of expert effort).
  2. Frequency-based (Co-occurrence): Faster, but lacks semantic depth. It treats words as isolated tokens rather than points in a semantic space.

Methodology: How WEB-SOBA Works

The core innovation of WEB-SOBA lies in its three-stage pipeline driven by word embeddings.

1. Term Selection via Semantic Density

Instead of just counting words, WEB-SOBA calculates a TermScore (TS) using:

  • Domain Similarity (DS): Comparing domain-specific vectors (Yelp) against general vectors (Google News) to find unique restaurant terminology.
  • Mention Class Similarity (MCS): Judging how close a word is to core concepts like "Food" or "Service" in the vector space.

2. Sentiment Refinement

A notorious weakness of word2vec is that "good" and "bad" often have similar vectors because they appear in the same contexts. WEB-SOBA solves this by sentiment-aware refinement, pushing vectors closer to known emotional anchors (using the E-ANEW lexicon) or further apart based on polarity.

3. Hierarchical Clustering

The system uses Average Linkage Clustering (ALC) to group terms into a hierarchy (e.g., "Fries" "Food"). This creates the "is-a" relationships necessary for ontological reasoning.

WEB-SOBA Mention Subclasses Figure 1: The skeletal structure of the Mention classes used as the foundation for the ontology.

Experiments & Results: Efficiency meets Accuracy

The authors evaluated WEB-SOBA on the SemEval-2016 restaurant dataset.

The Speed Advantage

WEB-SOBA requires significantly less "human-in-the-loop" time:

  • Manual: 420 mins
  • SOBA (Prior SOTA): 90 mins
  • WEB-SOBA: 40 mins

The Performance Leap

When used as part of the HAABSA framework (pairing the ontology with a Rotatory Attention neural network), WEB-SOBA outperformed even the handcrafted manual ontology.

Performance Comparison Table 1: Hybrid approach results showing WEB-SOBA achieving 87.16% accuracy.

The Welch t-test (p < 0.05) confirmed that this improvement isn't just noise—it's a statistically significant advancement. WEB-SOBA complements neural models better because its embedding-based construction captures the same semantic nuances the neural network is looking for.

Critical Insight & Conclusion

Why does it work?

WEB-SOBA succeeds because it treats ontology building as a mapping problem rather than a discovery problem. By starting with word embeddings, the system already "understands" which words are related; the user simply acts as a high-level filter to ensure the logic remains sound.

Limitations & Future Paths

  • Polysemy: Current word2vec models struggle with words that have multiple meanings. The authors suggest moving toward BERT-based contextual embeddings in future iterations.
  • Type-3 Sentiment: The system struggled to identify context-dependent sentiment (e.g., "cold" is positive for beer but negative for pizza) due to dataset sparsity.

Final Takeaway: WEB-SOBA is a major step toward Scalable Knowledge Engineering. It proves that we can build high-precision expert systems without the grueling manual labor traditionally associated with them.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize contextualized embeddings like BERT or RoBERTa for semi-automatic ontology construction in sentiment analysis.
  • Identify the seminal paper that introduced sentiment-aware word embedding refinement (e.g., retrofitting or counter-fitting) and how WEB-SOBA adapts it.
  • Explore how the WEB-SOBA methodology can be extended to multi-modal sentiment analysis where ontology nodes correspond to both text and image features.
Contents
WEB-SOBA: Streamlining Domain Ontologies with Word Embeddings for Hybrid ABSA
1. Executive Summary
2. The "Knowledge Bottleneck" in ABSA
3. Methodology: How WEB-SOBA Works
3.1. 1. Term Selection via Semantic Density
3.2. 2. Sentiment Refinement
3.3. 3. Hierarchical Clustering
4. Experiments & Results: Efficiency meets Accuracy
4.1. The Speed Advantage
4.2. The Performance Leap
5. Critical Insight & Conclusion
5.1. Why does it work?
5.2. Limitations & Future Paths