IOLDAIR: Bridging the Semantic Gap in Big Data Retrieval via Fuzzy Ontologies

Intelligent ontology based semantic information retrieval using feature selection and classification

2018-02-03
B. Selvalakshmi, M. Subramaniam
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces IOLDAIR, a hybrid Semantic Information Retrieval framework combining Ontology matching, Latent Dirichlet Allocation (LDA), and Fuzzy Rough Set theory. It aims to improve retrieval relevancy and speed in big data environments like social networks (Twitter, Facebook) by performing deep semantic analysis rather than simple keyword matching.

TL;DR

The paper proposes IOLDAIR (Intelligent Ontology and Latent Dirichlet Allocation based Information Retrieval), a sophisticated framework designed to tackle the "relevancy crisis" in big data. By combining Fuzzy Rough Sets for intelligent feature selection and Ontology matching for semantic depth, the system achieves a staggering 99%+ accuracy in document classification while significantly reducing computational latency compared to traditional LDA and SOR models.

Background & Positioning

In an era where Facebook and Twitter generate data at a velocity that traditional syntactic search engines cannot match, the "meaning" of a query often gets lost. Syntactic approaches (like PageRank) focus on link popularity and keyword overlap. This work positions itself as a semantic paradigm shift, moving from "what word did you type?" to "what concept are you looking for?" by using multi-level inheritance in ontologies.

Motivation: The Noise and Latecy Problem

Existing Information Retrieval (IR) systems face two major hurdles when dealing with social media:

  1. High Noise: Informal language, abbreviations, and irrelevant "jitter" in tweets.
  2. Semantic Ambiguity: The same keyword can have vastly different meanings depending on the context (e.g., "service" in a hotel vs. a software context).

The authors argue that neither pure probabilistic models (like LDA) nor pure rule-based systems are sufficient. They propose a hybrid that uses Soft Computing to handle the uncertainty of natural language.

Methodology: The IOLDAIR Architecture

The core of the work lies in a dual-algorithm approach designed for the MapReduce/Big Data era.

1. Intelligent Fuzzy Rough Set Feature Selection

Instead of keeping all terms, the system calculates a Significance Score () for each feature. It uses fuzzy lower and upper approximations to handle incomplete information.

  • Formula Logic: It identifies keywords, performs ontology matching to find the "importance" of a feature, and categorizes them into relevancy tiers (Low to High).
  • Result: Irrelevant features are pruned early, drastically reducing the dimensions for the second stage.

2. Semantic Analysis and IOLDAIR

The system doesn't just look for words; it builds a Query Ontology. System Architecture

  • Ontology Alignment: It groups web documents based on property inheritance (e.g., if a user searches for "Hotel," the system understands "Room," "AC," and "Service" are related via sub-class relations).
  • LDA Integration: It applies Latent Dirichlet Allocation to model the distribution of topics, but weights these topics based on the ontology hierarchy scores.

Experimental Performance: SOTA Results

The authors benchmark IOLDAIR against standard LDA and the Semantic Ontology Retrieval (SOR) algorithm across massive datasets of web docs, tweets, and Facebook comments.

Performance Metrics

  • Accuracy: The proposed classifier achieved 99.24% accuracy on 5,000 records, outperforming LDA (91.36%) and SOR (92.45%).
  • Speed: In the "Selected Features" mode, the execution time was nearly half that of standard LDA (0.29s vs 0.62s for large datasets).

Accuracy Comparison The figure above illustrates the consistent superiority of IOLDAIR in relevancy accuracy across varying document volumes compared to traditional baselines.

Critical Insight & Conclusion

Why it Works

The "secret sauce" is the pre-processing. By using Fuzzy Rough Sets, the system handles the uncertainty of human emotion and slang in social media before the heavy semantic lifting begins. The multi-level inheritance in their ontology allows the system to find documents that are semantically related even if they share zero identical keywords with the query.

Limitations & Future Work

While performance is high, the construction of the initial "Knowledge Base" and "Rule Manager" remains a potential bottleneck for general-purpose application. The authors suggest that moving toward Intelligent Multi-Agent Systems could further decentralize and speed up this processing for truly global-scale distributed retrieval.

Final Takeaway: This paper is a blueprint for developers building sentiment-aware search engines. It proves that combining the logic of Ontologies with the statistical power of LDA, filtered through Fuzzy sets, is the most robust way to handle the "V"s (Volume, Velocity, Variety) of Big Data.

Find Similar Papers

Try Our Examples

  • Find recent papers that combine Fuzzy Rough Set theory with Deep Learning for feature selection in high-dimensional text data.
  • What is the origin of the SOR (Semantic Ontology Retrieval) algorithm and how has it evolved since its 2017 proposal?
  • Explore current research applying the IOLDAIR framework or similar ontology-LDA hybrids to real-time sentiment analysis in multimedia big data.
Contents
IOLDAIR: Bridging the Semantic Gap in Big Data Retrieval via Fuzzy Ontologies
1. TL;DR
2. Background & Positioning
3. Motivation: The Noise and Latecy Problem
4. Methodology: The IOLDAIR Architecture
4.1. 1. Intelligent Fuzzy Rough Set Feature Selection
4.2. 2. Semantic Analysis and IOLDAIR
5. Experimental Performance: SOTA Results
5.1. Performance Metrics
6. Critical Insight & Conclusion
6.1. Why it Works
6.2. Limitations & Future Work