Beyond Numbers: Integrating Ontology and Trust into the Next Generation of Business Intelligence
Ontology and Trust based Data Warehouse in New Generation of Business Intelligence
The paper proposes a next-generation Business Intelligence (BI) framework that integrates structured organizational data with unstructured external data from social media. It introduces a dual-layer approach using Ontology-based semantic enrichment and a Trust-based evaluation metric to ensure the reliability and meaningfulness of external information like "Voice of the Customer" (VoC).
TL;DR
Modern Business Intelligence (BI) is hitting a wall by ignoring the "Voice of the Market" hidden in unstructured social data. This paper proposes a transformation of the traditional Data Warehouse into a "New Generation BI" system. It solves the twin problems of data unreliability and semantic ambiguity by layering a Trust-evaluation algorithm and an Ontology-based enrichment engine over the standard ETL pipeline.
The "Dark Matter" of Data: The Problem
While organizations have become experts at analyzing their own transactional SQL databases, they remain blind to approximately 80% of the world's data: the unstructured text from Twitter, Facebook, and blogs.
The authors identify two critical reasons why traditional BI fails to capitalize on this:
- The Context Gap: A word like "Coles" or "Australia" means nothing to a computer without a semantic map (Ontology) to explain relationships (is it a location, a company, or a sentiment?).
- The Credibility Gap: Social media is noisy. A tweet from a bot or a new account carries less weight than a post from a verified industry influencer. Traditional ETL tools treat all data strings as equal, leading to "Garbage In, Garbage Out."
Methodology: The Trust-Ontology Architecture
The researchers propose a 5-stage lifecycle for Big Data Analysis that shifts the focus from simple storage to intelligent assimilation.
1. The Trust Metric (CCCI)
Rather than accepting all external data, the system evaluates the "Source Reputation." By analyzing user attributes—such as follower count, account age, and the "Impact Factor" of their past posts—the system assigns a Trust Value to each data source. This ensures that the insights generated have a "Confidence Level" attached to them.
2. Ontological Enrichment
Instead of building a database from scratch, the system utilizes an Ontology Repository.
- Tokenization: Raw text is broken down into parts.
- Mapping: Words are matched against classes (concepts) and properties in a domain-specific ontology.
- Inference: By using rules (e.g., "If words X and Y appear, it is an 'Event'"), the system extracts structured RDF triples from unstructured sentences.
Figure 1: The proposed workflow for extracting and validating unstructured data before it hits the landing area.
Experiments & Results: Turning Tweets into Facts
The authors demonstrated the framework using two case studies: General hashtags (#Australia) and Corporate monitoring (#Coles).
- Fine-Grained Filtering: Using a "Travel/Job" ontology, the system could distinguish between a tweet about a vacation and a tweet about an employment opportunity, even if both used the same general hashtag.
- Integration with Fact Tables: The study showed how these "Social Dimensions" (Trust, Sentiment, Topic) could be joined with traditional "Sales Dimensions" in a star schema.
Figure 2: The schema integration showing how social insights become part of the traditional fact-and-dimension model.
Critical Analysis & Conclusion
Takeaway
The core contribution of this work is the realization that Data Integrity in the age of Social BI is not just about avoiding "Null" values; it is about Provenance (Trust) and Semantics (Ontology). By enriching raw text before it enters the warehouse, businesses can finally listen to the "Voice of the Customer" with mathematical rigor.
Limitations & Future Work
While the framework is robust, the paper relies heavily on manually defined ontologies and rules. In the modern era of 2026, the next step would be moving toward Self-Evolving Ontologies or using LLMs to perform the mapping stage automatically. Additionally, the temporal factor of trust—how a source's reputation changes over time—remains an area for deeper empirical study.
Keywords: Social Business Intelligence, Data Warehouse, ETL, Trust Metrics, Ontology.
