Beyond Numbers: Integrating Ontology and Trust into the Next Generation of Business Intelligence

Ontology and Trust based Data Warehouse in New Generation of Business Intelligence

Pornpit Wongthongtham, Abu Bilal, Salih
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes a next-generation Business Intelligence (BI) framework that integrates structured organizational data with unstructured external data from social media. It introduces a dual-layer approach using Ontology-based semantic enrichment and a Trust-based evaluation metric to ensure the reliability and meaningfulness of external information like "Voice of the Customer" (VoC).

TL;DR

Modern Business Intelligence (BI) is hitting a wall by ignoring the "Voice of the Market" hidden in unstructured social data. This paper proposes a transformation of the traditional Data Warehouse into a "New Generation BI" system. It solves the twin problems of data unreliability and semantic ambiguity by layering a Trust-evaluation algorithm and an Ontology-based enrichment engine over the standard ETL pipeline.

The "Dark Matter" of Data: The Problem

While organizations have become experts at analyzing their own transactional SQL databases, they remain blind to approximately 80% of the world's data: the unstructured text from Twitter, Facebook, and blogs.

The authors identify two critical reasons why traditional BI fails to capitalize on this:

  1. The Context Gap: A word like "Coles" or "Australia" means nothing to a computer without a semantic map (Ontology) to explain relationships (is it a location, a company, or a sentiment?).
  2. The Credibility Gap: Social media is noisy. A tweet from a bot or a new account carries less weight than a post from a verified industry influencer. Traditional ETL tools treat all data strings as equal, leading to "Garbage In, Garbage Out."

Methodology: The Trust-Ontology Architecture

The researchers propose a 5-stage lifecycle for Big Data Analysis that shifts the focus from simple storage to intelligent assimilation.

1. The Trust Metric (CCCI)

Rather than accepting all external data, the system evaluates the "Source Reputation." By analyzing user attributes—such as follower count, account age, and the "Impact Factor" of their past posts—the system assigns a Trust Value to each data source. This ensures that the insights generated have a "Confidence Level" attached to them.

2. Ontological Enrichment

Instead of building a database from scratch, the system utilizes an Ontology Repository.

  • Tokenization: Raw text is broken down into parts.
  • Mapping: Words are matched against classes (concepts) and properties in a domain-specific ontology.
  • Inference: By using rules (e.g., "If words X and Y appear, it is an 'Event'"), the system extracts structured RDF triples from unstructured sentences.

System Architecture for Unstructured Data Figure 1: The proposed workflow for extracting and validating unstructured data before it hits the landing area.

Experiments & Results: Turning Tweets into Facts

The authors demonstrated the framework using two case studies: General hashtags (#Australia) and Corporate monitoring (#Coles).

  • Fine-Grained Filtering: Using a "Travel/Job" ontology, the system could distinguish between a tweet about a vacation and a tweet about an employment opportunity, even if both used the same general hashtag.
  • Integration with Fact Tables: The study showed how these "Social Dimensions" (Trust, Sentiment, Topic) could be joined with traditional "Sales Dimensions" in a star schema.

Integration Process Figure 2: The schema integration showing how social insights become part of the traditional fact-and-dimension model.

Critical Analysis & Conclusion

Takeaway

The core contribution of this work is the realization that Data Integrity in the age of Social BI is not just about avoiding "Null" values; it is about Provenance (Trust) and Semantics (Ontology). By enriching raw text before it enters the warehouse, businesses can finally listen to the "Voice of the Customer" with mathematical rigor.

Limitations & Future Work

While the framework is robust, the paper relies heavily on manually defined ontologies and rules. In the modern era of 2026, the next step would be moving toward Self-Evolving Ontologies or using LLMs to perform the mapping stage automatically. Additionally, the temporal factor of trust—how a source's reputation changes over time—remains an area for deeper empirical study.


Keywords: Social Business Intelligence, Data Warehouse, ETL, Trust Metrics, Ontology.

Find Similar Papers

Try Our Examples

  • Search for recent papers on "Semantic ETL" processes that combine Knowledge Graphs with traditional Data Warehousing.
  • What are the current SOTA methods for calculating "Source Trustworthiness" in social media beyond the CCCI metrics mentioned in this 2016 paper?
  • Explore how Large Language Models (LLMs) are currently replacing or augmenting traditional Ontology-based enrichment for unstructured data in Business Intelligence.
Contents
Beyond Numbers: Integrating Ontology and Trust into the Next Generation of Business Intelligence
1. TL;DR
2. The "Dark Matter" of Data: The Problem
3. Methodology: The Trust-Ontology Architecture
3.1. 1. The Trust Metric (CCCI)
3.2. 2. Ontological Enrichment
4. Experiments & Results: Turning Tweets into Facts
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work