Unifying Chaos: Ontology-Based Data Reorganization for Efficient Open Sharing

An Ontology-Based Data Organization Method

2017-08-01
Qian Hao, Yue Li, Li-Min Wang, Mei Wang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a hybrid data organization method that integrates relational databases with domain ontologies to facilitate open data sharing. By mapping relational schemas to an ontology knowledge base and employing human-computer interaction for model refinement, the authors achieve a 79.9% reduction in data size and a 90% improvement in query performance.

TL;DR

The explosion of Big Data has led to a paradox: we have more data than ever, but its "shared" forms are often unusable due to rigid structures or over-simplified statistics. This paper introduces an Ontology-Based Data Organization Method that transforms messy, schema-dependent relational data into a semantic-rich knowledge base. By aligning data with real-world concepts, the authors achieved an 80% reduction in storage and a 10x speedup in query performance on medical datasets.

Problem & Motivation: The "Data-Driven" Paradox

In traditional environments (OLTP/OLAP), the people who build the database are the ones who use it—the structure follows the application. However, in Open Data Sharing, the users are unpredictable and often lack technical database knowledge.

The authors identify two critical pain points:

  1. Structural Inconsistency: Even within the same domain (e.g., hospitals), different providers use different table schemas.
  2. High Cognitive Load: Users must know exact field names and complex foreign key joins to extract value.

The insight? Use Ontology—a model directly related to the real world—to "reorganize" the underlying data so that the logical structure matches human intuition rather than just machine efficiency.

Methodology: From Relational Tables to Knowledge Trees

The proposed method doesn't just put a "label" on data; it fundamentally restructures it through a three-step pipeline.

1. Automated Extraction

The system views the database schema as a directed graph . It applies heuristic rules to convert tables into classes and foreign keys into hierarchical relations.

  • Rule 1: Each table node becomes an ontology class.
  • Rule 3/4: Foreign key directions dictate the "Parent-Child" hierarchy in the concept tree.

2. Human-in-the-Loop Refinement

Since automated extraction can't capture human nuances, the paper introduces three standardized operations:

  • Select: Pruning noisy or non-conceptual tables.
  • Add: Incorporating domain-specific constraints (e.g., valid ranges for lab results).
  • Semantic Extension: Building a lexicon of synonyms (e.g., mapping "Thyroid Stimulating Hormone" to "TSH") to enable natural language-like queries.

Process of Building Ontology Model

3. Physical Data Reconstruct

This is where the performance gains happen. Instead of just a virtual view, the authors re-materialize the data:

  • Relation Merging: Tables mapping to the same concept are joined physically to reduce future JOIN overhead.
  • Column Filtering: Removing "system noise" fields (e.g., serial numbers, equipment IDs).
  • Row Filtering: Using ontology range constraints to purge irrelevant records.

Experiments & Results: Efficiency Gains

The authors tested their framework on a 8.94GB thyroid disease dataset from a Chinese hospital.

Model Quality

The "Initial Model" (automated) covered 75% of domain terms. After just one round of human interaction, coverage jumped to 95%, proving the efficiency of the "Extract-then-Refine" approach over building from scratch.

Quantitative Performance

The physical reorganization led to dramatic improvements:

  • Storage: 8.94GB 1.79GB (-79.9%).
  • Query Speed: Average runtime across typical medical queries dropped by 90%.

Query Runtime Comparison

The speedup is attributed to two factors: the drastically reduced table size (less I/O scanning) and the pre-merging of tables which eliminated expensive SQL JOIN operations during runtime.

Critical Analysis & Conclusion

Takeaway

The core value of this work lies in its Data Reconstruct phase. While many ontology papers focus on the "semantic layer," this paper demonstrates that using that semantic knowledge to physically prune and merge datasets can yield massive performance dividends for Big Data applications.

Limitations & Future Work

While the query interface is "user-friendly," it still relies on a Name Parser and Relation Parser that find matches in a lexicon. The authors acknowledge that the next step is moving toward full Natural Language Processing (NLP) integration, allowing users to ask questions in plain English rather than structured "ontology-style" queries.

In an era where "Open Data" is often synonymous with "Useless Data," this ontology-driven approach provides a scalable roadmap for making big datasets truly accessible to the domain experts who need them most.

Find Similar Papers

Try Our Examples

  • Search for recent studies that automate the mapping between Relational Databases (RDB) and Web Ontology Language (OWL) for Big Data integration.
  • Which paper first introduced the methodology of "Learning Ontology from Relational Databases" (Li et al., 2005), and how has this specific paper improved upon its heuristic rules?
  • Explore research applying ontology-driven data reorganization to heterogeneous multi-modal datasets beyond medical informatics, such as smart city or IoT data sharing.
Contents
Unifying Chaos: Ontology-Based Data Reorganization for Efficient Open Sharing
1. TL;DR
2. Problem & Motivation: The "Data-Driven" Paradox
3. Methodology: From Relational Tables to Knowledge Trees
3.1. 1. Automated Extraction
3.2. 2. Human-in-the-Loop Refinement
3.3. 3. Physical Data Reconstruct
4. Experiments & Results: Efficiency Gains
4.1. Model Quality
4.2. Quantitative Performance
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work