Unifying Chaos: Ontology-Based Data Reorganization for Efficient Open Sharing
An Ontology-Based Data Organization Method
This paper presents a hybrid data organization method that integrates relational databases with domain ontologies to facilitate open data sharing. By mapping relational schemas to an ontology knowledge base and employing human-computer interaction for model refinement, the authors achieve a 79.9% reduction in data size and a 90% improvement in query performance.
TL;DR
The explosion of Big Data has led to a paradox: we have more data than ever, but its "shared" forms are often unusable due to rigid structures or over-simplified statistics. This paper introduces an Ontology-Based Data Organization Method that transforms messy, schema-dependent relational data into a semantic-rich knowledge base. By aligning data with real-world concepts, the authors achieved an 80% reduction in storage and a 10x speedup in query performance on medical datasets.
Problem & Motivation: The "Data-Driven" Paradox
In traditional environments (OLTP/OLAP), the people who build the database are the ones who use it—the structure follows the application. However, in Open Data Sharing, the users are unpredictable and often lack technical database knowledge.
The authors identify two critical pain points:
- Structural Inconsistency: Even within the same domain (e.g., hospitals), different providers use different table schemas.
- High Cognitive Load: Users must know exact field names and complex foreign key joins to extract value.
The insight? Use Ontology—a model directly related to the real world—to "reorganize" the underlying data so that the logical structure matches human intuition rather than just machine efficiency.
Methodology: From Relational Tables to Knowledge Trees
The proposed method doesn't just put a "label" on data; it fundamentally restructures it through a three-step pipeline.
1. Automated Extraction
The system views the database schema as a directed graph . It applies heuristic rules to convert tables into classes and foreign keys into hierarchical relations.
- Rule 1: Each table node becomes an ontology class.
- Rule 3/4: Foreign key directions dictate the "Parent-Child" hierarchy in the concept tree.
2. Human-in-the-Loop Refinement
Since automated extraction can't capture human nuances, the paper introduces three standardized operations:
- Select: Pruning noisy or non-conceptual tables.
- Add: Incorporating domain-specific constraints (e.g., valid ranges for lab results).
- Semantic Extension: Building a lexicon of synonyms (e.g., mapping "Thyroid Stimulating Hormone" to "TSH") to enable natural language-like queries.

3. Physical Data Reconstruct
This is where the performance gains happen. Instead of just a virtual view, the authors re-materialize the data:
- Relation Merging: Tables mapping to the same concept are joined physically to reduce future JOIN overhead.
- Column Filtering: Removing "system noise" fields (e.g., serial numbers, equipment IDs).
- Row Filtering: Using ontology range constraints to purge irrelevant records.
Experiments & Results: Efficiency Gains
The authors tested their framework on a 8.94GB thyroid disease dataset from a Chinese hospital.
Model Quality
The "Initial Model" (automated) covered 75% of domain terms. After just one round of human interaction, coverage jumped to 95%, proving the efficiency of the "Extract-then-Refine" approach over building from scratch.
Quantitative Performance
The physical reorganization led to dramatic improvements:
- Storage: 8.94GB 1.79GB (-79.9%).
- Query Speed: Average runtime across typical medical queries dropped by 90%.

The speedup is attributed to two factors: the drastically reduced table size (less I/O scanning) and the pre-merging of tables which eliminated expensive SQL JOIN operations during runtime.
Critical Analysis & Conclusion
Takeaway
The core value of this work lies in its Data Reconstruct phase. While many ontology papers focus on the "semantic layer," this paper demonstrates that using that semantic knowledge to physically prune and merge datasets can yield massive performance dividends for Big Data applications.
Limitations & Future Work
While the query interface is "user-friendly," it still relies on a Name Parser and Relation Parser that find matches in a lexicon. The authors acknowledge that the next step is moving toward full Natural Language Processing (NLP) integration, allowing users to ask questions in plain English rather than structured "ontology-style" queries.
In an era where "Open Data" is often synonymous with "Useless Data," this ontology-driven approach provides a scalable roadmap for making big datasets truly accessible to the domain experts who need them most.
