Decoding the Data Warehousing Process: An Ontology-Based Quest for Standards
14137_Toward Developing Data Warehousing Process Standards An Ontology-Based Review of Existing Methodologies.
The paper presents an ontological framework to analyze and categorize 30 commercial data warehousing process (DWP) methodologies. By developing a dual-hierarchy ontology (composition and classification), the researchers identify common industry patterns and group methodologies using hierarchical cluster analysis to provide a basis for future platform-independent DWP standards.
TL;DR
In a landscape cluttered with over 30 proprietary vendor methodologies (IBM, Oracle, Microsoft, etc.), this paper introduces an ontological model to categorize the Data Warehousing Process (DWP). By analyzing these methodologies through the lens of task decomposition and clustering, the authors reveal the "hidden" standards of the industry—identifying the dominant reign of the Inmon vs. Kimball paradigms and providing a roadmap for platform-independent DWP standards.
Background: The "Tower of Babel" in Data Warehousing
Data warehousing is the backbone of modern decision support, yet its construction (the DWP) remains a Wild West. Every vendor claims a superior methodology, but these are often platform-dependent and siloed. Previous attempts at standardization, such as the Common Warehouse Metamodel (CWM), focused primarily on how tools "talk" (metadata interchange) rather than the "how-to" of the process itself.
The authors argue that before we can build a universal standard, we must first map the "DNA" of existing practices.
Methodology: Building the DWP Ontology
The researchers utilized Protégé to build a dual-hierarchy ontological model:
- Composition Hierarchy: Breaking down DWP into subtasks (e.g., ETL Extract, Transform, Load).
- Classification Hierarchy: Listing the specific flavors of these tasks (e.g., Logical Design Relational, Dimensional, or Object-Oriented).
The Structural Backbone
The authors validated their model with 15 seasoned industry experts, mapping out tasks from Business Requirements to Change Management.
Figure 2: The refined and validated DWP composition hierarchy used as the basis for the study.
Patterns in the Chaos: Cluster Analysis
By coding 30 methodologies against 50+ ontological attributes, the authors performed Hierarchical Cluster Analysis. The results provided a stark visualization of how the industry has self-organized around specific "schools of thought."
Figure 4: Dendogram representing the natural grouping of methodologies based on data modeling techniques.
Key Findings & Industry Anchors
- The Kimball School (Cluster 1): Methodologies like Cognos and Oracle follow a Data Mart architecture using Dimensional Modeling and the Dimensional Business Lifecycle. This is a bottom-up, requirements-driven approach.
- The Inmon School (Cluster 2): Giants like IBM, SAP, and Sybase favor a Centralized Data Warehouse with dependent data marts, primarily using a Data-Driven Iterative (Spiral) approach.
- Technique Synergy: The research found that Joint Application Development (JAD) is the "soulmate" of iterative approaches for large-scale enterprise warehouses, while Interviews are the standard for smaller-scope data marts.
Deep Insight: The ETL Frontier
One of the most valuable segments of the paper is the breakdown of ETL Standards. The authors identified four distinct clusters of practice:
- Deferred Extraction: Dominant among methodologies like SAP and Kimball, relying on timestamping.
- Immediate Extraction: Preferred by Sagent and Informatica via database triggers, allowing for near real-time updates.
Table VI: Standard practices for ETL across different methodology groups.
Critical Analysis & Future Outlook
The paper successfully demonstrates that while DWP methodologies are diverse, they are not random. They cluster around the scope of the project:
- Broad Scope (Enterprise) Centralized Architecture + Relational Modeling + Iterative Process.
- Narrow Scope (Departmental) Data Mart Architecture + Dimensional Modeling + Phased Lifecycle.
Limitations: The study is a snapshot of commercial tools available in the mid-2000s. It precedes the explosion of NoSQL, Data Lakes, and Cloud-Native warehouses which introduce new paradigms like ELT (Extract-Load-Transform) instead of ETL.
Conclusion: This research provides the first empirical "periodic table" of DWP. It suggests that future standardization must not be one-size-fits-all but must cater to the distinct needs of Enterprise vs. Departmental data strategies.
