Turning Data Trash into Troubleshooting Gold: The Power of Ontology-Supported Refurbishing
An ontology-supported database refurbishing technique and its application in mining actionable troubleshooting rules from real-life databases
The paper introduces an ontology-supported database refurbishing technique combined with an enhanced association rule-mining algorithm (MSapriori) to automate troubleshooting in the GSM domain. By leveraging domain ontologies and text mining, the system converts unstructured remark fields into structured, actionable data, significantly outperforming traditional mining on original "noisy" databases.
Executive Summary
TL;DR: This paper presents a sophisticated framework for transforming messy, real-world customer service logs into a high-precision troubleshooting rule base. By utilizing domain-specific ontologies and a modified association rule algorithm, the authors overcome the "noisy data" and "rare item" problems that typically plague industrial data mining.
Positioning: This work serves as a practical bridge between Knowledge Engineering (Ontologies) and Data Mining (Association Rules), demonstrating that domain knowledge isn't just a byproduct of mining, but a prerequisite for quality results in noisy environments.
The "Other" Problem: Why Your CRM Data is Useless
If you've ever looked at a corporate database, you know the pain of the "Other" category. Under pressure, customer service agents often categorize complex technical issues as "Other" to save time, while burying the actual solution in unstructured, natural language "Remark" fields.
For a data miner, this is a catastrophe. Standard algorithms like Apriori see a sea of "Other" vs. "Other" associations, which offer zero actionable value. The authors identify two core challenges:
- Imprecise Categorization: Valuable information is trapped in text, not in the database fields.
- The Rare Item Paradox: Critical technical faults happen rarely but are highly predictable. If you set your mining threshold high, you miss them; set it low, and you get overwhelmed by "noise" from frequent, trivial events.
Methodology: The Art of Refurbishing
The core innovation is Database Refurbishing. Instead of just "cleaning" (deleting bad data), the authors "refurbish" (restore) it using three domain-specific ontologies: Symptoms (GSMSO), Causes (GSMCO), and Processes (GSMPO).
1. The Refurbishing Pipeline
The system uses a Text-Mining-Empowered approach to salvage the "Other" records:
- Remark Retriever: Groups natural language remarks based on their associated (even if broad) categories.
- Keyword Extractor: Uses the MMSEG algorithm for Chinese word segmentation, then filters keywords through the ontology to ensure they have physical meaning in the GSM domain.
- TFIDF Weighting: Scores and selects the most "significant" keywords to replace the generic "Other" labels.

2. Hunting for Rare Items with MSapriori
The authors don't just use standard Apriori. They implement MSapriori (Multiple Minimum Supports), but with a twist. They introduce a new parameter, , to modify the minimum item support (MIS) for multi-item sets. This ensures that a rare (but important) symptom doesn't get pruned just because it appears alongside a common one.
Experimental Proof: From 58% to 81%
Working with over 6,000 real-life records from a telecom carrier, the results were striking. By refurbishing the database, the system successfully re-labeled thousands of records that were previously "dead data."

Key Metrics:
- Inference Accuracy: Jumped from 58.29% (original DB) to 81.95% (refurbished DB).
- Rule Quantity: Discovered 3x more cause-process rules than previous benchmarks.
- Expert Validation: The use of a "Rule Verifier" ensured the mined rules were free of common logic anomalies like unreachability or conflict.
Critical Insight & Conclusion
The true value of this paper lies in its Inductive Bias. Instead of assuming the data is correct, the authors treat the database as a "corrupted signal" and use Ontologies as a "low-pass filter" to extract the true underlying logic.
Limitations: The system relies heavily on a manually constructed ontology. In today's landscape, one might wonder if a Large Language Model (LLM) could automate this ontology construction; however, this paper’s rigorous framework for "verification and validation" remains a gold standard for mission-critical systems where hallucinations are not an option.
Future Outlook: As we move toward more autonomous systems, the "Database Refurbishing" concept will be essential for training AI on historical "dark data" that corporations have been hoarding for decades.
