Automatic Ontology Generation: Taming the Chaos of Government Email Collections
Ontology generation for large email collections
The paper introduces an automated framework for generating concept ontologies from massive email collections, specifically targeting the U.S. government's "eRulemaking" process. By combining nominal N-gram mining, WordNet-guided restructuring, and supervised K-medoids clustering with pseudo-relevance feedback, the system achieves a state-of-the-art F1-measure of 0.94 on real-world datasets like the "Mercury" and "Polar Bear" rule comments.
TL;DR
Researchers at Carnegie Mellon University have developed a robust, fully automated system to transform massive, noisy email datasets into structured concept ontologies. By leveraging Nominal N-gram mining, WordNet, and a supervised clustering mechanism with pseudo-relevance feedback, they achieved an impressive 0.94 F1-measure in organizing complex public policy issues.
Context & Motivation: The eRulemaking Crisis
In the digital age, U.S. government agencies receive hundreds of thousands of public comments via email regarding proposed regulations (e.g., the "Polar Bear" or "Mercury" rules). By law, every substantive issue raised must be considered. However, the sheer volume of data makes manual inspection a bottleneck.
The authors identified that while tools for "duplicate detection" exist, we lack efficient ways to discover the "lay of the land"—the conceptual taxonomy of what people are actually talking about. Existing automatic methods often fail because email text is notoriously messy, filled with typos, and lacks the formal structure of academic papers or news.
Methodology: How to Build a Machine-Learned Hierarchy
The system operates through a sophisticated pipeline that moves from raw text to a "forest" of concepts, and finally to a unified tree.
1. Concept Mining and Web-Scale Cleaning
First, the system extracts nominal bigrams and trigrams. To combat the inevitable errors of POS taggers on messy email text, the authors used a clever Web-based heuristic: they queried Google for candidate concepts and discarded those that didn't appear frequently in snippets, effectively filtering out "gibberish" noun phrases like "bear bear."
2. From Fragments to a Forest
Using WordNet and head-noun matching, the system creates small, high-precision hierarchy fragments (e.g., grouping "water pollution" and "air pollution" under the parent "pollution").

3. Supervised Clustering (The Secret Sauce)
This is the core innovation. Usually, clustering is unsupervised. Here, the authors used the high-precision fragments already built as training data.
- Objective: Learn a similarity score function using Semi-Definite Programming (SDP).
- Mechanism: If two concepts are siblings in a small fragment, the model learns that their features imply a strong similarity. This knowledge is then "propagated" upward to cluster the roots of these fragments into even higher-level categories.
4. Automated Naming
To name the newly created parent clusters (which have no label), the system queries the Web with the child nodes (e.g., "Bush, Reagan") and extracts the most frequent common term (e.g., "President") from the search results.
Experimental Validation
The system was tested on two massive datasets of ~500,000 emails each.
| Component Added | Precision | Recall | F1 |
|---|---|---|---|
| Basic N-grams | 0.29 | 0.43 | 0.34 |
| + POS Corrector | 0.35 | 0.58 | 0.44 |
| + Supervised Clustering | 0.91 | 0.98 | 0.94 |
The "Ablation Study" (Table 4) shows that Supervised Clustering is the most critical component for precision, as it eliminates the "flatness" of unsupervised methods and enforces a logical hierarchy.

Critical Insight & Future Outlook
The genius of this work lies in Pseudo-Relevance Feedback. By assuming the lower-level linguistic relationships are "ground truth," the authors created a self-supervised loop that allows a machine to learn complex semantic abstractions without needing thousands of manually labeled examples.
Limitations:
- The system relies heavily on external search engine snippets, which might introduce bias or become a latency bottleneck.
- While effective for nouns, it may struggle with capturing "action-based" or "sentiment-based" ontologies.
Future Impact: This approach paves the way for "Self-Organizing Knowledge Bases" in any domain—from legal discovery to customer support—where the data is too large for humans but too messy for traditional rigid schemas.
