Scaling Knowledge: A Fast, SNA-Driven Approach to Engineering Ontology
A fast and economic ontology engineering approach towards improving capability matching: Application to an online engineering collaborative platform
This paper proposes a fast and economic ontology engineering methodology based on the SENSUS approach and Social Network Analysis (SNA). By leveraging search engine indices (Google Sets) and a "Snowball Sampling" technique, the method automatically generates domain-specific ontologies to improve need-and-resource matching for engineering collaborative platforms like WMCCM.
TL;DR
Building ontologies usually requires months of expert labor. This paper introduces a radical shortcut: using Google Sets and Social Network Analysis (SNA) to automatically "snowball" a handful of seed words into a massive, weighted network of domain concepts. When applied to the West Midlands Collaborative Commerce Marketplace (WMCCM), this method boosted automated tender matching accuracy from 51% to 77%.
The "Expert Bottleneck" in Matching Platforms
Online collaborative platforms live or die by their ability to match a business's "Needs" (tenders) with a company's "Resources" (capabilities). The bridge between these two is an Ontology.
However, traditional ontologies face a triple threat:
- Cost: They require constant input from expensive domain experts.
- Rigidity: Standard codes like SIC or UNSPSC are too high-level and lack "fuzzy" semantic links.
- Speed: They cannot adapt quickly to emerging technical terms in the engineering sector.
The authors argue that if we treat the web's collective linguistic patterns as a social network, we can extract a domain's structure without manual labor.
Methodology: The Snowball and the Network
The core innovation lies in a two-step process: Corpus Construction and Network Analysis.
1. Snowball Sampling via Google Sets
Instead of asking experts to list every possible engineering term, the researchers asked for just three pairs of "seeding words" (e.g., Drilling & Cutting). Using these as queries in Google Sets, the system retrieved semantically related clusters. By re-pairing these new results and querying again—a process called Snowball Sampling—they rapidly expanded a few seeds into a corpus of over 10,000 terms.
2. The Three Zones of Knowledge
Identifying terms is easy; structuring them is hard. The authors applied SNA to determine the "social position" of each word within the network:
- Definition Zone (The Core): Highly centralized terms (e.g., Machining, Welding) that define the domain.
- Description Zone (The Flesh): Terms that describe specific sub-processes (e.g., Gun Drilling, CNC Machining), connecting directly to the core.
- Connection Zone (The Boundary): Low-centrality "long tail" terms that provide the necessary fuzziness for natural language processing (e.g., Food Processing in an engineering context).
Figure 1: The proposed development cycle, moving from expert seeds to automated network analysis.
Mathematizing Semantic Closeness
Unlike simple taxonomies, this method calculates Directional Weights. Using a "Closeness" formula, the system determines the decisive power one word has over another. For instance, the term Centering might be related to both Turning and Milling, but the degree of relationship is quantified as a vector, allowing the matching algorithm to handle ambiguity more effectively.
Experimental Results: Real-World Impact
The researchers tested their SNA-derived ontology against the existing WMCCM system, which relied on SIC codes and manual expert refinement.
Figure 2: The new methodology significantly outperformed the traditional expert-led approach in both data capture and categorization accuracy.
Key Metrics:
- Trigger Rate: 91% of tenders were recognized (up from 82%).
- Accuracy: Correct categorization jumped to 77% (up from 51%).
- Richness: The automated method found 100x more internal relationships than the manual one.
Critical Insight & Conclusion
The true value of this research is the shift from Ontology as a Hierarchy to Ontology as a Weighted Network. By accepting "noise" in the connection zone, the system gains the "fuzziness" required to understand how humans actually describe engineering tasks in tenders.
While the reliance on Google Sets (now part of Google Sheets) highlights a dependency on external search indices, the underlying logic—using SNA centrality to automatically structure web-crawled data—remains a powerful template for building economic and scalable knowledge systems.
