Ontology-Driven Data Semantics: Solving the "Ad Hoc" Log Bottleneck in Cyber-Security
Ontology-Driven Data Semantics Discovery for Cyber-Security
This paper introduces an ontology-driven architecture for data semantics discovery in cyber-security, designed to extract structured information from ad hoc, human-readable log files without predefined formats. By integrating Knowledge Representation with Machine Learning (SVMs), the system automatically identifies file formats, tokenizes data using genetic programming-based regex generation, and maps entities to a comprehensive cyber-asset ontology.
TL;DR
In the high-stakes world of cyber-forensics, the inability to parse unknown or changing log formats (Ad Hoc data) is a critical failure point. This paper presents a hybrid architecture that uses Machine Learning to discover the structure of unknown files and an Ontology to give those structures semantic meaning. The result? A system that can turn a pile of undocumented text files into a searchable, unified knowledge base with ~98% file identification accuracy.
Context: Why "Grep" is No Longer Enough
Security analysts often find themselves in "log hell"—navigating thousands of files from diverse nodes, each with unique configurations. Standard tools like grep or regular-expression-based searches are brittle; they fail the moment a software update changes a log's delimiter or an analyst needs to correlate a "Network Address" that appears as an IPv4 in one file and a MAC address in another. The fundamental challenge is the semantic gap between raw text and actionable intelligence.
Methodology: The Hybrid Intelligence Engine
The paper's architecture is a sophisticated pipeline that bridges machine learning with formal logic.
1. Structure Discovery (The "How")
Instead of manually writing parsers, the system uses a multi-stage ML approach:
- Novelty Detection: A One-Class SVM determines if a file is "known." If it’s new, it triggers the Template Generator.
- Genetic Programming: The system identifies delimiters (like whitespace or punctuation) and uses genetic algorithms to "evolve" the perfect regular expression that fits the observed data tokens.
- Feature Engineering: It uses n-grams of space-delimited tokens, but with a twist—it replaces alphanumeric characters with generic labels (e.g., "192.168.1.1" becomes "NNN.NNN.N.N") to focus on structural patterns rather than specific values.
2. The Ontology (The "Why")
The ontology acts as the brain of the operation. It defines the hierarchy of Events (OS events, hardware events) and Objects (IP addresses, processes).
- Abstraction: If an analyst searches for a
NetworkAddress, the ontology knows to look forIPv4,IPv6, andMACAddresssimultaneously. - Guided Learning: The ontology stores paths to training samples, allowing the ML components to "self-correct" based on domain knowledge.
Figure 1: The proposed architecture showing the interaction between ML modules and the domain ontology.
Experiments: Quantifying the Semantic Lift
The researchers tested the system on 2,022 text files. The performance metrics highlight a high degree of precision in high-level categorization:
| Task | Precision | Recall | F-Measure |
|---|---|---|---|
| File Format ID | 0.9791 | 0.9808 | 0.9799 |
| Record Type ID | 0.8350 | 0.8438 | 0.8394 |
| Entity Type ID | 0.8279 | 0.7819 | 0.8042 |
While File Format identification is near-perfect, the slightly lower scores for Entity Type identification (0.80) reflect the inherent ambiguity of short text strings—is "1024" a Port Number, a User ID, or a File Size? However, the authors argue that the Ontology's hierarchy mitigates this: even a misclassified low-level entity might still be correctly grouped under a parent class, preserving query relevance.
Figure 2: Example of a DNS Query Record being mapped into the hierarchical ontology.
Critical Insight: Semantic Resiliency
The true value of this work lies in Query Independence. Because the data is mapped to a high-level ontology, an analyst can write a SPARQL query to find a "malicious attachment arrival followed by a DNS tunnel" without knowing the specific format of the mail logs or the DNS logs on a particular node. The system handles the "translation" from raw text to conceptual event.
Limitations & Future Work
Despite its strengths, the system relies on the assumption that non-alphanumeric characters serve as delimiters. In modern "structured-within-unstructured" logs (like nested JSON within a syslog), this heuristic might struggle. The authors' plan to explore semi-supervised learning is a necessary next step to reduce the burden of providing labeled training samples for every new file type.
Conclusion
This paper is a significant step toward "Zero-Touch" forensics. By moving away from brittle, human-authored parsers and toward an ontology-driven discovery model, it provides a blueprint for managing the data explosion in large-scale network security.
