RBA: Bridging the Gap Between Symbolic Traces and Continuous Analytics
A reference based analysis framework for analyzing system call traces
This paper introduces Reference Based Analysis (RBA), a data mining framework that transforms complex, discrete system call traces into a multivariate continuous representation. By mapping sequence data into a vector space based on a reference "normal" dataset, RBA enables effective visualization and state-of-the-art anomaly detection for host-based intrusion detection systems.
TL;DR
The paper introduces Reference Based Analysis (RBA), a framework that converts discrete, symbolic system call traces into a multivariate continuous format. This allows security analysts to "see" intrusions through visualization and use standard nearest-neighbor algorithms to achieve superior detection accuracy compared to traditional Hidden Markov Models or Finite State Automata.
Background: The "Invisibility" of System Calls
In host-based intrusion detection, we monitor the system calls a process makes (e.g., open, read, mmap). However, these traces are problematic:
- No Metric: Is
opencloser toreadorclose? There is no inherent distance. - Variable Length: One execution might yield 100 calls, another 10,000.
- Scale: Traces are too long for human eyes to spot subtle malicious deviations.
The authors' insight is to stop looking at the traces in isolation and instead look at them through the lens of a "normal" reference set.
Methodology: The RBA Transformation
The core of RBA is a feature extraction process that maps a sequence into a fixed-length vector.
- Sliding Windows: Break the sequence into -length windows.
- Referential Counting: For each window, find how often it (and its prefix) appeared in a known "normal" dataset.
- 2D Histogramming: These two frequencies ( and ) form coordinates. By binning these coordinates into a 2D histogram, the method captures the "surprisingness" of transitions.
- Flattening: The histogram is flattened into a continuous vector.
This creates a vector space representation where "normal" traces cluster together because they share frequency patterns relative to the reference set, while "intrusive" traces deviate into different regions of the manifold.
Figure: The RBA mapping effectively separates normal (reference) sequences from intrusive ones in the snd-cert dataset.
Experiments and Results
The authors tested RBA against the UNM (University of New Mexico) and DARPA BSM datasets. They used a simple nearest-neighbor anomaly detection technique on the RBA-generated features.
Key Findings:
- Visualization Power: As seen in the PCA plots, RBA reveals clear clusters of normal behavior and distinct outliers for attacks. For example, in the
snd-certdataset, the separation is nearly perfect. - Higher Precision: RBA achieved an average precision of 0.78, significantly outperforming HMMs (0.07) and standard RIPPER-based rules (0.48).
Table: Comparison of Precision. RBA (bottom row) consistently leads across different data environments.
Critical Analysis
Why does it work?
RBA succeeds because it converts a structural problem (sequence pattern matching) into a density problem (finding outliers in continuous space). By using the prefix frequency as a baseline for the window frequency, it implicitly captures the conditional probability of system calls without the computational overhead of training a full Markov Model.
Limitations
- Reference Dependency: The quality of the "normal" dataset is critical. If the reference set is "polluted" with attacks, the transformation becomes less effective.
- Granularity: The choice of window size and binning resolution are hyperparameters that likely require tuning for different operating systems.
Conclusion
RBA is a robust example of how dimensionality reduction and referential mapping can simplify complex data types. For cybersecurity, it transforms the "black box" of system call logs into an interpretable map, allowing for both automated detection and human-in-the-loop forensic analysis.
