Deciphering LLMs for Web Security: Robustness and Key Logic in Malicious Traffic Detection
7789_Learning and Applying Ontology for Machine Learning in Cyber Attack Detection.
The paper investigates the robustness and feature sensitivity of Large Language Models (LLMs) in detecting malicious web traffic. It evaluates detection accuracy under various log formatting permutations and analyzes the impact of specific log attributes on model performance.
TL;DR
This study evaluates the effectiveness of Large Language Models (LLMs) in classifying malicious web traffic. It discovers that LLMs are surprisingly resilient to changes in log formats ("disordered" logs) and identifies that "Request" strings and "NPWI" features are the primary drivers of successful detection, achieving over 96% accuracy.
Background & Motivation: Beyond Rigid Parsing
Traditionally, log analysis requires strict adherence to formats like the Standard Apache log. If a log is malformed or some fields are swapped, traditional rule-based or machine-learning systems often fail. The authors set out to determine if LLMs, with their inherent understanding of natural language, can bypass these rigid constraints and identify the "intent" of a log entry regardless of its structure.
Methodology: Testing the Limits of Disorder
The researchers devised two main testing frameworks:
- Structural Robustness: They created five "Disorder" variants by shuffling the position of IP addresses, timestamps, and request bodies.
- Attribute Sensitivity: By isolating specific attributes (IP, Time, Request, Package Size, etc.), they measured which part of a log entry provides the most signal for "maliciousness."
Figure 1: Conceptual framework for evaluating LLM robustness against disordered inputs.
Key Insights from Experiments
1. High Robustness to Formatting
One of the most impressive findings is that LLMs do not "break" when logs are messy. Even in the worst-case disorder scenario, the accuracy only dropped by about 3.1%. This suggests the model is performing deep semantic analysis rather than simple pattern matching on fixed positions.
| Input Format | Accuracy | Error vs. Standard |
|---|---|---|
| Standard Apache | 0.962 | — |
| Disorder 3 (Max Shift) | 0.931 | -0.031 |
2. The Dominance of the "Request" Attribute
When the model was asked to classify logs based only on one attribute, the results were stark. The "Request" field (the actual URL and parameters) achieved 0.964 accuracy, which is actually higher than the full standard log in some cases. Conversely, the User-Agent was almost useless (0.1148 accuracy).
Figure 2: Performance breakdown by individual log attributes.
3. Feature Importance (NPWI vs. Metadata)
The study highlights NPWI as a critical feature. Combining Method (M), Status Code (S), and Package Size (P) only gets the model so far. Adding NPWI into the mix pushes the feature importance and model accuracy to near-perfect levels (1.0).
Figure 3: Ablation study showing the impact of combining different features.
Critical Analysis & Conclusion
The research confirms that LLMs are powerful tools for cybersecurity because they focus on the "content" (the Request) and are largely indifferent to the "container" (the Log Format).
Key Takeaways:
- Resilience: You don't need perfect log parsers to use LLMs for threat hunting.
- Efficiency: If compute is limited, focusing solely on the "Request" and "NPWI" features provides the highest ROI for detection.
Limitations: While the model is robust to disorder, the study does not deeply explore "adversarial" disorder—where an attacker intentionally crafts logs to mislead the LLM's semantic understanding. Future work should investigate if these models can be easily fooled by obfuscated payloads that mimic legitimate traffic.
