SetConv & OAML: Decoding Stealthy APT Tactics via Imbalanced Graph Learning

Identifying Tactics of Advanced Persistent Threats with Limited Attack Traces

2021-01-01
Khandakar Ashrafi Akbar, Yigong Wang, Md Shihabul Islam, Anoop Singhal, Latifur Khan, Bhavani Thuraisingham
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a graph-based machine learning framework to identify Advanced Persistent Threat (APT) tactics using GraphSAGE for embeddings and novel classifiers like SetConv and OAML. It achieves state-of-the-art performance, reaching up to 94.49% accuracy on highly imbalanced provenance datasets by leveraging domain-specific log filtration.

TL;DR

Advanced Persistent Threats (APTs) are the "shadows" of the cyber world—long-term, stealthy, and extremely difficult to pin down. This paper presents a sophisticated pipeline that transforms system logs into Provenance Graphs, filters them using NLP-based domain knowledge, and employs a revolutionary SetConv architecture to identify attack tactics even when malicious data is extremely scarce.

Background Positioning: This work moves beyond simple "anomaly detection" into "tactic identification" (e.g., Persistence vs. Privilege Escalation), filling a gap between raw log collection and high-level kill-chain analysis.

The "Needle in a Haystack" Problem

Detecting APTs is traditionally a losing game for two reasons:

  1. Data Imbalance: In a real-world system, 99.9% of events are benign. Traditional ML models get "lazy" and simply predict everything as benign to achieve high accuracy.
  2. Noise: Modern system logs are massive. Investigating every read, write, or socket call is computationally prohibitive.

The authors' insight? Use Domain Knowledge to pre-filter the logs. By creating "Base Sentences" for specific TTPs (Techniques, Tactics, and Procedures), they can calculate similarity scores and keep only the parts of the log that actually "look" like an attack.

Methodology: From Traces to Graph Embeddings

The system architecture follows a clean, modular flow:

1. Provenance Graph Construction

Raw logs (from Sysdig) are converted into nodes (Processes, Files, Sockets) and edges (System Calls). Unlike basic graphs, these include edge weights representing the volume of data transferred (bytes read/written), adding a layer of behavioral intensity to the model.

2. GraphSAGE & Aggregation

To turn a complex graph into a mathematical vector, the authors use GraphSAGE. This inductive approach learns how to aggregate neighbor information, allowing the model to handle new, unseen nodes—a must for dynamic enterprise environments.

System Architecture Figure 1: The overall architecture from attack generation to ML classification.

3. The Secret Sauce: SetConv

Standard convolutions work on grids. SetConv operates on sets of features. It uses Episodic Training, where the model is shown "episodes" that preserve the imbalance ratio.

  • It learns a "Class Representative" for each tactic.
  • During inference, a query sample is simply compared against these class anchors.
  • This ensures the classifier isn't biased by the sheer volume of benign data.

SetConv Operation Figure 2: The SetConv training and inference procedure.

Experimental Battleground

The model was tested against two primary datasets: DAPT 2020 and a custom Graph Dataset.

Hard-Won Results

The "Re-sampling" setup (improving the encoder) showed that SetConv is exceptionally dominant. It achieved a Macro F1-score of 0.9406, significantly outperforming traditional Support Vector Machines (SVM) and standard Multilayer Perceptrons (MLP).

ModelAccuracyMacro F1
SVM0.90630.8963
MLP0.93570.9357
SetConv0.94490.9406

Note: Macro F1 is used here because it treats all classes (tactics) as equally important, preventing the majority "Benign" class from masking poor performance on rare "Attack" classes.

Critical Analysis & Future Outlook

Strengths: The use of NLP (Sentence Encoders) to filter logs for "Tactical keywords" is a brilliant bridge between cybersecurity domain expertise and machine learning. It provides a way to "crop" the history to only the relevant parts of an attack.

Limitations: The current approach relies on predefined keywords. If an attacker develops a completely novel technique that doesn't use those keywords, the pre-filtration step might accidentally discard the evidence.

Future Work: The authors suggest moving toward a shifting window-based approach, which would allow the model to monitor system logs in real-time, effectively sliding across the timeline to detect an attack as it unfolds, rather than analyzing post-mortem traces.

Conclusion

This research proves that with the right graph representations and a smart approach to class imbalance (SetConv), we can identify specific adversarial tactics with high precision. For security operations centers (SOCs), this means fewer false positives and a much clearer picture of what an intruder is actually trying to accomplish.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) or GraphSAGE specifically for MITRE ATT&CK tactic classification in enterprise logs.
  • What is the original theoretical basis for Set Convolution (SetConv) as an operator for imbalanced learning, and how does it compare to standard SMOTE techniques?
  • Explore if there are studies applying Universal Sentence Encoders or NLP embeddings to filter system call sequences before provenance graph generation in cybersecurity.
Contents
SetConv & OAML: Decoding Stealthy APT Tactics via Imbalanced Graph Learning
1. TL;DR
2. The "Needle in a Haystack" Problem
3. Methodology: From Traces to Graph Embeddings
3.1. 1. Provenance Graph Construction
3.2. 2. GraphSAGE & Aggregation
3.3. 3. The Secret Sauce: SetConv
4. Experimental Battleground
4.1. Hard-Won Results
5. Critical Analysis & Future Outlook
6. Conclusion