FindMal: Leveraging File-to-File Social Networks for Robust Malware Detection

KNOWLEDGE‐BASED SYSTEMS

2024-01-10
Lieven Dubois, Philippe Mack
Summary
Problem
Method
Results
Takeaways
Abstract

FindMal is a graph-based malware detection framework that leverages a file-to-file social network and semi-supervised learning. By combining a k-Nearest Neighbors (kNN) relation graph, Label Propagation (LP), and an active learning strategy, it achieves state-of-the-art performance in classifying malicious and benign files.

TL;DR

FindMal is a novel framework that moves away from analyzing files in isolation. Instead, it treats files as nodes in a "social network" constructed from their co-occurrence patterns on user machines. By combining Label Propagation with Active Learning and specialized graph features, it detects malware that traditional content-based scanners might miss, achieving superior performance on massive, real-world datasets.

Problem & Motivation: The Limits of Isolation

Traditional malware detection relies on signatures or behavioral features extracted from the file itself (static/dynamic analysis). However, these methods struggle with:

  1. Polymorphism/Metamorphism: Malware that changes its binary structure to evade content-based detection.
  2. Zero-day Threats: New malware for which no signature exists.
  3. High Labeling Cost: Requiring human experts to manually analyze millions of files is unfeasible.

The authors' core insight is that malware rarely travels alone. Malicious files often co-exist with specific other files or variants (e.g., a downloader and its payload). By modeling these relationships as a graph, we can use "guilt-by-association"—if a file is closely related to a known malicious sample in the network, it is likely malicious itself.

Methodology: The FindMal Framework

FindMal operates in four distinct phases:

1. Graph Construction (kNN Relation Graph)

The system builds a graph where nodes are files. The edge weight is calculated using the Jaccard Similarity of the user machines where the files appear. To reduce noise, the authors use a k-Nearest Neighbors (kNN) approach, only keeping edges between the most similar files.

2. Graph-Based Feature Extraction

To identify which files are most "influential" or "representative" for labeling, the authors propose three metrics:

  • Degree Centrality (DC): Measures local connectivity.
  • Closeness Centrality (CC): Measures how fast a label can spread from this node to the whole graph.
  • Weighted Local Clustering Coefficient (LCC): Quantifies the "community" strength, helping identify clusters of malware variants.

Model Architecture Fig 1. Overview of the FindMal Framework

3. Label Propagation (LP)

Once a few nodes are labeled, the LP algorithm iteratively "pushes" these labels to neighbors based on edge weights. The probability of an unlabeled file being malicious is calculated by: where is the transition matrix and are the initial labels.

4. Active Learning with Maximum Network Gain

Instead of labeling random files, FindMal uses a Batch Mode Active Learning strategy. It selects a set of files that maximize "Network Gain" (NG)—a balance between selecting uncertain files (high entropy) and ensuring diversity/representativeness in the batch to avoid redundancy.

Experimental Performance

The framework was tested on a dataset of 69,165 real samples. Key findings include:

  • kNN Advantage: Using kNN to filter the graph improved the Recall by over 13% by removing noisy, weak associations.
  • SOTA Comparison: FindMal (F1: 62.5%) significantly outperformed established baselines like AESOP (F1: 10.9%) and SVM (F1: 20.3%). The standard SVM failed because it couldn't capture the topological information inherent in the graph.
  • Active Learning Efficiency: The Batch mode strategy showed a steeper performance gain curve compared to random sampling, proving that choosing the right nodes to label is as important as the detection algorithm itself.

Experimental Results Fig 2. Performance comparison between Label Propagation and Baselines

Critical Analysis & Conclusion

Takeaways

FindMal demonstrates that the context of a file's existence (who its neighbors are) is a robust signal against evasion techniques. The integration of active learning makes the system economically viable for anti-malware vendors.

Limitations

  • Homogeneity: Currently, the model only considers co-occurrence. It treats all relations as the same type.
  • Adversarial Graphs: Intelligent attackers might attempt to "pollute" the graph by making malware co-exist with popular benign files (e.g., chrome.exe) to lower their malicious probability.

Future Work

The authors suggest moving toward Heterogeneous Graphs to incorporate file-to-archive or file-to-machine relations and exploring more complex non-linear models like GNNs to fuse content features with relational data.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize Heterogeneous Information Networks (HIN) for malware detection to capture multiple relationship types beyond co-occurrence.
  • Which paper first proposed the "Guilt by Association" principle in large-scale malware detection, and how does the FindMal framework specifically refine its node-ranking mechanism?
  • Explore how Graph Neural Networks (GNNs) have been applied to file-relation graphs to replace or enhance traditional Label Propagation and hand-crafted centrality features.
Contents
FindMal: Leveraging File-to-File Social Networks for Robust Malware Detection
1. TL;DR
2. Problem & Motivation: The Limits of Isolation
3. Methodology: The FindMal Framework
3.1. 1. Graph Construction (kNN Relation Graph)
3.2. 2. Graph-Based Feature Extraction
3.3. 3. Label Propagation (LP)
3.4. 4. Active Learning with Maximum Network Gain
4. Experimental Performance
5. Critical Analysis & Conclusion
5.1. Takeaways
5.2. Limitations
5.3. Future Work