PM vs. VM: Is Virtualization Killing Your Malware Detection Performance?
Abstract: With the fast development of online education, the volume of education data traffic increased dramatically. Security information is potential to be mined from it. We can use data mining with some cloud computing platform for malware detection because the data volume is huge. The online education institutions need to virtualize their data centers and build cloud infrastructure for better using resources. So they should move data centers from physical machines(PMs) to virtual machines(VMs) for implementing the virtualization. But there are some risks such as the loss of computing ability, performance decline and so on. In this paper, we do a series of experiments to test performance of data mining algorithm based on Hadoop in physical machines and virtual machines. Through these experiments, we find that the performance of data mining algorithm based on Hadoop depends on disk I/O performance of Hadoop. The disk I/O performance of Hadoop deployed in PMs is better than that in VMs .Some iterative algorithms like k-means need more disk I/O, so we don't advise using VMs for computing. Other basic algorithms like Bayes classification need less disk I/O, so we advise computing in the VMs
This paper investigates the performance of Hadoop-based malware detection algorithms—specifically K-Means and Naive Bayes—across Physical Machines (PMs) and Virtual Machines (VMs). By analyzing traffic from educational networks, the study identifies critical performance bottlenecks in virtualized environments related to specific algorithm behaviors.
TL;DR
With the surge in online education data, malware detection has moved to the cloud. However, this study reveals a stark reality: migrating Hadoop-based data mining to Virtual Machines (VMs) can slow down iterative algorithms like K-Means by up to 600%, while simpler algorithms like Naive Bayes remain unaffected. The culprit? A massive discrepancy in disk writing speed between physical and virtual environments.
Context: The Push for Virtualized Education Networks
Educational institutions are rapidly adopting cloud infrastructures to manage the "Big Data" of online learning. While virtualization promises agility and better resource management, it introduces an abstraction layer that can severely hamper computational performance. This paper asks a critical question: Is it always worth migrating Hadoop clusters from Physical Machines (PMs) to VMs?
Problem & Motivation: The Hidden Cost of the Hypervisor
Existing research often treats "the cloud" as a homogenous resource. However, virtualization techniques (such as those used in CloudStack or VMware) introduce overhead in CPU scheduling and, more critically, in Disk I/O. For security tasks like malware detection, where datasets grow into the gigabytes, these overheads can turn a minutes-long detection task into an hours-long liability.
Methodology: A Fair Ground for Comparison
To ensure scientific rigor, the authors used a controlled environment where the PM and VM shared identical specifications:
- CPU: Two cores @ 2.5GHz
- Memory: 1GB
- Disk: 50GB
- Framework: Hadoop-1.2.1 (HDFS block size 64M)
They extracted 15 key features from real-world educational network traffic (e.g., HTTP response codes, attachment sizes, and packet ratios) to train two distinct types of models:
- K-Means: An iterative clustering algorithm (intensive read/write).
- Naive Bayes: A probabilistic classifier (minimal disk interaction).
Table: Hardware configuration parity between PM and VM.
The "Smoking Gun": Disk I/O Performance
The core insight of this paper lies in the benchmark of Hadoop's I/O throughput.
- Reading: Both PMs and VMs performed similarly (80-120 MB/s).
- Writing: PMs sustained 70-80 MB/s, while VMs plummeted to 10-20 MB/s.
Fig 5: The dramatic drop in write performance within Virtual Machines.
Experimental Results: Iteration is the Enemy
The impact of this write bottleneck is seen directly in the algorithm execution times:
- K-Means (The Loser in VM): Because K-Means is iterative, it frequently writes intermediate cluster centers and re-reads data. When traffic data reached 1GB, the execution time in the VM was 6 times longer than in the PM.
- Naive Bayes (The Winner in VM): Since Naive Bayes generally calculates probabilities in a single pass with minimal disk writes, the performance between PM and VM was nearly identical.
Fig 2: Running time comparison for K-Means—VM latency scales poorly with data volume.
Critical Insight & Conclusion
Takeaway
If your security pipeline relies on iterative algorithms (K-Means, iterative SVM, or deep learning with frequent checkpointing), keep your Hadoop cluster on Physical Machines. If your pipeline is primarily classification-based (Naive Bayes), migration to VMs is highly recommended to save costs and power without sacrificing speed.
Limitations & Future Work
The study utilizes Hadoop 1.2.1, which is now legacy. Modern frameworks like Apache Spark minimize disk I/O by keeping data in memory (RDDs), which might bridge the performance gap between PM and VM. Future research should investigate whether memory-resident computing can "mask" the poor disk I/O of virtualized environments.
