ANN Performance in the Cloud: Quantifying the Memory vs. Disk Trade-off
An Evaluation of Neural Networks Performance for Job Scheduling in a Public Cloud Environment
This paper evaluates the performance of a Multilayer Perceptron Classifier (MLPC) implemented via Apache Spark's MLlib for job scheduling tasks in a public cloud environment (AWS). It specifically quantifies the performance gap between in-memory (MEMORY_ONLY) and disk-based (DISK_ONLY) data persistence when processing large-scale Google Cluster Datasets.
TL;DR
Artificial Neural Networks (ANNs) are increasingly hungry for data, but what happens when that data no longer fits in your cloud instance's RAM? This paper provides an empirical look at the performance of Multilayer Perceptron Classifiers (MLPC) using Apache Spark on AWS. By testing against the Google Cluster Dataset, the authors reveal that while disk storage enables big data processing, it incurs a roughly 20-25% time penalty compared to in-memory processing, turning a 38-hour job into a 48-hour endurance test.
Context & Motivation: The Memory Wall
In the landscape of Big Data analytics, ANNs have moved from niche academic tools to the backbone of predictive modelling in Cloud and IoT systems. However, a significant "Memory Wall" exists. Tools like WEKA often require the entire dataset to be loaded into main memory, causing crashes or extreme slowdowns when datasets reach the gigabyte or terabyte scale.
The authors tackle a fundamental infrastructure question: How much performance are we actually sacrificing when we move from "In-Memory" to "Distributed Disk" storage in a public cloud?
Methodology: Spark MLlib on AWS EMR
The research utilized the Apache Spark MLlib, leveraging its Resilient Distributed Dataset (RDD) abstraction. RDDs are critical because they allow developers to choose "Storage Levels."
The Setup:
- Platform: Amazon Web Services (AWS) EMR.
- Compute: 2x
c4.8xlargenodes (72 vCores total, 120 GB RAM). - Algorithm: MLPC with 2 hidden layers (70 and 50 neurons).
- Storage Tiers: Compared
MEMORY_ONLY(deserialized in JVM) vs.DISK_ONLY.
Figure 1: The AWS EMR Master-Worker architecture used for the distributed training simulations.
Deep Dive into Results
The study used the Google Cluster Trace 2011, a gold-standard dataset for job scheduling tasks. The researchers focused on two types of data: Task Usage and Task Events.
1. The Cost of Disk I/O
The results were consistent across all tests: reading from disk is a heavy tax. For the largest Task Usage dataset (3.5 GB):
- Memory Only: 138,359 seconds (~38.4 hours)
- Disk Only: 172,904 seconds (~48.0 hours)
- Delta: ~9.6 hours of additional compute time.
2. Scalability Trends
As shown in the charts below, the execution time scales with data size, but the "Persistence Gap" remains a significant constant.
Figure 2: Execution time comparison for Task Usage dataset. Note the consistent gap between memory (blue/shorter) and disk (orange/longer).
Similarly, for the Task Events dataset, even at smaller scales (0.5 GB), the DISK_ONLY persistence required significantly more time, illustrating that the overhead is not merely about capacity but the throughput limits of distributed file systems (HDFS/EMRFS).
Figure 3: Performance results for Task Events data, reinforcing the impact of persistence types.
Critical Insight: Why Does This Matter?
The "why" behind the performance hit is the iterative nature of ANN training. In Spark, MLPC uses optimization routines like L-BFGS. These algorithms require multiple passes over the data. If the data is kept in memory, subsequent passes are near-instant. If the data is on disk, every single iteration triggers a massive I/O cycle across the network and local EBS volumes.
Conclusion & Future Outlook
This paper serves as a pragmatic guide for practitioners managing cloud budgets.
- Key Takeaway: If your training is taking days, check your RDD persistence. Moving to an instance with higher RAM (to maintain
MEMORY_ONLY) may actually be cheaper in the long run than paying for 10+ extra hours of compute on a "cheaper" instance that spills to disk. - Limitations: The study focuses on MLPC. The performance dynamics might shift with Convolutional Neural Networks (CNNs) where GPU-to-Memory bandwidth is the primary bottleneck rather than Disk-to-CPU.
- Future Work: The authors suggest exploring Google File System (GFS) or specialized storage layers to bridge the performance gap between volatile RAM and persistent disk.
Academic References
- Spark Implementation: Apache Spark MLlib documentation on MLPC.
- Dataset: Google Cluster Trace 2011 (ClusterData2011_2).
- Optimization: L-BFGS as the routine for logistic loss optimization in Spark.
