ANN Performance in the Cloud: Quantifying the Memory vs. Disk Trade-off

An Evaluation of Neural Networks Performance for Job Scheduling in a Public Cloud Environment

2018-06-18
Klodiana Goga, Fatos Xhafa, Olivier Terzo
Summary
Problem
Method
Results
Takeaways
Abstract

This paper evaluates the performance of a Multilayer Perceptron Classifier (MLPC) implemented via Apache Spark's MLlib for job scheduling tasks in a public cloud environment (AWS). It specifically quantifies the performance gap between in-memory (MEMORY_ONLY) and disk-based (DISK_ONLY) data persistence when processing large-scale Google Cluster Datasets.

TL;DR

Artificial Neural Networks (ANNs) are increasingly hungry for data, but what happens when that data no longer fits in your cloud instance's RAM? This paper provides an empirical look at the performance of Multilayer Perceptron Classifiers (MLPC) using Apache Spark on AWS. By testing against the Google Cluster Dataset, the authors reveal that while disk storage enables big data processing, it incurs a roughly 20-25% time penalty compared to in-memory processing, turning a 38-hour job into a 48-hour endurance test.

Context & Motivation: The Memory Wall

In the landscape of Big Data analytics, ANNs have moved from niche academic tools to the backbone of predictive modelling in Cloud and IoT systems. However, a significant "Memory Wall" exists. Tools like WEKA often require the entire dataset to be loaded into main memory, causing crashes or extreme slowdowns when datasets reach the gigabyte or terabyte scale.

The authors tackle a fundamental infrastructure question: How much performance are we actually sacrificing when we move from "In-Memory" to "Distributed Disk" storage in a public cloud?

Methodology: Spark MLlib on AWS EMR

The research utilized the Apache Spark MLlib, leveraging its Resilient Distributed Dataset (RDD) abstraction. RDDs are critical because they allow developers to choose "Storage Levels."

The Setup:

  • Platform: Amazon Web Services (AWS) EMR.
  • Compute: 2x c4.8xlarge nodes (72 vCores total, 120 GB RAM).
  • Algorithm: MLPC with 2 hidden layers (70 and 50 neurons).
  • Storage Tiers: Compared MEMORY_ONLY (deserialized in JVM) vs. DISK_ONLY.

System Architecture Figure 1: The AWS EMR Master-Worker architecture used for the distributed training simulations.

Deep Dive into Results

The study used the Google Cluster Trace 2011, a gold-standard dataset for job scheduling tasks. The researchers focused on two types of data: Task Usage and Task Events.

1. The Cost of Disk I/O

The results were consistent across all tests: reading from disk is a heavy tax. For the largest Task Usage dataset (3.5 GB):

  • Memory Only: 138,359 seconds (~38.4 hours)
  • Disk Only: 172,904 seconds (~48.0 hours)
  • Delta: ~9.6 hours of additional compute time.

2. Scalability Trends

As shown in the charts below, the execution time scales with data size, but the "Persistence Gap" remains a significant constant.

Performance Comparison - Task Usage Figure 2: Execution time comparison for Task Usage dataset. Note the consistent gap between memory (blue/shorter) and disk (orange/longer).

Similarly, for the Task Events dataset, even at smaller scales (0.5 GB), the DISK_ONLY persistence required significantly more time, illustrating that the overhead is not merely about capacity but the throughput limits of distributed file systems (HDFS/EMRFS).

Performance Comparison - Task Events Figure 3: Performance results for Task Events data, reinforcing the impact of persistence types.

Critical Insight: Why Does This Matter?

The "why" behind the performance hit is the iterative nature of ANN training. In Spark, MLPC uses optimization routines like L-BFGS. These algorithms require multiple passes over the data. If the data is kept in memory, subsequent passes are near-instant. If the data is on disk, every single iteration triggers a massive I/O cycle across the network and local EBS volumes.

Conclusion & Future Outlook

This paper serves as a pragmatic guide for practitioners managing cloud budgets.

  • Key Takeaway: If your training is taking days, check your RDD persistence. Moving to an instance with higher RAM (to maintain MEMORY_ONLY) may actually be cheaper in the long run than paying for 10+ extra hours of compute on a "cheaper" instance that spills to disk.
  • Limitations: The study focuses on MLPC. The performance dynamics might shift with Convolutional Neural Networks (CNNs) where GPU-to-Memory bandwidth is the primary bottleneck rather than Disk-to-CPU.
  • Future Work: The authors suggest exploring Google File System (GFS) or specialized storage layers to bridge the performance gap between volatile RAM and persistent disk.

Academic References

  • Spark Implementation: Apache Spark MLlib documentation on MLPC.
  • Dataset: Google Cluster Trace 2011 (ClusterData2011_2).
  • Optimization: L-BFGS as the routine for logistic loss optimization in Spark.

Find Similar Papers

Try Our Examples

  • Identify recent comparative studies on the performance of the L-BFGS optimizer versus Stochastic Gradient Descent (SGD) in distributed Spark MLlib environments for large datasets.
  • What are the underlying architectural differences between HDFS and EMRFS that lead to the performance variances mentioned in cloud-based neural network training studies?
  • Explore newer techniques in "Gradient Compression" or "Communication-Efficient Learning" that aim to reduce the I/O and networking overhead highlighted in this paper's findings.
Contents
ANN Performance in the Cloud: Quantifying the Memory vs. Disk Trade-off
1. TL;DR
2. Context & Motivation: The Memory Wall
3. Methodology: Spark MLlib on AWS EMR
3.1. The Setup:
4. Deep Dive into Results
4.1. 1. The Cost of Disk I/O
4.2. 2. Scalability Trends
5. Critical Insight: Why Does This Matter?
6. Conclusion & Future Outlook
6.1. Academic References