OS-ELM: Real-Time Job Failure Prediction in the Cloud

Predicting of Job Failure in Compute Cloud Based on Online Extreme Learning Machine: A Comparative Study

2017-01-01
Chunhong Liu, Jingjing Han, Yanlei Shang, Chuanchang Liu, Bo Cheng, Junliang Chen
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes an online job failure prediction method for cloud clusters using Online Sequential Extreme Learning Machine (OS-ELM). By processing Google cluster traces, the method achieves 93% prediction accuracy with an ultra-fast model update time of 0.01 seconds.

TL;DR

Cloud data centers waste massive amounts of resources on jobs that eventually fail. This paper introduces a prediction framework based on Online Sequential Extreme Learning Machine (OS-ELM) that can identify potential job failures in real-time. Achieving 93% accuracy with a staggering 0.01s update time, it crushes traditional SVM-based methods in both speed and scalability.

Problem & Motivation: The "Streaming" Reality of Cloud

In large-scale compute clusters like Google's, roughly 40% of jobs terminate abnormally (killed or failed). Predicting these failures early allows for proactive resource reclamation. However, traditional researchers faced a "velocity vs. accuracy" trade-off:

  • Offline models (ELM, SVM) require re-training on the entire history to learn new patterns, making them useless for streaming data.
  • Online SVM (OS-SVM) involves complex quadratic programming that slows down as the number of "support vectors" grows over time.

The authors recognized that cloud logs are a continuous stream. What was needed was a model that could learn incrementally—absorbing new information without forgetting the old or slowing to a crawl.

Methodology: The Power of OS-ELM

The core innovation lies in the use of OS-ELM. Unlike traditional neural networks that use backpropagation (which is iterative and slow), ELM-based methods randomly fix input weights and only analytically solve for output weights.

The OS-ELM Workflow

  1. Initialization: Use a small initial batch of data to establish the base model.
  2. Sequential Learning: As new job data arrives, the model updates its weights using a recursive least-squares approach. It does not reuse previous samples, saving massive amounts of memory and compute.

Architecture of Online Job Failure Prediction

The feature set includes static characteristics like scheduling class, task count, and requested resources (CPU, RAM, Disk), which are extracted before the job begins execution to enable truly "early" prediction.

Experiments & Results: Speed is King

The authors tested their method against Google Cluster traces (29 days of data). The results highlight a massive disparity in performance metrics:

Performance Comparison

MetricOS-ELMOS-SVM
Update Time~0.01 s~33.04 s
Accuracy~93.10%~81.05%
Precision~94.62%~78.15%

OS-ELM is not just slightly faster—it is orders of magnitude more efficient. While OS-SVM struggles with the KKT conditions and increasing complexity, OS-ELM remains stable.

Influences of Hidden Nodes on OS-ELM Figure: The study on Hidden Nodes (HN) shows that OS-ELM reaches peak stability and accuracy with a relatively small number of neurons (50-100), further reducing the computational footprint.

Critical Insight: Why Does It Work?

The effectiveness of OS-ELM in this context stems from its Inductive Bias. Cloud failures often follow patterns related to resource over-commitment or scheduling priority. Because OS-ELM uses a Moore-Penrose generalized inverse to find the minimum norm solution, it achieves excellent generalization even with few samples.

The 0.01s update time is the breakthrough. In a production environment, a delay of 33 seconds (as seen in OS-SVM) might mean the failing job has already consumed significant resources before the system decides to kill it.

Conclusion & Perspective

This paper presents a strong case for "Simple but Fast" algorithms in the era of Big Data. While deep learning dominates the headlines, the OS-ELM's ability to handle streaming data with minimal latency makes it a far more practical choice for system-level middleware in cloud computing.

Future Work: The authors suggest moving toward a distributed online environment, which would be the next logical step to handle the petabyte-scale logs of modern hyperscalers.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Online Sequential Extreme Learning Machine (OS-ELM) for anomaly detection in distributed cloud systems after 2020.
  • Which paper first proposed the OS-ELM algorithm, and what were the primary mathematical improvements it made over the original Extreme Learning Machine (ELM)?
  • Explore research that applies OS-ELM or similar online incremental learning techniques to proactive energy-aware scheduling in green data centers.
Contents
OS-ELM: Real-Time Job Failure Prediction in the Cloud
1. TL;DR
2. Problem & Motivation: The "Streaming" Reality of Cloud
3. Methodology: The Power of OS-ELM
3.1. The OS-ELM Workflow
4. Experiments & Results: Speed is King
4.1. Performance Comparison
5. Critical Insight: Why Does It Work?
6. Conclusion & Perspective