OS-ELM: Real-Time Job Failure Prediction in the Cloud
Predicting of Job Failure in Compute Cloud Based on Online Extreme Learning Machine: A Comparative Study
This paper proposes an online job failure prediction method for cloud clusters using Online Sequential Extreme Learning Machine (OS-ELM). By processing Google cluster traces, the method achieves 93% prediction accuracy with an ultra-fast model update time of 0.01 seconds.
TL;DR
Cloud data centers waste massive amounts of resources on jobs that eventually fail. This paper introduces a prediction framework based on Online Sequential Extreme Learning Machine (OS-ELM) that can identify potential job failures in real-time. Achieving 93% accuracy with a staggering 0.01s update time, it crushes traditional SVM-based methods in both speed and scalability.
Problem & Motivation: The "Streaming" Reality of Cloud
In large-scale compute clusters like Google's, roughly 40% of jobs terminate abnormally (killed or failed). Predicting these failures early allows for proactive resource reclamation. However, traditional researchers faced a "velocity vs. accuracy" trade-off:
- Offline models (ELM, SVM) require re-training on the entire history to learn new patterns, making them useless for streaming data.
- Online SVM (OS-SVM) involves complex quadratic programming that slows down as the number of "support vectors" grows over time.
The authors recognized that cloud logs are a continuous stream. What was needed was a model that could learn incrementally—absorbing new information without forgetting the old or slowing to a crawl.
Methodology: The Power of OS-ELM
The core innovation lies in the use of OS-ELM. Unlike traditional neural networks that use backpropagation (which is iterative and slow), ELM-based methods randomly fix input weights and only analytically solve for output weights.
The OS-ELM Workflow
- Initialization: Use a small initial batch of data to establish the base model.
- Sequential Learning: As new job data arrives, the model updates its weights using a recursive least-squares approach. It does not reuse previous samples, saving massive amounts of memory and compute.

The feature set includes static characteristics like scheduling class, task count, and requested resources (CPU, RAM, Disk), which are extracted before the job begins execution to enable truly "early" prediction.
Experiments & Results: Speed is King
The authors tested their method against Google Cluster traces (29 days of data). The results highlight a massive disparity in performance metrics:
Performance Comparison
| Metric | OS-ELM | OS-SVM |
|---|---|---|
| Update Time | ~0.01 s | ~33.04 s |
| Accuracy | ~93.10% | ~81.05% |
| Precision | ~94.62% | ~78.15% |
OS-ELM is not just slightly faster—it is orders of magnitude more efficient. While OS-SVM struggles with the KKT conditions and increasing complexity, OS-ELM remains stable.
Figure: The study on Hidden Nodes (HN) shows that OS-ELM reaches peak stability and accuracy with a relatively small number of neurons (50-100), further reducing the computational footprint.
Critical Insight: Why Does It Work?
The effectiveness of OS-ELM in this context stems from its Inductive Bias. Cloud failures often follow patterns related to resource over-commitment or scheduling priority. Because OS-ELM uses a Moore-Penrose generalized inverse to find the minimum norm solution, it achieves excellent generalization even with few samples.
The 0.01s update time is the breakthrough. In a production environment, a delay of 33 seconds (as seen in OS-SVM) might mean the failing job has already consumed significant resources before the system decides to kill it.
Conclusion & Perspective
This paper presents a strong case for "Simple but Fast" algorithms in the era of Big Data. While deep learning dominates the headlines, the OS-ELM's ability to handle streaming data with minimal latency makes it a far more practical choice for system-level middleware in cloud computing.
Future Work: The authors suggest moving toward a distributed online environment, which would be the next logical step to handle the petabyte-scale logs of modern hyperscalers.
