Predicting High-Performance Computing Job Performance: A Machine Learning Approach

Machine Learning Based Performance Analysis and Prediction of Jobs on a HPC Cluster

2019-12-01
Zhengxiong Hou, Shuxin Zhao, Chao Yin, Yunlan Wang, Jianhua Gu, Xingshe Zhou
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a machine learning-based framework for predicting the performance (CPU time) of parallel jobs on an HPC cluster. By leveraging historical job logs from Northwestern Polytechnical University, the authors compare multiple models—Multivariate Linear Regression, Polynomial Regression, and Neural Networks—achieving a prediction accuracy of over 83% for specific user applications.

TL;DR

Researchers at Northwestern Polytechnical University have developed a framework to predict the CPU time of HPC jobs by mining years of historical logs. By training specialized models (Linear, Polynomial, and Neural Networks) per user and application, they successfully predicted the performance of new scientific jobs (like VASP) with over 83% accuracy, even when critical memory features were initially unknown.

Background & Positioning

In the world of High-Performance Computing (HPC), job scheduling is the "brain" that determines efficiency. Most schedulers rely on user-estimated walltimes, which are often exaggerated to prevent jobs from being killed. This work moves away from user intuition and toward Data-Driven Performance Modeling, situating itself as a practical application of regression and neural networks to real-world infrastructure management.

The Problem: The "Informational Gap"

Existing methods fail primarily because:

  1. Data Complexity: The relationship between requested cores, memory, and actual CPU time is rarely linear.
  2. Missing Features: Some of the most predictive features, such as UsedMemory, are only known after a job completes, making them useless for pre-execution prediction unless they can be estimated.

The authors' insight was to stop looking for a universal model and instead focus on user-specific behaviors and application-specific "seeds" (found in input files) to predict these missing features.

Methodology: Specialized Model Architecture

The workflow involves rigorous data cleaning of Torque-based logs, followed by the deployment of four distinct model types:

  1. Multivariate Linear Regression: Establishing a baseline by assuming linear weights for features like StartTime and CoreNumber.
  2. Multivariate Polynomial Regression: Capturing non-linearities using Taylor expansion-style terms (up to degree 4).
  3. Linear Neural Networks: A single-layer approach for fast iteration.
  4. BP Neural Networks (Back-Propagation): A three-layer (input, hidden, output) architecture designed to find complex, hidden correlations.

Bridging the Feature Gap

To solve the "missing feature" problem for new VASP jobs, the authors extracted the Mesh Matrix from VASP input files. They discovered that the computational volume () calculated from these matrices significantly correlates with memory usage.

Need to replace with Figure: Model Architecture or Input Table Table 1: The 10 key features extracted from job logs used for training.

Experiments & Core Insights

The researchers found an "Overfitting Threshold" in polynomial models. While increasing the degree reduced training error, it often caused spikes in testing error (e.g., degree 4 for user xuepy led to an MSE of 39.979 compared to 1.576 at degree 2).

Performance Comparison

The results confirm the necessity of "Model Selection per User":

  • User USPEX: BP Neural Network was the clear winner (MSE 3.638).
  • User xuepy: Polynomial Regression (Degree 2) was most effective (MSE 1.576).

Experimental Results Comparison Figure 1: Distribution of measured (red) vs. predicted (blue) values for specific users.

Critical Analysis & Conclusion

Takeaway

The study proves that even with "messy" real-world logs, ML can achieve high accuracy if we leverage application-specific domain knowledge (like parsing VASP input files). This shifts HPC management from reactive to proactive.

Limitations

  • Feature Sensitivity: The dependency on specific input file formats (like VASP's OUTCAR) makes the method difficult to generalize to "black-box" proprietary software.
  • Static Nature: The models are trained on historical data and may need frequent retraining as cluster hardware or compiler versions change.

Future Work

The authors aim to extend this to GPU-based jobs, where the performance bottlenecks (PCIe bandwidth, VRAM) are significantly different from CPU-bound tasks. This research paves the way for "Intelligent Schedulers" that can automatically adjust job priority based on predicted resource footprints.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning and Recurrent Neural Networks (RNNs) to predict HPC job execution times based on temporal sequence patterns in job logs.
  • Which paper first proposed the use of the "Parallel Workloads Archive" for standardizing HPC job metadata, and how does contemporary research handle the data quality issues identified there?
  • Explore research that applies runtime prediction models to the dynamic scheduling of GPU-accelerated jobs in heterogeneous HPC environments.
Contents
Predicting High-Performance Computing Job Performance: A Machine Learning Approach
1. TL;DR
2. Background & Positioning
3. The Problem: The "Informational Gap"
4. Methodology: Specialized Model Architecture
4.1. Bridging the Feature Gap
5. Experiments & Core Insights
5.1. Performance Comparison
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Work