RFMTL: Bridging the Scale Gap in HPC Performance Prediction
Using Small-Scale History Data to Predict Large-Scale Performance of HPC Application
The paper introduces a novel two-level machine learning framework, RFMTL, designed to predict large-scale HPC application performance using only small-scale historical data. By combining Random Forest interpolation with Multi-task Lasso extrapolation, it achieves state-of-the-art accuracy in cross-scale performance modeling.
TL;DR
Predicting how an application will perform on 512 processors using data from only 16 to 128 processors is a notorious "extrapolation" challenge. This paper presents a Two-Level Model (RFMTL) that uses Random Forest for parameter interpolation and Multi-task Lasso for scalability extrapolation. The result? A significant reduction in prediction error compared to traditional DL and regression methods.
Background: The IID Fallacy in HPC
In the world of Machine Learning, we usually assume our training and testing data come from the same distribution (IID). However, in High-Performance Computing (HPC), we often train on "cheap" small-scale runs and want to predict "expensive" large-scale executions. This is a classic out-of-distribution (OOD) extrapolation problem where the IID hypothesis breaks down, causing standard models like Neural Networks or Random Forests to fail spectacularly.
The Core Insight: Decompression of Complexity
The authors argue that performance prediction involves two distinct types of complexity:
- Parameter Sensitivity: How do input variables (e.g., grid size, particles) affect time?
- Scalability Logic: How does execution time diminish as we add processors ()?
By separating these, we can use a high-capacity black-box model (Random Forest) for the parameters and a structured, theory-driven model (PMNF) for the scaling.
Methodology: The Two-Level Architecture
Level 1: Interpolation (The "What")
The model first uses a Random Forest (RF). Since RF is an ensemble of decision trees, it excels at handling non-numeric features and non-linear relationships. Within the small-scale range, IID holds, so RF provides a "clean" baseline for how parameters interact.
Level 2: Extrapolation (The "How it Scales")
Predicting large requires the Performance Model Normal Form (PMNF): Instead of fitting this for every run individually—which would be sensitive to noise—the authors use Multi-task Lasso with K-means clustering.
- Clustering: Groups similar parameter sets that likely share scaling behaviors.
- Multi-task Lasso: Learns coefficients for all tasks in a cluster simultaneously. This "data augmentation" effect allows the model to ignore random noise in the interpolation predictions and focus on the underlying scaling trend.
Figure 1: The workflow showing the transition from parameter-space interpolation to scaling-space extrapolation.
Experimental Validation
The authors tested their approach on two major benchmarks: MCB (Monte Carlo) and Kripke (Particle Transport).
SOTA Comparison
As shown in the table below, the RFMTL method consistently achieves the lowest MAPE, MAE, and RMSE.
- Random Forest (Baseline): Fails because it cannot predict values outside its training range (it simply predicts the maximum seen value).
- MLP (Neural Network): While it can extrapolate, its "guesses" for large-scale data are often wild and unconstrained by the physics of parallel algorithms.
Table 1: Comparison of prediction errors across different ML methods.
The Power of Multi-tasking
One of the most compelling findings is the comparison between Single-Task (ST) and Multi-Task (MT) extrapolation. Single-task models are highly susceptible to "jitter" in the training data. Multi-task learning acts as a regularizer, ensuring the scalability curve follows a logical trajectory.
Figure 2: Visualizing how Multi-task learning (MT) stays closer to the ground truth compared to noisy Single-task (ST) predictions.
Critical Insight & Conclusion
The significance of this work lies in its hybrid philosophy. Purely analytical models require too much expert knowledge, while pure ML models lack physical "common sense" regarding scalability. RFMTL strikes a balance: it uses ML to handle the "messy" application parameters and a constrained linear framework to handle the "structured" laws of parallel scaling.
Future Outlook: While K-means is a solid start for clustering, future work could involve Graph Neural Networks (GNNs) to better capture the relationship between parameter tasks, potentially pushing the extrapolation boundary even further.
Takeaway: If you are building performance models for systems where data is expensive to collect at scale, stop using vanilla Regression or RF. Move to a two-level approach that respects the laws of scalability.
