Heterogeneity-Aware Scheduling: Bridging the HPC Theory-Practice Gap with Machine Learning
A New Approach for Scheduling Job with the Heterogeneity-Aware Resource in HPC Systems
The paper proposes a dynamic job scheduling approach for heterogeneous HPC systems (CPU/Coprocessor) using a combination of simulation and nonlinear regression. By extracting features from workload logs, the authors generate a machine-learning-based scheduling function that outperforms traditional heuristics like FCFS and WFP3 in terms of "bounded slowdown."
TL;DR
Researchers have developed a dynamic scheduling framework for heterogeneous HPC systems that uses simulation-driven machine learning to derive custom priority functions. By incorporating features like Xeon Phi (MIC) card requirements and user-estimated due dates, their approach significantly reduces the Average Bounded Slowdown, outperforming standard heuristics like FCFS and WFP3 across various real-world workload traces.
Background: The Problem of Static Scheduling
In the world of High-Performance Computing (HPC), scheduling is the art of deciding "who goes next." However, most clusters today rely on static policies like First-Come-First-Serve (FCFS) or Shortest Processing Time (SPT). While simple, these methods ignore the nuanced "heterogeneity" of modern workloads—where one job might need 100 CPUs while another requires specialized coprocessors like the Intel Xeon Phi.
The authors argue that the divergence between theoretical scheduling and real-world performance stems from ignoring user behavior and hardware diversity. They propose that instead of guessing a policy, we should "learn" it from the system's own history.
Methodology: From Logs to Learning
The core of this research is a two-phase pipeline designed to extract an optimal scheduling function :
1. The Simulation Stage (Dataset Generation)
Using SimGrid, the authors replayed historical job logs under thousands of different permutations. They calculated the "Bounded Slowdown" for each scenario.
- The Score: Each job was assigned a priority score based on how its position in the queue affected the total system performance.
- The Features: Five variables were tracked: Processing time (), CPU cores (), Submit time (), MIC cards (), and Due date ().
Figure 1: The job submission and scheduling model used in the research.
2. Nonlinear Regression (Function Discovery)
Instead of a "black-box" model, the authors used nonlinear regression to find a mathematical expression that fits the data. They tested combinations of base functions (logarithm, square root, inverse) and operators to find the best-performing equation.
Experiments & Real-World Validation
The authors didn't just stop at simulations; they tested their findings on the SuperNode-XP, a real heterogeneous cluster.
Key Comparison:
The study compared their learned functions (C1, C2, C3) against:
- Standard Heuristics: FCFS, SPT, LPT.
- Smart Ad-hoc Policies: WFP3 and UNICEF.
- SOTA Reference: SC_ref (a function previously optimized for homogeneous systems).
Figure 2: Performance comparison on the ANL Intrepid workload.
Performance Gains:
- On the CTC SP2 log, the learned function C2 reduced the median slowdown to 4.58, while FCFS lingered at a staggering 438.37.
- On physical hardware (SuperNode-XP), C1 emerged as the champion, showing that while optimal functions vary by trace, the ML-based approach consistently finds a winner where standard heuristics fail.
Deep Insight: Why Why Does This Work?
Traditional heuristics are "one-size-fits-all." By contrast, the learned functions capture the Inductive Bias of the specific cluster. For example, some functions might discover that jobs requiring MIC cards should be prioritized during specific submit-time windows to maximize throughput. By treating the scheduler as a dynamic entity that adapts to "Heterogeneity-Aware" resources, the system reduces the idling of specialized hardware.
Critical Analysis & Conclusion
Takeaway
The research successfully demonstrates that the "optimal" scheduler is not a static rule but a hidden function within the system's workload logs. By using nonlinear regression, they provide a transparent, deployable mathematical formula that can be easily integrated into production managers like PBS Pro.
Limitations
- Estimation Accuracy: The model relies on user-provided "walltime" estimates. If users provide wildly inaccurate estimates, the scheduling function's efficiency may degrade.
- Trace Sensitivity: As shown in the results, the "best" function (C1 vs C2) changes depending on the workload, suggesting that models may need periodic "retraining" as user patterns evolve.
Future Work
The authors aim to incorporate more "application-specific" features (like memory bandwidth or network intensity) to further refine the internal logic of the scheduler, potentially moving toward a fully self-optimizing HPC environment.
