Beyond Packet Switching: Optimizing DML Iterations with Optical Circuit Schedulers
KNOWLEDGE‐BASED SYSTEMS
This paper introduces a specialized online scheduling framework for Distributed Machine Learning (DML) in Optical Circuit Switch (OCS) networks. It proposes two novel heuristic algorithms, Heaviest-Load-First (HLF) for intra-job flow scheduling and Shortest Weighted Remaining Time First (SWRTF) for inter-job scheduling, achieving SOTA performance in reducing job completion times.
TL;DR
Distributed Machine Learning (DML) is often bottlenecked by communication. While Optical Circuit Switches (OCS) offer high bandwidth and low power, their reconfiguration delay makes traditional scheduling inefficient. This paper introduces a specialized scheduler—HLF (Intra-job) and SWRTF (Inter-job)—that treats DML training as a series of interleaved stages, cutting communication time by up to 64.97% and significantly lowering total job completion time.
Academic Context
In the hierarchy of data center networking, we are moving from static architectures to Demand-Aware Networks. This paper sits at the intersection of Physical Layer reconfiguration and Application Layer awareness. It tackles the NP-hard problem of multi-job scheduling within OCS constraints, moving beyond the "Coflow" abstraction to a "Multi-stage Iterative" model.
The Problem: The "Iteration" Blind Spot
Most existing OCS schedulers (like Sunflow or Solstice) treat traffic as a Coflow—a collection of parallel flows that start and end together. However, DML is an iterative beast:
- Iterative nature: Hundreds of thousands of cycles of computation communication.
- Interleaving: While one worker computes, the network could be serving another worker's communication.
- Port Constraints: In OCS, one port can only talk to one other port at a time. If the scheduler is "blind" to which port is the bottleneck, the whole iteration hangs.
Methodology: HLF & SWRTF
1. Heaviest-Load-First (HLF)
The authors argue that the completion time of an iteration is dictated by the heaviest load port. If you schedule light flows first, the "heavy" port remains busy long after others finish, wasting overall bandwidth.
- Insight: Prioritize the port with the most data to send/receive. This allows smaller flows to "fill in the gaps" without extending the total ICT.
Figure 1: Comparison of traditional packet switches (a) vs OCS-based reconfigurable topology (b).
2. Shortest Weighted Remaining Time First (SWRTF)
For multi-job environments, the challenge is when to switch between Jobs A and B.
- The "Computation Gap": Traditional SJF (Shortest Job First) might keep the OCS circuits locked even when a job is in its computation phase.
- The Fix: SWRTF re-evaluates priorities () every time a job enters its computation stage, releasing circuits to higher-priority jobs.
Experimental Validation
The researchers tested their algorithms against Sunflow (the previous gold standard for OCS coflow scheduling) using VGG-16 models and synthetic DML workloads.
- ICT Reduction: HLF outperformed Sunflow consistently, especially as "Iterationflow Density" (the complexity of the traffic) increased.
- WJCT Performance: SWRTF showed its strength in online scenarios, proving that considering the "Remaining Completion Time" is far superior to simple weight-based or FIFO scheduling.
Figure 2: WJCT Comparison across different ratios of computation vs communication (Beta).
Critical Insight & Future Outlook
The core value of this work is the rejection of the "one-size-fits-all" network abstraction. By admitting that DML is different from standard web traffic or big data shuffles, the authors unlock massive efficiency.
Limitations: The HLF algorithm has a higher time complexity than random schedulers, which might be a hurdle for ultra-fast, nanosecond-scale switching in future hardware.
Takeaway: As AI models grow, the network must become a first-class citizen in the training loop. This paper provides the mathematical and algorithmic blueprint for building OCS-based AI clusters that don't just provide "fat pipes," but "smart pipes."
Summary Table
| Metric | Improvement over SOTA |
|---|---|
| Iteration Comm. Time (ICT) | Up to 64.97% (vs Sunflow) |
| Weighted Job Completion Time | 42.9% - 54.2% (vs SJF/Baraat) |
| Sensitivity to Delay | Highly robust to OCS reconfiguration latency |
