DeepEE: Bridging the Gap Between IT Scheduling and Mechanical Cooling via PADQN
DeepEE: Joint Optimization of Job Scheduling and Cooling Control for Data Center Energy Efficiency Using Deep Reinforcement Learning
This paper introduces DeepEE, a model-free framework using Deep Reinforcement Learning (DRL) to jointly optimize server job scheduling and cooling unit airflow control in data centers. By employing a novel Parameterized Action Space Deep Q-Network (PADQN), it achieves SOTA energy efficiency, outperforming traditional siloed and joint optimization methods.
TL;DR
DeepEE is a model-free Deep Reinforcement Learning (DRL) framework designed to solve the energy efficiency puzzle in data centers. It tackles the challenge of managing discrete job scheduling alongside continuous cooling airflow control. By introducing the PADQN algorithm and a two-time-scale mechanism, it reduces total energy consumption by up to 15% compared to existing siloed approaches.
Background & Motivation: The Silicon-Mechanical Mismatch
Modern data centers consume massive amounts of electricity, with IT systems (56%) and cooling units (30%) being the primary culprits. Historically, these two are managed in siloes:
- IT Scheduling: Focuses on task latency and CPU utilization.
- Cooling Control: Focuses on preventing "hot-spots" via conservative over-provisioning.
The problem is that these systems are fundamentally different. IT tasks respond in seconds, while cooling fans and thermal inertia have a "thermal lag" that takes minutes to stabilize. Combining them creates a hybrid action space (Discrete Server ID + Continuous Airflow Rate) that most standard RL algorithms struggle to solve without losing precision through discretization.
Methodology: The PADQN Core
The authors proposed PADQN (Parameterized Action Space Deep Q-Network) to navigate this complexity.
1. Hybrid Action Processing
Instead of simply discretizing the airflow control (which leads to coarse-grained, unstable power spikes), the authors utilized a dual-network approach:
- Policy Network: Outputs the continuous cooling action (airflow rate).
- Q-Network: Takes both the state and the policy network's output to calculate energy-cost/reward values for the discrete task scheduling.
2. Coordination via Two-Time-Scale Control
To resolve the temporal mismatch, a time factor () was introduced. While the IT scheduler makes decisions every few seconds, the cooling control only activates at designated intervals (e.g., every 5 minutes). This prevents the cooling system from "chasing" temporary IT spikes, which would lead to mechanical wear and energy oscillation.

Experimental Results: Real-Trace Validation
The framework was tested using high-fidelity CFD (Computational Fluid Dynamics) 3D models of the National Supercomputing Centre Singapore and real-world workloads.
Performance Gains
- Energy Efficiency: Saved 15% more energy than siloed cooling-only or IT-only controllers.
- PUE (Power Usage Effectiveness): Achieved a consistently lower PUE compared to traditional JCO (Joint Control Optimizers).
- Thermal Stability: Successfully kept outlet temperatures close to the threshold (30°C) without crossing into the "hot-spot" zone, whereas simpler controllers either over-cooled (wasting energy) or allowed dangerous temperature spikes.

The Importance of t_cool
A critical finding in the ablation study was that setting the cooling interval correctly ( epochs / 5 mins) outperformed both high-frequency control (which causes air fluctuation) and low-frequency control (which fails to react to thermal changes).
Critical Insight & Conclusion
DeepEE proves that model-free DRL can handle the messy, nonlinear thermodynamics of a 3D data center better than simplified queueing or linear models. The real "secret sauce" is the handling of the hybrid action space—allowing the agent to pick a server knowing how the airflow will be adjusted simultaneously.
Limitations: The study primarily targets homogeneous clusters. Extending this to heterogeneous "Cloud" environments with GPUs and diverse cooling (liquid cooling + air cooling) remains a fertile ground for future research.
Takeaway: Effective Green AI isn't just about better code; it's about making sure your software knows the thermal cost of the hardware it runs on.
