DeepEE: Bridging the Gap Between IT Scheduling and Mechanical Cooling via PADQN

DeepEE: Joint Optimization of Job Scheduling and Cooling Control for Data Center Energy Efficiency Using Deep Reinforcement Learning

2019-07-01
Yongyi Ran, Han Hu, Xin Zhou, Yonggang Wen
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces DeepEE, a model-free framework using Deep Reinforcement Learning (DRL) to jointly optimize server job scheduling and cooling unit airflow control in data centers. By employing a novel Parameterized Action Space Deep Q-Network (PADQN), it achieves SOTA energy efficiency, outperforming traditional siloed and joint optimization methods.

TL;DR

DeepEE is a model-free Deep Reinforcement Learning (DRL) framework designed to solve the energy efficiency puzzle in data centers. It tackles the challenge of managing discrete job scheduling alongside continuous cooling airflow control. By introducing the PADQN algorithm and a two-time-scale mechanism, it reduces total energy consumption by up to 15% compared to existing siloed approaches.

Background & Motivation: The Silicon-Mechanical Mismatch

Modern data centers consume massive amounts of electricity, with IT systems (56%) and cooling units (30%) being the primary culprits. Historically, these two are managed in siloes:

  1. IT Scheduling: Focuses on task latency and CPU utilization.
  2. Cooling Control: Focuses on preventing "hot-spots" via conservative over-provisioning.

The problem is that these systems are fundamentally different. IT tasks respond in seconds, while cooling fans and thermal inertia have a "thermal lag" that takes minutes to stabilize. Combining them creates a hybrid action space (Discrete Server ID + Continuous Airflow Rate) that most standard RL algorithms struggle to solve without losing precision through discretization.

Methodology: The PADQN Core

The authors proposed PADQN (Parameterized Action Space Deep Q-Network) to navigate this complexity.

1. Hybrid Action Processing

Instead of simply discretizing the airflow control (which leads to coarse-grained, unstable power spikes), the authors utilized a dual-network approach:

  • Policy Network: Outputs the continuous cooling action (airflow rate).
  • Q-Network: Takes both the state and the policy network's output to calculate energy-cost/reward values for the discrete task scheduling.

2. Coordination via Two-Time-Scale Control

To resolve the temporal mismatch, a time factor () was introduced. While the IT scheduler makes decisions every few seconds, the cooling control only activates at designated intervals (e.g., every 5 minutes). This prevents the cooling system from "chasing" temporary IT spikes, which would lead to mechanical wear and energy oscillation.

DeepEE Overall Architecture

Experimental Results: Real-Trace Validation

The framework was tested using high-fidelity CFD (Computational Fluid Dynamics) 3D models of the National Supercomputing Centre Singapore and real-world workloads.

Performance Gains

  • Energy Efficiency: Saved 15% more energy than siloed cooling-only or IT-only controllers.
  • PUE (Power Usage Effectiveness): Achieved a consistently lower PUE compared to traditional JCO (Joint Control Optimizers).
  • Thermal Stability: Successfully kept outlet temperatures close to the threshold (30°C) without crossing into the "hot-spot" zone, whereas simpler controllers either over-cooled (wasting energy) or allowed dangerous temperature spikes.

Performance Comparison - PUE and Reward

The Importance of t_cool

A critical finding in the ablation study was that setting the cooling interval correctly ( epochs / 5 mins) outperformed both high-frequency control (which causes air fluctuation) and low-frequency control (which fails to react to thermal changes).

Critical Insight & Conclusion

DeepEE proves that model-free DRL can handle the messy, nonlinear thermodynamics of a 3D data center better than simplified queueing or linear models. The real "secret sauce" is the handling of the hybrid action space—allowing the agent to pick a server knowing how the airflow will be adjusted simultaneously.

Limitations: The study primarily targets homogeneous clusters. Extending this to heterogeneous "Cloud" environments with GPUs and diverse cooling (liquid cooling + air cooling) remains a fertile ground for future research.

Takeaway: Effective Green AI isn't just about better code; it's about making sure your software knows the thermal cost of the hardware it runs on.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Parameterized Action Space MDPs (PAMDP) for resource allocation in cyber-physical systems.
  • Which original research established the theoretical framework for Deep Deterministic Policy Gradient (DDPG) and how does PADQN integrate its principles with DQN?
  • Investigate studies applying multi-agent reinforcement learning (MARL) to coordinate power management between IT systems and variable-frequency cooling equipment in edge data centers.
Contents
DeepEE: Bridging the Gap Between IT Scheduling and Mechanical Cooling via PADQN
1. TL;DR
2. Background & Motivation: The Silicon-Mechanical Mismatch
3. Methodology: The PADQN Core
3.1. 1. Hybrid Action Processing
3.2. 2. Coordination via Two-Time-Scale Control
4. Experimental Results: Real-Trace Validation
4.1. Performance Gains
4.2. The Importance of t_cool
5. Critical Insight & Conclusion