Interpretable Intelligence: Enhancing Cloud Workload Prediction via Construction Functions
8027_A feature generation framework for Google trace analysis.
This paper introduces a function-based feature generation framework for cloud workload analysis, focusing on Google cluster trace logs. By utilizing a set of "construction functions" (such as volatility, level number, and fairness index), the method achieves a state-of-the-art MSE improvement of approximately 8.7% in workload prediction compared to traditional feature sets.
TL;DR
In the realm of cloud computing, predicting workload (CPU/Memory usage) is critical for cost-efficiency. This paper moves away from opaque "black-box" features, proposing a Function-based Feature Generation Framework. By applying mathematical construction functions like Volatility and Fairness to Google Cluster Trace data, the authors improved prediction accuracy (MSE) by 8.7% while maintaining full technical interpretability.
Problem & Motivation: The "Black-Box" Struggle
Modern cloud data centers (like those at Google) handle millions of tasks. Efficient scheduling requires accurate workload prediction. However, researchers face a dilemma:
- Raw Features: Simple averages or max values are too "thin" and miss temporal dynamics.
- Deep Learning Features: While powerful, these "hidden layer" representations are uninterpretable, leaving system admins in the dark about why a prediction was made.
The authors' insight is profound yet simple: Bridge the gap using "Construction Functions" that mirror human physical intuition about machine behavior—how much a signal "vibrates" (Volatility) or how "even" the load is (Fairness).
Methodology: The Core Framework
The framework operates by taking raw trace logs and passing them through three distinct "Construction Function" layers:
1. Simple Construction Functions
These handle basic statistics but with a focus on windowed observations, such as cpu_mean_usage and task_length.
2. Physical Construction Functions
This is the heart of the paper. Instead of just looking at the value, they look at the state:
- Volatility Index: Measures the degree of fluctuation. High volatility often precedes system instability.
- Fairness Index: Derived from networking principles, it evaluates if resource usage is balanced or spike-heavy within a window.
- Level Number: Counts distinct resource usage states, identifying the complexity of the workload.
3. Integrated Construction Functions
These combine multiple metrics (e.g., cpu_memory_rate) to capture the correlation between different resource dimensions.
Figure 1: The proposed Function-based Feature Generation Framework.
Experiments & Results: Proving the Value
The authors validated their framework on the Google Cluster Trace (v2011), a gold standard in the industry. They used the Bayesian Algorithm as the predictor to ensure that the performance gains came from the features themselves rather than an overly complex model.
Key Performance Gains
Compared to the baseline feature set established in prior SOTA works (e.g., [8] in the paper):
- MSE Reduction: Decreased from 3.282 to 2.999.
- Relative Improvement: ~8.7% across the entire trace log.
Robustness across Scales
The study found that as the Observation Window increases, the Physical Construction Functions (like Level Number and Volatility) significantly outperform traditional mean/max features. This suggests that the "intrinsic character" of a workload is better captured over longer temporal spans.
Figure 2: Performance comparison showing lower MSE for the proposed feature set.
Critical Analysis & Conclusion
Takeaway
The biggest contribution of this paper isn't just the 8.7% boost; it's the validation of interpretability. By using functions like "Fairness" and "Volatility," the model provides features that tell a story: "We expect a spike because the fairness index just dropped, indicating a resource-heavy task onset."
Limitations
- Function Selection: Currently, functions are predefined. A more advanced approach might involve using Genetic Programming to discover these construction functions automatically.
- Dynamic Load: The paper focuses on trace logs; real-time application in a live production environment with shifting distributions remains a future challenge.
Future Work
The authors aim to combine this framework with Feature Learning (Deep Learning) while maintaining the interpretability constraints, potentially creating a "Best of Both Worlds" scenario for cloud orchestration.
Academic Reference: The study was supported by the National Natural Science Foundation of China (No. 61572231) and related provincial innovation projects.
