Adaptive Resource Provisioning: Leveraging Job History for Optimized Hybrid Cloud Computing
An Adaptive Resource Provisioning Method Using Job History Learning Technique in Hybrid Infrastructure
This paper introduces an adaptive resource provisioning model for hybrid infrastructures (Cluster and Cloud) using a Multi-Layer Perceptron (MLP) with error back-propagation to learn from job history. The method achieves 77.8% prediction accuracy for optimal VM selection and utilizes horizontal/vertical scaling to satisfy Service Level Agreements (SLA).
TL;DR
Scientists often struggle to provision the "right" amount of resources for large-scale experiments in hybrid environments. This paper presents a machine-learning-driven model that analyzes past job history to predict optimal Virtual Machine (VM) configurations. By combining an MLP-based predictor with dynamic horizontal and vertical scaling, the system ensures that scientific applications meet strict cost and deadline constraints (SLA) even during system failures.
Background: The Hybrid Infrastructure Challenge
In the world of High-Throughput Computing (HTC) and Many-Task Computing (MTC), researchers utilize hybrid infrastructures that bridge local clusters (like Sun Grid Engine) and private/public clouds (like OpenStack). However, the "heterogeneity" of these resources makes manual provisioning a nightmare. Users typically choose between two extremes:
- Cost-Minimum (CM): Scalable but often too slow, violating deadlines.
- Performance-Maximum (PM): Fast but prohibitively expensive.
The authors argue that the missing link is Job History. By treating past execution patterns as a training set, we can move from "guessing" to "predicting" the ideal resource footprint.
Methodology: MLP Learning and Adaptive Scaling
The core of the proposed solution is a two-step process: Predictive Provisioning and Adaptive Management.
1. The Learning Model
The model uses a Multi-Layer Perceptron (MLP) with 11 input nodes and 17 hidden layers. It considers parameters such as application type (e.g., CPU utilization), system status, VM specifications, and historical costs.

2. Service Architecture
The architecture is divided into three layers, with the Middleware Layer (HTCaaS) acting as the broker between the Service Layer and the underlying physical/virtual hardware.

3. Dual-Mode Scaling
Unlike many systems that only add more VMs (Horizontal Scaling), this method also supports Vertical Scaling (adjusting CPU/Memory of existing VMs). When the monitoring algorithm detects a "Deadline Violation" or "System Failure," it triggers a scaling decision to re-balance the workload.
Experimental Proof: Autodock and PYTHIA
The researchers tested the model using two CPU-intensive scientific applications: Autodock (molecular docking) and PYTHIA (physics simulations).
- Accuracy: The MLP model achieved a prediction accuracy of 77.78%, providing a solid foundation for initial resource allocation.
- SLA Satisfaction: For Autodock, the proposed Learning-Based (LB) method successfully balanced resources (using 15 t2.medium instances) to satisfy both the $4.40 budget and the 12,000s deadline. In contrast, the Cost-Minimum approach failed the deadline, and the Performance-Maximum approach was unnecessarily expensive.

Resilience via Simulation
In scenarios involving system failures, the "Non-Adaptive" (NAP) methods inevitably missed deadlines because they couldn't react to lost nodes. The authors' Adaptive Provisioning (AP) used horizontal/vertical scaling to compensate, proving that elasticity is the key to scientific reliability.

Critical Insight & Conclusion
The true value of this work lies in its holistic view of resources. By treating clusters and clouds as a single pool and applying MLP to historical data, it removes the "trial-and-error" aspect of scientific computing.
Limitations: The current model focuses heavily on CPU-intensive tasks. Future iterations will need to address Data-Intensive tasks where network I/O and storage latency might skew the MLP's predictions.
Takeaway: For DevOps engineers and Researchers, this paper highlights that your "Job History" is an untapped goldmine for cost-optimization.
