Self-Learning IDPS: Solving the Data Scarcity Problem in Healthcare Cybersecurity

A Self-Learning Approach for Detecting Intrusions in Healthcare Systems

2021-06-01
Panagiotis I. Radoglou-Grammatikis, Panagiotis G. Sarigiannidis, Georgios Efstathopoulos, Thomas Lagkas, George F. Fragulis, Antonios Sarigiannidis
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a self-learning Intrusion Detection and Prevention System (IDPS) specifically designed for the Internet of Medical Things (IoMT) and healthcare environments. It utilizes an Active Learning framework to dynamically re-train supervised classifiers (Decision Tree and Random Forest) to detect multi-vector attacks across HTTP and Modbus/TCP protocols.

TL;DR

The rapid growth of the Internet of Medical Things (IoMT) has expanded the attack surface of healthcare organizations. Traditional Intrusion Detection Systems (IDS) struggle because medical data is too sensitive to share, leading to a lack of training datasets. This paper proposes a Self-Learning IDPS that uses Active Learning to adapt to new threats on the fly, achieving over 94% accuracy in detecting both IT (HTTP) and OT (Modbus/TCP) medical protocol attacks.

The Motivation: Why Static Models Fail in Hospitals

Healthcare is currently the most targeted critical infrastructure sector, yet it is also the least prepared. There are two fundamental barriers to effective AI-driven security in this space:

  1. The Privacy Paradox: Machine Learning (ML) requires massive labeled datasets, but hospital network traffic contains sensitive patient data that cannot be publicly disclosed.
  2. Protocol Heterogeneity: Healthcare environments mix standard ICT protocols (like HTTP for Electronic Health Records) with industrial protocols (like Modbus/TCP for medical devices), creating a complex traffic profile that static models cannot easily master.

The researchers recognized that instead of waiting for a "perfect" dataset, the system should learn from the environment it protects.

Methodology: The Active Learning Loop

The core innovation is the shift from Passive Learning (training once on a fixed dataset) to Active Learning. In this framework, the model acts as an "informed seeker" of knowledge.

1. The Architecture

The system consists of three modules:

  • Flow Monitoring: Captures raw traffic via SPAN and converts it into bidirectional flows using CICFlowMeter.
  • Intrusion Detection Engine: Houses the classifiers (Decision Tree for HTTP, Random Forest for Modbus).
  • Notification & Active Learning: Identifies "uncertain" traffic and involves a security expert for labeling.

Overall Architecture

2. The "Uncertainty" Logic

The system doesn't ask for help with every packet. It uses Entropy-based Uncertainty Sampling. If a network flow results in a high entropy value (meaning the model is "confused" about whether it is a DoS, SQL injection, or normal traffic), that flow is prioritized for manual labeling and subsequent re-training.

Experimental Performance

The authors evaluated several algorithms, including Support Vector Machines (SVM), Naive Bayes, and Deep Neural Networks (DNN).

HTTP & Modbus/TCP Results

  • HTTP Protocol: The Decision Tree outperformed others with an accuracy of 96.44% and an F1-score of 0.91.
  • Modbus/TCP Protocol: Random Forest was the winner, achieving 94.45% accuracy and a perfect True Positive Rate (TPR = 1).

The "Self-Learning" Proof

The most compelling evidence is the accuracy progression. As the Active Learner processed more rounds of data, the accuracy didn't just fluctuate—it steadily climbed, proving that the model was successfully "learning" the specific nuances of the local healthcare environment.

Decision Tree Accuracy Increment Fig: Notice the distinct steps in accuracy as the model undergoes re-training phases.

Critical Insights & Takeaways

  • Adaptability is Key: In healthcare, a model trained on Dataset A will likely fail in Hospital B. This paper's approach allows the IDPS to be "calibrated" to a specific hospital's traffic patterns through the active learning loop.
  • Human-in-the-Loop: While the system is highly automated, it retains a "Security Expert" role. This ensures that the labels used for re-training are ground truths, preventing "label noise" from degrading the model over time.
  • Limitations: The reliance on a security expert for labeling could be a bottleneck in high-traffic environments. Future work might explore "Self-Supervised" techniques to further reduce the human workload.

Conclusion

This research provides a robust solution to the "cold start" problem in healthcare cybersecurity. By treating intrusion detection as a dynamic, evolving process rather than a static classification task, the proposed IDPS offers a scalable way to protect the future of IoMT.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply Active Learning or Transfer Learning to improve Intrusion Detection Systems specifically within the Internet of Medical Things (IoMT) domain.
  • Which original study proposed the Entropy-based Uncertainty Sampling strategy in Active Learning, and how does this paper adapt that criteria for real-time network flow analysis?
  • Explore research that integrates Federated Learning with Active Learning to address data privacy and labeling challenges in Critical Infrastructure protection.
Contents
Self-Learning IDPS: Solving the Data Scarcity Problem in Healthcare Cybersecurity
1. TL;DR
2. The Motivation: Why Static Models Fail in Hospitals
3. Methodology: The Active Learning Loop
3.1. 1. The Architecture
3.2. 2. The "Uncertainty" Logic
4. Experimental Performance
4.1. HTTP & Modbus/TCP Results
4.2. The "Self-Learning" Proof
5. Critical Insights & Takeaways
5.1. Conclusion