Beyond Surveys: Leveraging Machine Learning to Decipher Patient Satisfaction
Application of data mining techniques to determine patient satisfaction
This paper presents a machine learning-based methodology to determine patient satisfaction by analyzing a multi-centric dataset encompassing patient perceptions, nurse work environments, and organizational hospital attributes. By employing Naïve Bayes and AdaBoost classifiers, the authors achieved a 87% classification accuracy in predicting satisfaction levels, outperforming traditional statistical-only approaches.
TL;DR
In the healthcare sector, patient satisfaction—the ultimate metric of service quality—has long been measured via static surveys and basic correlations. This paper by Galatas et al. (PETRA '13) shifts the paradigm by treating the hospital as a service provider and applying Data Mining to predict satisfaction with 87% accuracy. By merging patient feedback, nursing environment data, and hospital metrics, the study identifies the "hidden drivers" of healthcare quality.
Problem & Motivation: The Static Survey Trap
While patient satisfaction is a critical indicator for hospital accreditation and market performance, it remains notoriously difficult to pin down. Traditional statistical analysis (like linear regression) often overlooks the complex interactions between high-dimensional data points, such as nurse burnout vs. technological infrastructure.
The authors argue that existing research is often siloed—focusing only on clinical outcomes—while ignoring the "Customer Service" aspect of healthcare. Their goal was to prove that machine learning could not only predict satisfaction but also rank the organizational factors that influence it most.
Methodology: The Data Mining Pipeline
The authors employed a dataset consisting of:
- Patient Perceptions: Professionalism, environment, and discharge processes.
- Nurse Perceptions: Working environment and perceived quality of care.
- Organizational Attributes: Hospital size, type (University vs. General), and infrastructure.
The Stack
- Feature Selection: Using a Greedy Forward Selection wrapper to filter out noise and retain features with high information gain.
- Classifiers:
- Naïve Bayes: Chosen for its efficiency with feature outcome frequency pairs.
- AdaBoost: Used to combine weak learners into a strong predictive model.
- J48 Decision Trees: Utilized for its interpretability (though slightly lower in accuracy).
Note: The study utilized a multi-centric data approach to ensure the model captures diversity across different hospital types.
Experiments and Key Findings
The experimental results were striking, with Naïve Bayes and AdaBoost reaching 87% and 86.96% accuracy respectively.
The Determinants of Satisfaction
The feature selection process revealed a "Hierarchy of Satisfaction":
- Top Predictors: Quality of nursing care, technological infrastructure, and hospital specialization.
- The Surprise: Nurse burnout and perceived job dissatisfaction among staff did not correlate significantly with patient satisfaction. The authors attribute this to the high level of professionalism in nursing staff who "provide a high level of care regardless of circumstances."

Critical Analysis & Conclusion
Takeaway
The paper successfully validates that Data Mining is a viable and more powerful alternative to traditional statistics in the healthcare administrative sector. By treating patients as "customers" in a service model, hospitals can prioritize technical and organizational investments (like infrastructure) that have the highest impact on perceived care.
Limitations & Future Work
The study’s primary limitation lies in its data format—relying on manually processed surveys rather than real-time electronic health records (EHR). Furthermore, as an early work from 2013, it lacks the application of modern Deep Learning techniques which could potentially handle the "missing information" problem more effectively.
As healthcare moves toward a more digitalized future, models like this will be essential for real-time quality management and hospital resource allocation.
