Predictive Integrity: Leveraging Machine Learning to Measure Corruption Risk in Civil Service
Using Political Party Affiliation Data to Measure Civil Servants' Risk of Corruption
This paper presents a machine learning framework for quantifying corruption risk among civil servants using political party affiliation data. By evaluating Bayesian Networks, SVM, Random Forest, and ANN, the study identifies Random Forest as the optimal model, achieving over 90% precision in identifying corrupt individuals.
Executive Summary
TL;DR: Researchers from the Brazilian Office of the Comptroller General (CGU) have developed a machine learning approach to predict the risk of corruption among civil servants based solely on their political party affiliation. Using a cost-sensitive Random Forest model, they achieved a high precision rate that significantly outperforms traditional expert-designed heuristics, revealing that shorter affiliation durations and specific cancellation motives are strong indicators of corruptibility.
Academic Context: This work transitions corruption detection from subjective, expert-led rule sets to an objective, data-driven classification task, positioning itself as a pioneer in using individual political behavior for risk modeling.
The "Precision" Paradox in Anti-Corruption
In the realm of government oversight, the cost of a False Positive (wrongly accusing an innocent official) is exceptionally high—not just in terms of wasted investigative resources, but also regarding moral and political damage. The Brazilian CGU requires at least 90% precision before launching a formal inquiry.
Existing expert models used by the Department of Research and Strategic Information (DIE) were based on two assumptions:
- Higher number of party affiliations indicates higher risk.
- Certain cancellation motives are more suspicious.
However, these rules were never statistically validated. This paper addresses this gap by asking: Can machine learning provide a more rigorous, high-precision alternative to expert intuition?
Methodology: From Raw Data to Risk Scores
The researchers integrated three primary databases: SIAPE (human resources), TSE (electoral court data), and CEAF (registry of expelled officials).
1. Statistical Foundation
Before modeling, a Chi-Square Hypothesis Test () was conducted. The calculated value () far exceeded the critical value (), statistically proving that political affiliation and corruption are not independent variables for civil servants.
2. Feature Engineering & Discretization
The paper emphasizes the importance of data formatting. They selected three key features:
- Sum of days affiliated: Total time spent in political parties.
- Maximum days in a single party: The longest tenure in one organization.
- Cancellation motive code: Ranging from voluntary choice to judicial cancellation.
To handle continuous data for algorithms like Bayesian Networks, they tested three discretization methods: Multi-interval (MI), Equal-Frequency Binning (EQF), and Proportional k-Interval (PKI).
3. Cost-Sensitive Learning
To meet the 90% precision requirement, all classifiers were wrapped in MetaCost. A cost matrix was applied, assigning a weight of 5.0 to false positives versus 1.0 for false negatives, forcing the models to prioritize accuracy over the "Corrupt" classification.
Table: Comparison shows Random Forest (RF) effectively doubling the recall of Expert models while maintaining identical precision.
Key Insights and Results
The Random Forest (RF) algorithm emerged as the winner. When compared to the Experts' model on a separate test dataset:
- Recall increased by 15% (from 0.17 to 0.32), meaning the automated model catches nearly twice as many corrupt individuals.
- Mean Absolute Error (MAE) dropped by 12%.
- Kappa Statistic (agreement beyond chance) nearly doubled from 0.14 to 0.27.
Challenging Expert Intuition
One of the most striking findings was the rejection of the "Number of Parties" attribute. Experts believed that joining many parties was a red flag. However, feature selection algorithms discarded this attribute as unrepresentative. Instead, the model revealed a non-obvious insight: Shorter total affiliation time is a stronger predictor of corruption.
Physical Intuition: Corrupt individuals often use political ties for temporary gain or to demonstrate "trustworthiness" to specific figures only during their period of illicit activity, rather than maintaining long-term ideological commitment.
Figure 1: Example of a Decision Tree generated within the Random Forest, showcasing the importance of Motive Codes and Affiliation Days.
Critical Analysis & Future Work
While the precision is high, the Recall (0.32-0.36) remains relatively low. This suggests that while the model is very sure when it labels someone "high risk," it currently misses a large portion of corrupt individuals who do not exhibit these specific political patterns.
Limitations:
- The model relies on historical expulsion data, which might not capture the "smartest" corrupt individuals who are never caught.
- The features are limited to affiliation; functional data (salary, position, seniority) could significantly add predictive power.
Conclusion: This research proves that even "thin" data—like political status—contains significant signals for governance. For agencies like the CGU, this provides a statistically sound foundation to prioritize investigations, moving away from subjective "hunches" toward evidence-based risk management.
