Automated Complaint Analysis: Navigating Economic Activities with Machine Learning
Automatic Identification of Economic Activities in Complaints
2019-01-01
Summary
Problem
Method
Results
Takeaways
Abstract
This paper presents an automated system for classifying citizen complaints into economic activity categories for the Portuguese Economic and Food Safety Authority (ASAE). By framing the problem as a multi-class text categorization task and utilizing a Linear SVM, the authors achieve a SOTA accuracy of 0.8164 for top-1 and 0.9474 for top-3 predictions.
## TL;DR
The Portuguese Economic and Food Safety Authority (ASAE) receives over 20,000 complaints a year. Managing this manually is a resource drain. This paper introduces an NLP-driven system that automatically categorizes these complaints into economic sectors (e.g., Restoration, Retail, Industry) using a **Linear SVM**, achieving **81.6% accuracy** and a staggering **94.7% accuracy** when providing a top-3 recommendation list.
## The Administrative Bottleneck: Why This is Hard
Governments are digitizing, but "digital" usually just means free-form text fields in contact forms. For ASAE, this results in:
- **High Volume**: 20k+ complaints annually.
- **Noisy Data**: Complaints include HTML tags, meta-data, mixed languages (Portuguese/English), or references to previous attachments.
- **Class Imbalance**: Over 40% of complaints fall into "Restoration and Beverages," while others have almost zero data.
- **Resource Constraints**: Portuguese is a "low-resource" language in the NLP world compared to English, making standard tools less reliable.
## Methodology: The Search for the Best Pipeline
The authors didn't just pick a model; they conducted a comprehensive "battle royale" of NLP libraries and algorithms.
### 1. Preprocessing Excellence
They compared **NLTK**, **spaCy**, and **StanfordNLP**. While NLTK showed slight numerical leads, the authors opted for **StanfordNLP** due to its superior Part-of-Speech (PoS) tagging and punctuation identification, which proved critical for cleaning official documents.
### 2. The Model Architecture
The system uses a classical but robust pipeline:
- **Feature Engineering**: TF-IDF (Term Frequency-Inverse Document Frequency) to weigh the importance of words.
- **Decision Engine**: Evaluation of 8+ classifiers including Random Forests, Naive Bayes, and K-Nearest Neighbors.

*Table: Cross-comparison of NLP libraries and ML classifiers. Linear SVM consistently dominates.*
## Key Insights: When "Less" is "More"
The paper provides several counter-intuitive findings that challenge common ML wisdom:
1. **LDA/PCA Failed**: Attempting to reduce dimensionality via Latent Dirichlet Allocation (LDA) or Principal Component Analysis (PCA) significantly *decreased* accuracy. The raw sparsity of TF-IDF was more informative.
2. **Balancing Hurdles**: Neither Random Over-Sampling (ROS) nor Random Under-Sampling (RUS) improved the SVM. The model naturally handled the distribution better without synthetic intervention.
3. **The Power of Rankings**: In an administrative context, a model doesn't need to be 100% right on the first try. By offering a "Top-3" list, the system becomes a high-reliability assistant for human operators.

*Table: The jump from Acc@1 (80.7%) to Acc@3 (94.7%) demonstrates the model's utility as a decision-support tool.*
## Error Analysis: The Human Element
The authors performed a manual deep-dive into the failures. They found that the model often gets confused by:
- **Semantic Overlap**: "Industry" vs. "Retail" often share the same vocabulary in a complaint about a faulty product.
- **Empty Content**: Some "complaints" are just citizens asking for status updates on previous files, containing no sector-specific keywords.
- **Human Error**: Interestingly, the ML model identified complaints where the *human* operators had originally assigned the wrong category.
## Conclusion & Future Outlook
The study concludes that for Portuguese administrative text, **Linear SVM with TF-IDF** is the champion. It’s fast, explainable, and highly accurate.
**What's next?** The authors plan to move into the era of **Deep Learning** and **Word Embeddings** (like BERT or FastText). However, this paper serves as a vital benchmark: if a sophisticated Neural Network can't beat a 94.7% Top-3 accuracy, the classical approach remains the optimal choice for production-grade government tools.
### Takeaway for Practitioners
If you are dealing with imbalanced, noisy, domain-specific text in a non-English language, don't rush to use a massive Transformer. A well-tuned SVM with a solid NLP preprocessing pipeline might just be your most cost-effective and accurate solution.
