Automated Complaint Analysis: Navigating Economic Activities with Machine Learning

Automatic Identification of Economic Activities in Complaints

2019-01-01
Luís Barbosa, João Filgueiras, Gil Rocha, Henrique Lopes Cardoso, Luís Paulo Reis, João Pedro Machado, Ana Cristina Caldeira, Ana Maria Oliveira
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an automated system for classifying citizen complaints into economic activity categories for the Portuguese Economic and Food Safety Authority (ASAE). By framing the problem as a multi-class text categorization task and utilizing a Linear SVM, the authors achieve a SOTA accuracy of 0.8164 for top-1 and 0.9474 for top-3 predictions.

    ## TL;DR
    The Portuguese Economic and Food Safety Authority (ASAE) receives over 20,000 complaints a year. Managing this manually is a resource drain. This paper introduces an NLP-driven system that automatically categorizes these complaints into economic sectors (e.g., Restoration, Retail, Industry) using a **Linear SVM**, achieving **81.6% accuracy** and a staggering **94.7% accuracy** when providing a top-3 recommendation list.

    ## The Administrative Bottleneck: Why This is Hard
    Governments are digitizing, but "digital" usually just means free-form text fields in contact forms. For ASAE, this results in:
    - **High Volume**: 20k+ complaints annually.
    - **Noisy Data**: Complaints include HTML tags, meta-data, mixed languages (Portuguese/English), or references to previous attachments.
    - **Class Imbalance**: Over 40% of complaints fall into "Restoration and Beverages," while others have almost zero data.
    - **Resource Constraints**: Portuguese is a "low-resource" language in the NLP world compared to English, making standard tools less reliable.

    ## Methodology: The Search for the Best Pipeline
    The authors didn't just pick a model; they conducted a comprehensive "battle royale" of NLP libraries and algorithms.

    ### 1. Preprocessing Excellence
    They compared **NLTK**, **spaCy**, and **StanfordNLP**. While NLTK showed slight numerical leads, the authors opted for **StanfordNLP** due to its superior Part-of-Speech (PoS) tagging and punctuation identification, which proved critical for cleaning official documents.

    ### 2. The Model Architecture
    The system uses a classical but robust pipeline:
    - **Feature Engineering**: TF-IDF (Term Frequency-Inverse Document Frequency) to weigh the importance of words.
    - **Decision Engine**: Evaluation of 8+ classifiers including Random Forests, Naive Bayes, and K-Nearest Neighbors.

    ![Model Performance Comparison](https://cdn.atominnolab.com/wisdoc/tables/20260527-48caa062-5bd8-436b-bfa6-c6cba40aeafd/page_004_block_006.png)
    *Table: Cross-comparison of NLP libraries and ML classifiers. Linear SVM consistently dominates.*

    ## Key Insights: When "Less" is "More"
    The paper provides several counter-intuitive findings that challenge common ML wisdom:
    1.  **LDA/PCA Failed**: Attempting to reduce dimensionality via Latent Dirichlet Allocation (LDA) or Principal Component Analysis (PCA) significantly *decreased* accuracy. The raw sparsity of TF-IDF was more informative.
    2.  **Balancing Hurdles**: Neither Random Over-Sampling (ROS) nor Random Under-Sampling (RUS) improved the SVM. The model naturally handled the distribution better without synthetic intervention.
    3.  **The Power of Rankings**: In an administrative context, a model doesn't need to be 100% right on the first try. By offering a "Top-3" list, the system becomes a high-reliability assistant for human operators.

    ![Top-K Accuracy Result](https://cdn.atominnolab.com/wisdoc/tables/20260527-48caa062-5bd8-436b-bfa6-c6cba40aeafd/page_006_block_003.png)
    *Table: The jump from Acc@1 (80.7%) to Acc@3 (94.7%) demonstrates the model's utility as a decision-support tool.*

    ## Error Analysis: The Human Element
    The authors performed a manual deep-dive into the failures. They found that the model often gets confused by:
    - **Semantic Overlap**: "Industry" vs. "Retail" often share the same vocabulary in a complaint about a faulty product.
    - **Empty Content**: Some "complaints" are just citizens asking for status updates on previous files, containing no sector-specific keywords.
    - **Human Error**: Interestingly, the ML model identified complaints where the *human* operators had originally assigned the wrong category.

    ## Conclusion & Future Outlook
    The study concludes that for Portuguese administrative text, **Linear SVM with TF-IDF** is the champion. It’s fast, explainable, and highly accurate. 

    **What's next?** The authors plan to move into the era of **Deep Learning** and **Word Embeddings** (like BERT or FastText). However, this paper serves as a vital benchmark: if a sophisticated Neural Network can't beat a 94.7% Top-3 accuracy, the classical approach remains the optimal choice for production-grade government tools.

    ### Takeaway for Practitioners
    If you are dealing with imbalanced, noisy, domain-specific text in a non-English language, don't rush to use a massive Transformer. A well-tuned SVM with a solid NLP preprocessing pipeline might just be your most cost-effective and accurate solution.

Find Similar Papers

Try Our Examples

  • Search for recent papers using Transformer-based models like BERTimbau for Portuguese text classification in government or legal domains.
  • Which paper first established the "Economic Activity" taxonomy used in European regulatory bodies, and how does ASAE's version differ?
  • Explore research applying Large Language Models (LLMs) to zero-shot or few-shot classification of noisy administrative complaints in non-English languages.
Contents
Automated Complaint Analysis: Navigating Economic Activities with Machine Learning
1. TL;DR
2. The Administrative Bottleneck: Why This is Hard
3. Methodology: The Search for the Best Pipeline
3.1. 1. Preprocessing Excellence
3.2. 2. The Model Architecture
4. Key Insights: When "Less" is "More"
5. Error Analysis: The Human Element
6. Conclusion & Future Outlook
6.1. Takeaway for Practitioners