Multilingual Pornography Detection on Twitter: A Comparative Machine Learning Approach

Twitter Pornography Multilingual Content Identification Based on Machine Learning

2017-01-01
Edo Barfian, Bambang Heru Iswanto, Sani Muhamad Isa
Summary
Problem
Method
Results
Takeaways
Abstract

The paper explores "Twitter Pornography Multilingual Content Identification" using machine learning techniques (Decision Tree, Naive Bayes, and SVM). It focuses on identifying pornographic vs. non-pornographic tweets in Indonesian, English, and a combination of both, utilizing TF-IDF for feature extraction.

    ## TL;DR
    This research addresses the critical challenge of identifying pornographic content on Twitter across multiple languages (Indonesian and English). By benchmarking **Decision Trees**, **Naive Bayes**, and **Support Vector Machines (SVM)** using a TF-IDF framework, the study reveals that while Naive Bayes is highly effective for Indonesian text (92.78% accuracy), SVM with a Linear Kernel is superior for handling the complexities of English and mixed-language environments.

    ## Problem & Motivation
    The proliferation of adult content on social media presents a significant risk to the morale and safety of minors. Prior works often relied on static **URL Blocking** or simple **Keyword Filtering**. However, Twitter poses a unique challenge:
    *   **Dynamic Language**: Natural language on Twitter is irregular, filled with slang, and frequently skips standard grammar.
    *   **Multilingualism**: Users often switch between Indonesian and English, complicating traditional monolingual filters.
    *   **Evasion**: Content creators use creative text variations to bypass simple keyword-based blocks.

    The authors' insight was to move beyond simple filters toward a **Machine Learning classification** approach that analyzes the statistical significance of words (Unigrams) within documents to differentiate between "Positive" (Safe) and "Negative" (Pornographic) content.

    ## Methodology
    The system follows a classic NLP pipeline optimized for the noisy nature of social media data:

    1.  **Pre-processing**: This involves aggressive cleaning (removing links, hashtags, and special characters) followed by **Porter Stemming** (adapted for Indonesian) to reduce words to their base forms.
    2.  **Feature Engineering**: The team utilized **TF-IDF (Term Frequency-Inverse Document Frequency)** to weight the importance of words. They focused on **Unigrams** (single-word features) like "sex" or "porn."
    3.  **Classification**: Three main algorithms were compared:
        *   **Decision Tree**: A flowchart-like structure for attribute testing.
        *   **Naive Bayes**: A probabilistic classifier based on the assumption of independence between variables.
        *   **Support Vector Machines (SVM)**: A high-dimensional mapping technique aimed at finding the optimal hyperplane for classification.

    ![System Architecture](https://cdn.atominnolab.com/wisdoc/images/20260603-fe8e25a9-15a8-4c9a-9d7a-a02639e26f5b/page_002_block_004.png)

    ## Experiments & Results
    The researchers used a dataset of 600 tweets (equally split between pornographic and non-pornographic across three categories: Indonesian, English, and Combined).

    ### Performance Breakdown
    The results highlighted a fascinating discrepancy between language performance:
    *   **Indonesian Dataset**: **Naive Bayes** dominated with an average accuracy of **92.78%**. The authors noted that Indonesian grammar, while complex, showed high regularity in the specific pornographic corpus used.
    *   **English Dataset**: **SVM** took the lead with **83.33%**.
    *   **Combined Dataset**: Performance dropped to **72.72%** (SVM), illustrating the difficulty of "code-switching" where users mix languages in a single post.

    ![Accuracy Comparison Table](https://cdn.atominnolab.com/wisdoc/tables/20260603-fe8e25a9-15a8-4c9a-9d7a-a02639e26f5b/page_004_block_005.png)

    ### The Kernel Selection (SVM)
    A deep dive into SVM kernels for the Indonesian dataset proved that the **Linear Kernel** is most effective for text classification (87.60% average). More complex kernels like Radial or Polynomial actually degraded performance, suggesting that text features are often linearly separable in a high-dimensional TF-IDF space.

    ## Critical Analysis & Conclusion
    ### Takeaway
    The study confirms that a "one-size-fits-all" algorithm is rarely optimal for multilingual NLP. **Naive Bayes** remains a powerful, lightweight tool for specific linguistic structures (like Indonesian), whereas **SVM** provides the necessary robustness for the more fragmented English-speaking "Twitter-verse."

    ### Limitations
    1.  **Dataset Size**: 600 tweets is relatively small for a modern machine learning task; larger datasets would likely reveal more edge-case failures.
    2.  **Context**: Unigrams cannot capture sarcasm or nuanced meanings that depend on word order (Bigrams/Trigrams).

    ### Future Outlook
    With the rise of Large Language Models (LLMs), the next step for this research would be to apply **transformers** to capture the semantic context of tweets, potentially closing the gap in the "Multilingual Combined" category where accuracy currently struggles.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize Deep Learning or Transformers (e.g., mBERT, XLM-R) for multilingual pornography detection on social media to compare against traditional ML baselines.
  • Which study first introduced the Porter Stemmer adaptation for the Indonesian language, and how have modern Indonesian NLP libraries improved upon this for informal Twitter text?
  • Search for research exploring how multimodal fusion (combining the text features from this paper with image recognition) improves the accuracy of adult content filtering systems on Twitter.
Contents
Multilingual Pornography Detection on Twitter: A Comparative Machine Learning Approach
1. TL;DR
2. Problem & Motivation
3. Methodology
4. Experiments & Results
4.1. Performance Breakdown
4.2. The Kernel Selection (SVM)
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook