Personalizing Public Health: Identifying Target Audiences for Breast Cancer Screening via Social Media Mining

Preliminary Modeling of Gender Identification in the Social Community

2018-07-01
Kuo-Chung Chu, Peng-Hua Jiang, Huan-Yu Syu, Yan-Ling Chen
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a preliminary framework for identifying user demographics and personalities—specifically gender—by applying text mining and the Classification and Regression Tree (CART) machine learning algorithm to anonymous social media comments (PTT, Mobile01, Facebook) concerning breast cancer screening. The core goal is to enable personalized health promotion policies to improve the alarmingly low screening rates in Taiwan.

TL;DR

To combat low participation in breast cancer screening in Taiwan, researchers are turning to social media anonymity. By using Text Mining and the CART (Classification and Regression Tree) algorithm, this study develops a model to identify the demographic and personality traits of users discussing health online. A preliminary simulation achieved 77% accuracy in gender identification, paving the way for targeted, personality-driven health promotion campaigns.

Background: The "Privacy Barrier" in Healthcare

Breast cancer is the leading cancer among women in Taiwan, yet its screening rate is paradoxically the lowest among free government tests. The root cause is often psychological: shyness and the private nature of the exam. Traditional research—surveys and phone calls—often fails because people are reluctant to discuss sensitive health issues with strangers.

The authors' central insight is that anonymity breeds honesty. On platforms like Facebook and PTT, users express their "True Beliefs." By mining this data, health authorities can understand the specific fears and lifestyles of different personality groups without invasive questioning.

Methodology: From Unstructured Text to Personality Profiles

The research framework follows a sophisticated pipeline from raw social data to actionable insights:

  1. Data Collection: Crawling comments and replies from major Taiwanese social communities (PTT, Mobile01, Facebook).
  2. Feature Extraction: Using NLP to identify keywords linked to the Big Five Personality Traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism).
  3. CART Decision Tree: Utilizing the Gini Index to split data into homogeneous groups, the model categorizes users based on their linguistic patterns.

Conceptual Framework Figure 1: The proposed conceptual framework for personality and gender identification via social media mining.

The Logic of the Decision Tree

The authors opted for the CART algorithm because of its interpretability. In medical policy, "Black Box" models are often rejected. CART provides clear rules:

  • Gini Index: Used to determine the best "split" at each node to reduce impurity.
  • Pruning: Used after building the tree to prevent Overfitting, ensuring the model generalizes well to new social media posts.

Experimental Case Study: Gender Identification

Before tackling complex personality types, the authors validated the model on a gender identification task using the NLTK corpus. By analyzing structural features of names (such as the last two letters), the CART model successfully mapped linguistic markers to gender.

ModelAccuracyParameters
CART0.77criterion="gini", max_depth=5, min_samples_split=10

Model Performance Table 1: Preliminary results showing the feasibility of name-based identification using CART.

Critical Insight: Why This Matters

The shift from "what was said" to "who is saying it" is a game-changer for public health. For example:

  • If a user shows Neurotic traits (high anxiety) in their comments, a campaign might focus on "painless screening" and emotional support.
  • If a user is identified as Conscientious, the campaign could emphasize the "efficiency" and "preventative success rates" of the procedure.

Conclusion & Future Outlook

While current results focus on gender, the roadmap is clear: scaling this to identifying the "Big Five" traits in Chinese-language text. The limitations of this study include its current reliance on structural name features rather than full semantic sentiment. However, as an initial model, it demonstrates that Machine Learning can effectively bypass the "shyness" barrier in East Asian healthcare, allowing for a future where health advertisements are as precisely targeted as e-commerce recommendations.

Takeaway for Researchers

Precision medicine is not just about genetics; it is about Precision Communication. This study proves that social digital footprints are a viable proxy for psychological screening.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize the Big Five personality traits and machine learning to predict health-seeking behaviors from social media text.
  • Which paper originally established the use of CART (Classification and Regression Trees) for text-based demographic identification, and how does this study refine that approach for the medical domain?
  • Explore how deep learning models like BERT or RoBERTa have improved upon the CART algorithm's 77% accuracy in gender and personality identification tasks within Chinese social media contexts.
Contents
Personalizing Public Health: Identifying Target Audiences for Breast Cancer Screening via Social Media Mining
1. TL;DR
2. Background: The "Privacy Barrier" in Healthcare
3. Methodology: From Unstructured Text to Personality Profiles
3.1. The Logic of the Decision Tree
4. Experimental Case Study: Gender Identification
5. Critical Insight: Why This Matters
6. Conclusion & Future Outlook
6.1. Takeaway for Researchers