Bridging Linguistics and Health: Ontology-Based Surveillance of Turkish Twitter

Ontology-based automatic identification of public health-related Turkish tweets

2017-02-05
Emine Ela Küçük, Kürsad Yapar, Dilek Küçük, Dogan Küçük
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces an ontology-based system for the automatic identification of public-health related Turkish tweets. It leverages a semi-automatically constructed public health ontology, enhanced with linguistic relaxation and inflectional expansion, to filter the Twitter stream and provide real-time surveillance capabilities.

TL;DR

Researchers have developed the first specialized system for tracking public health trends in Turkish tweets. By combining a semi-automatically built ontology with sophisticated linguistic expansion—handling Turkish's complex grammar and "informal" spelling—the system filters millions of tweets to identify disease outbreaks, symptoms, and medication trends with over 70% precision.

Background & Motivation: Why Turkish Twitter?

Public health surveillance has historically relied on slow, manual reporting. Social media changed this for English-speaking regions, but Turkish remained a frontier. Turkish is an agglutinative language, meaning a single root word like hastane (hospital) can take many forms (hastanede, hastaneyle, etc.). Furthermore, Turkish users often drop diacritics (e.g., writing 'seker' instead of 'ÅŸeker'), rendering standard keyword searches ineffective.

Methodology: The Anatomy of a Linguistic-Aware Ontology

The core of this research is not just a list of words, but a structured Knowledge Base. The development followed a 7-stage pipeline:

  1. Core Concepts: Defining hierarchies for Disease, Symptom, Medication, and General Public Health.
  2. Web Mining: Extracting terms from Wikipedia and medical sites.
  3. Metaphor Filtering: A crucial manual step where terms like "paralyzed" (often used for traffic) were removed to prevent false positives.
  4. Diacritics Expansion: Automatically generating synonyms to account for informal typing (e.g., grip → grıp).
  5. Inflectional Expansion: This is the "secret sauce." Since Turkish uses suffixes for cases (Ablative, Dative, etc.), the authors implemented a rule-based system to generate all possible grammatical forms of each health term.

Overall Ontology Development Process The semi-automated pipeline showing the blend of manual expertise and automated expansion.

Experiments and Insights

The system was tested on two massive datasets of 1 million tweets each, representing different seasons (Late Winter vs. Late Summer).

  • Volume: The system identified approximately 0.15% of the total stream as health-related.
  • Precision: It maintained a solid 71.5% overall precision.
  • Seasonality: The system successfully captured shifts in public concern—flu and "cold" terms dominated the winter set, while "stomach ache" and "obesity" saw relative increases in the summer set.

Performance by Concept Analysis showing that 'Disease' and 'General Health' concepts provided the highest precision, while 'Symptom' terms introduced more noise.

Beyond Keywords: The Hybrid SVM Approach

The authors didn't stop at a rule-based tool. They used the high-quality data filtered by their ontology to train a Support Vector Machine (SVM) classifier. By using the ontology as a "teacher," the machine learning model achieved an F-Measure of 94.8%. This proves that high-quality linguistic knowledge is the best foundation for training robust AI models.

Critical Analysis & Limitations

While the precision is impressive, the system faces challenges typical of social media:

  • Metaphorical Language: Despite filtering, some phrases like "I have tuberculosis" are still used as idioms for extreme sadness, triggering false positives.
  • Ambiguity: Terms related to hospitals often appear in "chatter" following sports matches (e.g., "The referee belongs in a hospital").
  • Recall: With a recall of approximately 18.5%, there is significant room to improve coverage by including more informal slang and misspellings.

Conclusion

This study serves as a vital blueprint for public health experts in non-English speaking countries. It demonstrates that you don't need a massive manually labeled dataset to start—you can build a linguistically intelligent ontology and let it generate the data for you.

The project also provided a web-based interface for public health staff, moving the research from academic theory into a practical tool for real-world epidemic tracking.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize deep learning or transformer-based models (such as BERTurk) for public health surveillance in Turkish social media.
  • Which research first established the methodology for semi-automatic ontology construction from Wikipedia for low-resource or agglutinative languages?
  • How have state-of-the-art systems addressed the problem of metaphorical language use (e.g., "traffic is paralyzed") in medical text mining for social media?
Contents
Bridging Linguistics and Health: Ontology-Based Surveillance of Turkish Twitter
1. TL;DR
2. Background & Motivation: Why Turkish Twitter?
3. Methodology: The Anatomy of a Linguistic-Aware Ontology
4. Experiments and Insights
5. Beyond Keywords: The Hybrid SVM Approach
6. Critical Analysis & Limitations
7. Conclusion