Towards Accurate Question Identification: Decomposing Healthcare Forum Posts in Bahasa Indonesia

Towards question identification from online healthcare consultation forum post in bahasa

2017-12-01
Rahmad Mahendra, Abid Nurul Hakim, Mirna Adriani
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a question identification and decomposition module specifically for Bahasa Indonesia healthcare forum posts. It utilizes a hybrid approach—combining SVM-based sentence classification with rule-based chunking—to extract clean, answerable questions from narrative medical inquiries, achieving a 89.9% F1-score in classification.

TL;DR

Researchers from the University of Indonesia have developed a specialized pipeline to transform messy, narrative healthcare forum posts in Bahasa Indonesia into structured questions. By combining SVM classification with rule-based syntactic chunking, they successfully filter out irrelevant "noise" and split complex sentences into distinct questions, achieving nearly 90% F1-score in sentence classification.

Context & Motivation

When patients post on forums like Alodokter or Klikdokter, they don't just ask a question; they tell a story. A typical post might include a child's age, a recent immunization history, specific symptoms, and finally, a question.

Traditional Question Answering (QA) systems are designed for simple queries like "What causes a fever?". They fail when faced with a paragraph of Indonesian text where the actual question is buried under greetings and background context. To assist doctors and improve automated responses, the authors identified the need to decompose these narrative posts.

Methodology: The Three-Stage Architecture

The proposed module consists of a logical progression from raw text to refined questions:

1. Rule-Based Sentence Splitting

The system first leverages HTML structures (like <div> or <br>) and punctuation (periods, question marks) to break the narrative into manageable sentence segments.

2. Sentence Classification (SVM)

Each sentence is categorized into one of several classes:

  • BACKGROUND: Information about symptoms or medical history.
  • QUESTION: The core inquiry.
  • IGNORE: Preamble, greetings, or thanks (e.g., "Thanks, Doc").
  • MULTI-QUESTION: Sentences containing more than one inquiry.

The classifier uses four feature groups: N-grams, sentence position/length, question-specific attributes (like the clitic -kah), and a medical dictionary of symptoms and diseases.

System Decomposition Strategy

3. Multi-Question Splitting (The Core Engine)

This is where the system handles complexity. When a sentence contains multiple questions (e.g., "What is the danger and how to treat it?"), a rule-based chunking approach is used. By applying Part-of-Speech (POS) tagging, the system identifies coordinating conjunctions and WH-words to split the sentence into atomic questions.

Chunking Rules for Multi-Questions

Experiments and Performance

The researchers tested their approach on 1,000 instances (3,811 sentences) from various Indonesian health forums.

  • Classification Accuracy: The SVM model reached an F1-measure of 89.9% using Strategy (b), which involves specifically handling sentences that mix background info and questions.
  • Splitting Accuracy: The rule-based chunker achieved 67.2% accuracy for exact matches in splitting complex questions.

Performance Results Contrast

Critical Insight: Why it Works

The success of this method lies in its position-aware classification. The authors noted that "IGNORE" information (greetings) usually appears at the very beginning or end, while "BACKGROUND" information is clustered in the middle. By incorporating the relative position of a sentence as a feature, the SVM can distinguish a symptom described as background from a symptom mentioned within a question.

Future Outlook & Limitations

While highly effective, the system currently relies on manual rules for chunking. The authors suggest that moving toward Machine Learning-based chunking (such as using sequence labeling models like CRFs or Bi-LSTMs) could improve the splitting of multi-questions. Additionally, better handling of the interaction between background and question segments within a single sentence remains a frontier for Bahasa Indonesia NLP.

In conclusion, this work provides a vital bridge for Indonesian healthcare AI, transforming "noisy" human conversation into "clean" data that automated systems can actually understand.

Find Similar Papers

Try Our Examples

  • Search for recent papers on question decomposition in medical QA systems for low-resource languages other than Bahasa Indonesia.
  • Which paper first established the "BACKGROUND-QUESTION-IGNORE" annotation schema for consumer health questions, and how does this study adapt it?
  • Explore newer studies that replace rule-based chunking with transformer-based (BERT/IndoBERT) models for Indonesian medical text decomposition.
Contents
Towards Accurate Question Identification: Decomposing Healthcare Forum Posts in Bahasa Indonesia
1. TL;DR
2. Context & Motivation
3. Methodology: The Three-Stage Architecture
3.1. 1. Rule-Based Sentence Splitting
3.2. 2. Sentence Classification (SVM)
3.3. 3. Multi-Question Splitting (The Core Engine)
4. Experiments and Performance
5. Critical Insight: Why it Works
6. Future Outlook & Limitations