Towards Accurate Question Identification: Decomposing Healthcare Forum Posts in Bahasa Indonesia
Towards question identification from online healthcare consultation forum post in bahasa
The paper introduces a question identification and decomposition module specifically for Bahasa Indonesia healthcare forum posts. It utilizes a hybrid approach—combining SVM-based sentence classification with rule-based chunking—to extract clean, answerable questions from narrative medical inquiries, achieving a 89.9% F1-score in classification.
TL;DR
Researchers from the University of Indonesia have developed a specialized pipeline to transform messy, narrative healthcare forum posts in Bahasa Indonesia into structured questions. By combining SVM classification with rule-based syntactic chunking, they successfully filter out irrelevant "noise" and split complex sentences into distinct questions, achieving nearly 90% F1-score in sentence classification.
Context & Motivation
When patients post on forums like Alodokter or Klikdokter, they don't just ask a question; they tell a story. A typical post might include a child's age, a recent immunization history, specific symptoms, and finally, a question.
Traditional Question Answering (QA) systems are designed for simple queries like "What causes a fever?". They fail when faced with a paragraph of Indonesian text where the actual question is buried under greetings and background context. To assist doctors and improve automated responses, the authors identified the need to decompose these narrative posts.
Methodology: The Three-Stage Architecture
The proposed module consists of a logical progression from raw text to refined questions:
1. Rule-Based Sentence Splitting
The system first leverages HTML structures (like <div> or <br>) and punctuation (periods, question marks) to break the narrative into manageable sentence segments.
2. Sentence Classification (SVM)
Each sentence is categorized into one of several classes:
- BACKGROUND: Information about symptoms or medical history.
- QUESTION: The core inquiry.
- IGNORE: Preamble, greetings, or thanks (e.g., "Thanks, Doc").
- MULTI-QUESTION: Sentences containing more than one inquiry.
The classifier uses four feature groups: N-grams, sentence position/length, question-specific attributes (like the clitic -kah), and a medical dictionary of symptoms and diseases.

3. Multi-Question Splitting (The Core Engine)
This is where the system handles complexity. When a sentence contains multiple questions (e.g., "What is the danger and how to treat it?"), a rule-based chunking approach is used. By applying Part-of-Speech (POS) tagging, the system identifies coordinating conjunctions and WH-words to split the sentence into atomic questions.

Experiments and Performance
The researchers tested their approach on 1,000 instances (3,811 sentences) from various Indonesian health forums.
- Classification Accuracy: The SVM model reached an F1-measure of 89.9% using Strategy (b), which involves specifically handling sentences that mix background info and questions.
- Splitting Accuracy: The rule-based chunker achieved 67.2% accuracy for exact matches in splitting complex questions.

Critical Insight: Why it Works
The success of this method lies in its position-aware classification. The authors noted that "IGNORE" information (greetings) usually appears at the very beginning or end, while "BACKGROUND" information is clustered in the middle. By incorporating the relative position of a sentence as a feature, the SVM can distinguish a symptom described as background from a symptom mentioned within a question.
Future Outlook & Limitations
While highly effective, the system currently relies on manual rules for chunking. The authors suggest that moving toward Machine Learning-based chunking (such as using sequence labeling models like CRFs or Bi-LSTMs) could improve the splitting of multi-questions. Additionally, better handling of the interaction between background and question segments within a single sentence remains a frontier for Bahasa Indonesia NLP.
In conclusion, this work provides a vital bridge for Indonesian healthcare AI, transforming "noisy" human conversation into "clean" data that automated systems can actually understand.
