Building the First Chinese Argumentation Corpus: A Crowdsourcing Triumph in Review Mining
Crowdsourcing argumentation structures in Chinese hotel reviews
This paper introduces the first Chinese argumentation corpus for hotel reviews, comprising 4,814 argument components and 411 relations. The authors propose an extended "Premise-Claim" model and utilize a crowdsourcing-based annotation workflow integrated with a K-means clustering algorithm for label aggregation.
Executive Summary
TL;DR: This paper presents the first large-scale Chinese argumentation corpus derived from hotel reviews. By leveraging crowdsourcing and an innovative K-means clustering aggregation method, the authors successfully mapped complex discourse structures—claims, premises, and relations—within the noisy, informal landscape of customer feedback.
Positioning: This work bridges the gap between expert-heavy argumentation theory and practical, large-scale data engineering. It moves beyond the traditional "legal/essay" datasets into the commercially vital but technically challenging domain of user-generated content (UGC).
Problem & Motivation: The "Chaos" of Customer Reviews
Argumentation mining (AM) is the art of teaching machines to understand why someone holds an opinion. While we have robust corpora for English persuasive essays, Chinese customer reviews present unique hurdles:
- Informal Structure: Unlike legal briefs, reviews lack standard punctuation and consistent logical flow.
- Implicit Logic: Users often state facts (e.g., "The subway is 5 mins away") that serve as evidence for unstated claims ("The location is good").
- Annotation Subjectivity: What one person sees as a "Premise," another might see as a "Claim," leading to low Inter-Rater Agreement (IRA).
The authors recognized that to build a Chinese AM system, they first needed a data pipeline that could handle the inherent messiness of crowdsourced labor and subjective interpretation.
Methodology: The Extended Argumentation Model
The authors didn't just use a basic "Claim-Premise" structure. They extended it to fit the nuances of reviews:
- MajorClaim: Overall sentiment (e.g., "Highly recommended!").
- Claim: Specific attributes (e.g., "The bed was soft").
- Premise: Facts supporting a claim.
- PSIC (Premise Supporting Implicit Claim): Facts supporting a claim that isn't explicitly written.
Architecture of Annotation Aggregation
To solve the problem of conflicting human labels, the authors treated annotation as a clustering problem. Instead of simple voting, they represented each character's label as a vector and used K-means clustering to find the "centroid" of opinion. This naturally handled boundary disputes (where a sentence starts/ends) and provided a confidence score for every label.
Figure 1: The proposed argumentation model depicting legal relations between MajorClaims, Claims, and Premises.
Experiments & Results: Quality Through Filtering
The study involved 388 students. To ensure quality, the authors used "Gold Standard" reviews to identify and remove "less-devoted" annotators (those who missed easy examples).
Key Findings:
- Quality Boost: Removing low-quality workers improved the agreement score (αU) significantly across the board.
- The "Easy vs. Controversial" Split: The authors cleverly split the results into an "Easy Reviews Corpus" (where agreement was naturally high) and a "Less-Controversial Sentences Corpus" (extracting clear nuggets from otherwise messy reviews).
Table 3: IRA scores confirming that for "Easy" reviews, the crowdsourced quality is comparable to expert-annotated English datasets.
Error Analysis
The most significant confusion occurred between PSIC (Implicit Premises) and Non-Argumentative (NA) text. This highlights the "interpretation gap"—some annotators see a mention of a "nearby bar" as relevant evidence for hotel quality, while others see it as irrelevant chatter.
Critical Insight & Conclusion
The true value of this paper lies in its probabilistic approach to truth. By introducing confidence scores for annotations, the authors acknowledge that human language logic isn't always binary.
Takeaway: If you are building a dataset for a subjective task, don't force a consensus where none exists. Use clustering to find the center of gravity and maintain confidence scores to help your downstream model understand which examples are "hard cases."
Limitations: The reliance on university students (novice annotators) still requires heavy post-processing. Future work could benefit from providing workers with a predefined list of "Implicit Claims" to reduce the confusion between PSIC and NA labels.
