CHITL: Solving the "Unknown Unknowns" in Relation Extraction with Human-in-the-Loop Denoising

A Crowdsourcing Based Human-in-the-Loop Framework for Denoising UUs in Relation Extraction Tasks

2019-07-01
Mengting Li, Jian Jin, Wen Wu, Yan Yang, Liang He, Jing Yang
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces CHITL, a crowdsourcing-based human-in-the-loop framework designed to denoise "Unknown Unknowns" (UUs) in relation extraction. It utilizes an entity-pair level attention-based LSTM as a selector to identify high-confidence but incorrect labels, achieving state-of-the-art performance on the New York Times (NYT) dataset.

TL;DR

Relation Extraction (RE) under distant supervision is plagued by "Unknown Unknowns" (UUs)—noisy labels where models are confidently wrong. The CHITL (Crowdsourcing Based Human-In-The-Loop) framework tackles this by using an attention-based LSTM to "sniff out" suspicious data points and employing human crowd-workers to fix them, leading to a significant 10% jump in precision for standard CNN models.

The "Confidence Trap": Why Distant Supervision Fails

In the world of NLP, labeling data is expensive. Distant Supervision (DS) was the industry's "free lunch"—it automatically labels text by aligning it with Knowledge Bases (like Freebase). For example, if (Trump, USA) has a "nationality" relation in a database, DS assumes every sentence mentioning both entities expresses that relation.

However, this creates two deadly types of noise:

  1. Wrong Labels: "Trump visited the Apple Inc. in the US" (No actual nationality relation expressed).
  2. Missing Labels: The KB doesn't know "Trump was born in New York," so it labels it as 'NA'.

These are Unknown Unknowns (UUs). They are particularly dangerous because current SOTA models (RL and Adversarial-based) often just "guess" which samples are noisy without any objective ground truth to correct them.

Methodology: The CHITL Framework

The core philosophy of CHITL is simple: Automate the discovery, humanize the correction. It splits the task into three distinct modules:

1. The Intelligent Selector (LSTM + Attention)

Instead of checking every sentence, the framework uses an attention-based LSTM. It calculates Entity-pair Embedding () to focus the model's attention on words that actually define the relationship. If the model's prediction differs from the KB label with high confidence, it flags it as a "Potential UU."

CHITL Framework Architecture

2. Strategic Crowdsourcing

To prevent human error and keep costs low, CHITL uses two filters:

  • Participant Quality: Mixing "Golden Data" (pre-verified answers) into the tasks to screen out low-quality workers.
  • Majority Vote: Ensuring consensus before a label is updated.

3. Iterative Feedback Loop

The newly cleaned data is fed back into the training set. This doesn't just help the final classifier; it also helps the Selector get better at finding more subtle UUs in the next epoch.

Experiments: Superior Denoising Performance

The authors tested CHITL on the classic New York Times (NYT) dataset. The results prove that even a small amount of human intervention can outperform "pure" algorithmic denoising.

Key Performance Gains:

  • Precision @ 100: Jumped from 75.2% to 85.2% for CNN models.
  • Versatility: When added to existing denoising frameworks like ADV (Adversarial) or RL (Reinforcement Learning), CHITL consistently boosted their AUC and F1 scores.

Performance Comparison on NYT Dataset

The table below highlights that for the /people/person/place_lived relation, the selector achieved an impressive hit rate—identifying 162 real UUs out of 280 potential flags.

Annotated Data Table

Critical Insight: Data Quality over Model Complexity

The most striking takeaway is that human-in-the-loop (HITL) frameworks are model-independent. Whether you are using a simple CNN or a complex BiRNN, cleaning the "UUs" provides a performance floor that no amount of hyperparameter tuning can match.

Moreover, CHITL identifies "hidden" relations (Missing Labels) that the model would previously have ignored, effectively expanding the Knowledge Base through the extraction process itself.

Conclusion & Future Outlook

While CHITL effectively bridges the gap between machine confidence and human truth, the reliance on crowdsourcing still incurs a financial cost. Future research could look into Active Learning to further minimize the number of human queries needed or leverage Large Language Models (LLMs) as "proxy humans" to label these UUs at scale.


Keywords: Relation Extraction, Distant Supervision, Human-in-the-loop, Crowdsourcing, Unknown Unknowns, Denoising.

Find Similar Papers

Try Our Examples

  • Find recent papers that extend the concept of Unknown Unknowns (UUs) from relation extraction to other NLP tasks like Named Entity Recognition or Link Prediction.
  • What are the seminal works on the 'Human-in-the-loop' paradigm for denoising distantly supervised datasets before the year 2024?
  • Research other crowdsourcing strategies beyond Majority Vote that are specifically designed for improving the training quality of Large Language Models (LLMs).
Contents
CHITL: Solving the "Unknown Unknowns" in Relation Extraction with Human-in-the-Loop Denoising
1. TL;DR
2. The "Confidence Trap": Why Distant Supervision Fails
3. Methodology: The CHITL Framework
3.1. 1. The Intelligent Selector (LSTM + Attention)
3.2. 2. Strategic Crowdsourcing
3.3. 3. Iterative Feedback Loop
4. Experiments: Superior Denoising Performance
4.1. Key Performance Gains:
5. Critical Insight: Data Quality over Model Complexity
6. Conclusion & Future Outlook