From Messy Text to Clear Intent: Mastering Elicitation via Crowdsourcing Knowledge Graphs

Crowdsourcing service requirement oriented requirement pattern elicitation method

2019-10-19
Zhiying Tu, Mengyao Lv, Xiaofei Xu, Zhongjie Wang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel requirement pattern elicitation method oriented towards crowdsourcing services. By constructing a large-scale Knowledge Graph (KG) from unstructured requirement texts on platforms like Freelancer.com, the method extracts frequent demand sequences and domain-oriented association rules to enhance cognitive services' ability to understand and predict customer intentions.

TL;DR

Understanding what a customer actually wants from a fuzzy online post is the "Holy Grail" of cognitive AI. This paper presents a systematic method to mine Requirement Patterns from crowdsourcing platforms (like Freelancer.com). By building a domain-aware Knowledge Graph and applying advanced embedding models (CPTransE), the authors turn unstructured "word clouds" into actionable, predictive requirement templates.

The "Fuzzy" Problem: Why Intent is Hard to Catch

Current virtual assistants can handle simple commands like "Play music," but they stumble when faced with complex, multi-domain requests. The data researchers need is hidden in crowdsourcing platforms, but it comes with three major headaches:

  1. Sparsity: Valuable insights are buried in "non-structural" natural language.
  2. Heterogeneity: Different users use different words for the same thing (Synonyms/Ambiguity).
  3. The Long-Tail: Most systems prioritize popular requests, ignoring unique specialized needs (the long-tail) that are often high-value.

Methodology: The Architecture of Understanding

The authors don't just "read" the text; they reconstruct it. The workflow is divided into a robust offline pipeline and an intelligent online service.

1. Building the Foundation: Ontology & Fusion

First, they extract triplets (Head, Relation, Tail) using Stanford CoreNLP. To bridge the gap between different posts, they developed a hybrid alignment strategy:

  • Element-based: Using an evolved WordNet similarity to handle semantic overlaps.
  • Structure-based: Using Shared Nearest Neighbor (SNN) to verify if two entities are the same based on their "friends" (neighbors) in the graph.

Methodology Framework

2. Mining the "Gold": Pattern Elicitation

The paper defines two key patterns:

  • Link Patterns: Sequences of demands (e.g., "Website" -> "Needs Logo" -> "Needs HTML").
  • Cluster Patterns: Neighborhoods of related concepts.

To handle the Long-Tail, the authors introduced a dynamic threshold function . Instead of a fixed frequency cut-off, it uses a logarithmic standard to ensure that rare but highly correlated requirement pairs are not discarded.

3. Deep Learning Intent: CPTransE

The "secret sauce" is CPTransE. It builds on the TransE model () but adds:

  • Path Reliability: Evaluation of multi-step connections.
  • Relation Clustering: Since crowdsourcing has thousands of unique predicates (relations), they use Affinity Propagation to cluster them, reducing noise and boosting prediction accuracy.

Experimental Showdown

The method was tested against nearly 10,000 project requirements. The results were telling:

  • Alignment Accuracy: Their hybrid method achieved 95% precision, significantly higher than standard WordNet approaches (69%).
  • Relation Prediction: By using relation clustering (the "C" in CPTransE), the model's ability to accurately predict missing requirements (Hit@1) jumped to 60.2%.

Graph Synthesis Process

Case Study: From "I want a website" to a Full Spec

In a real-world demo, a user enters: "I want to create a website."

  1. The system identifies the Domain Term: "website."
  2. It triggers Cluster Patterns: Suggesting "Logo," "PHP," and "SQL Database."
  3. It follows Link Patterns: Warning the user that "Website" usually implies "Homepage" which requires "Expert in HTML."

The fuzzy request is automatically expanded into a professional requirement template.

Critical Insight & Conclusion

The true value of this work lies in its Multi-Granularity approach. It doesn't treat all requirements as equal; it uses Information Entropy to find domain-specific "Anchor" terms and then builds a web of probability around them.

Limitations: While the precision is high, the "Recall" in diverse domains like "IT" and "Mobile" is lower due to heavy overlapping of terms. Future work will likely need to incorporate Attention Mechanisms or Cause-Effect Aspects to better navigate these dense, multi-functional domains.

Final Takeaway: This is a blueprint for moving AI from a passive "listener" to an active "co-author" in the requirement engineering process.

Find Similar Papers

Try Our Examples

  • Search for recent papers using Knowledge Graphs for automatic software requirement elicitation or intention recognition in crowdsourcing environments.
  • Which paper first proposed the PTransE (Path-based TransE) model for knowledge graph embedding, and how does this paper adapt it for requirement engineering?
  • Explore research that applies the "long-tail" requirement analysis or Information Entropy-based domain term extraction to other NLP tasks in the service computing domain.
Contents
From Messy Text to Clear Intent: Mastering Elicitation via Crowdsourcing Knowledge Graphs
1. TL;DR
2. The "Fuzzy" Problem: Why Intent is Hard to Catch
3. Methodology: The Architecture of Understanding
3.1. 1. Building the Foundation: Ontology & Fusion
3.2. 2. Mining the "Gold": Pattern Elicitation
3.3. 3. Deep Learning Intent: CPTransE
4. Experimental Showdown
5. Case Study: From "I want a website" to a Full Spec
6. Critical Insight & Conclusion