From Messy Text to Clear Intent: Mastering Elicitation via Crowdsourcing Knowledge Graphs
Crowdsourcing service requirement oriented requirement pattern elicitation method
This paper introduces a novel requirement pattern elicitation method oriented towards crowdsourcing services. By constructing a large-scale Knowledge Graph (KG) from unstructured requirement texts on platforms like Freelancer.com, the method extracts frequent demand sequences and domain-oriented association rules to enhance cognitive services' ability to understand and predict customer intentions.
TL;DR
Understanding what a customer actually wants from a fuzzy online post is the "Holy Grail" of cognitive AI. This paper presents a systematic method to mine Requirement Patterns from crowdsourcing platforms (like Freelancer.com). By building a domain-aware Knowledge Graph and applying advanced embedding models (CPTransE), the authors turn unstructured "word clouds" into actionable, predictive requirement templates.
The "Fuzzy" Problem: Why Intent is Hard to Catch
Current virtual assistants can handle simple commands like "Play music," but they stumble when faced with complex, multi-domain requests. The data researchers need is hidden in crowdsourcing platforms, but it comes with three major headaches:
- Sparsity: Valuable insights are buried in "non-structural" natural language.
- Heterogeneity: Different users use different words for the same thing (Synonyms/Ambiguity).
- The Long-Tail: Most systems prioritize popular requests, ignoring unique specialized needs (the long-tail) that are often high-value.
Methodology: The Architecture of Understanding
The authors don't just "read" the text; they reconstruct it. The workflow is divided into a robust offline pipeline and an intelligent online service.
1. Building the Foundation: Ontology & Fusion
First, they extract triplets (Head, Relation, Tail) using Stanford CoreNLP. To bridge the gap between different posts, they developed a hybrid alignment strategy:
- Element-based: Using an evolved WordNet similarity to handle semantic overlaps.
- Structure-based: Using Shared Nearest Neighbor (SNN) to verify if two entities are the same based on their "friends" (neighbors) in the graph.

2. Mining the "Gold": Pattern Elicitation
The paper defines two key patterns:
- Link Patterns: Sequences of demands (e.g., "Website" -> "Needs Logo" -> "Needs HTML").
- Cluster Patterns: Neighborhoods of related concepts.
To handle the Long-Tail, the authors introduced a dynamic threshold function . Instead of a fixed frequency cut-off, it uses a logarithmic standard to ensure that rare but highly correlated requirement pairs are not discarded.
3. Deep Learning Intent: CPTransE
The "secret sauce" is CPTransE. It builds on the TransE model () but adds:
- Path Reliability: Evaluation of multi-step connections.
- Relation Clustering: Since crowdsourcing has thousands of unique predicates (relations), they use Affinity Propagation to cluster them, reducing noise and boosting prediction accuracy.
Experimental Showdown
The method was tested against nearly 10,000 project requirements. The results were telling:
- Alignment Accuracy: Their hybrid method achieved 95% precision, significantly higher than standard WordNet approaches (69%).
- Relation Prediction: By using relation clustering (the "C" in CPTransE), the model's ability to accurately predict missing requirements (Hit@1) jumped to 60.2%.

Case Study: From "I want a website" to a Full Spec
In a real-world demo, a user enters: "I want to create a website."
- The system identifies the Domain Term: "website."
- It triggers Cluster Patterns: Suggesting "Logo," "PHP," and "SQL Database."
- It follows Link Patterns: Warning the user that "Website" usually implies "Homepage" which requires "Expert in HTML."
The fuzzy request is automatically expanded into a professional requirement template.
Critical Insight & Conclusion
The true value of this work lies in its Multi-Granularity approach. It doesn't treat all requirements as equal; it uses Information Entropy to find domain-specific "Anchor" terms and then builds a web of probability around them.
Limitations: While the precision is high, the "Recall" in diverse domains like "IT" and "Mobile" is lower due to heavy overlapping of terms. Future work will likely need to incorporate Attention Mechanisms or Cause-Effect Aspects to better navigate these dense, multi-functional domains.
Final Takeaway: This is a blueprint for moving AI from a passive "listener" to an active "co-author" in the requirement engineering process.
