Collective Intelligence: Automating the "Daunting Task" of Constraint Elicitation
Towards Automatic Constraint Elicitation in Test Design: Preliminary Evaluation Based on Collective Intelligence
This paper introduces a collective intelligence-based approach to automate constraint elicitation in combinatorial test design, specifically for pair-wise testing. By utilizing web search engine hit counts to measure the "rate" of value combinations, the authors identify potential constraints without requiring manual domain expertise.
TL;DR
Determining what combinations of inputs are invalid for a software system is notoriously difficult. This paper proposes a clever shortcut: using the "Collective Intelligence" of the entire web. By analyzing Google search results, the authors can automatically predict which parameter values (like OS and Browser combinations) are valid or represent constraints, reducing the need for manual expert intervention.
Background: The Constraint Bottleneck
In Pair-wise Testing, we aim to test all possible pairs of parameter values. However, not all pairs are possible. For instance, testing "Internet Explorer on Linux" is a waste of time—it doesn't exist. These are Constraints.
The problem? Identifying these constraints is usually a manual, "daunting task" that requires deep domain knowledge. If a test designer misses one, the resulting test suite will contain unexecutable cases, leading to costly re-works.
Methodology: The "Rate" of Truth
The researchers' core insight is that the web reflects real-world constraints. If a combination of values (e.g., "Mac" and "Safari") appears frequently on the web, it is likely a valid, common pairing. If it rarely appears (e.g., "Linux" and "IE"), it likely marks a constraint.
The Formula
They calculate the Rate of a value pair as follows:

- High Rate (>0.5): Indicates a strong dependency (one value often forces the other).
- Low Rate (<0.1): Indicates an uncommon or impossible pairing (a potential exclusion constraint).
The Process Flow
The workflow moves from defining parameters to querying the collective intelligence of the web to eke out rules.

Experimental Insights
The authors tested this on two scenarios: Cross-browser testing and an ATM system.
Case 1: Browser Testing
The results were highly intuitive. For the "Linux" OS, the rate for the "Chrome" browser was 0.524, significantly higher than others, correctly identifying the rule: OS = "Linux" ⇒ Browser = "Chrome". Conversely, IE on Linux had a very low hit count, correctly signaling an incompatibility.
Case 2: ATM System
This was more challenging. While the system correctly identified that "Withdrawal" transactions are highly linked to "Account (B) = Not Selected" (since you only need one account to withdraw), it struggled with generic phrases like "not selected," which appear in many unrelated web contexts.

Critical Analysis: Is Search Enough?
The study proves a vital concept: Software constraints are often public knowledge. However, the "Search Hit" method has two main limitations:
- Semantic Noise: Common words like "not selected" create false positives because they aren't unique to the software domain.
- Over-Restriction: Just because "Chrome on Windows" is the most popular search doesn't mean "Chrome on Mac" is a constraint, even if the rate difference is huge.
Conclusion & Future Work
The paper successfully demonstrates that we can identify constraints without specific domain expertise by tapping into collective intelligence. Moving forward, the authors plan to:
- Refine the query mechanism to reduce noise.
- Merge this "Value-level" extraction with their previous "Parameter-level" work on Coupling Strength.
This research opens the door for self-configuring test design tools that "learn" the boundaries of a system by simply reading the web.
