PHCI: Leveraging Collective Intelligence to Detect Falsehoods in Witness Statements
Identifying Corroborated and Contradicted Claims Among Witness Statements Using Post-Hoc Collective Intelligence
This paper introduces a framework called Post-Hoc Collective Intelligence (PHCI) applied to witness statement analysis. It leverages Natural Language Processing (NLP) to extract claims and aggregate "votes" from multiple independent statements to identify contradictions, corroborations, and potential falsehoods.
TL;DR
Researchers at the University of the West Indies have developed Post-Hoc Collective Intelligence (PHCI), a framework that uses NLP to "mine" the wisdom of the crowd hidden in witness statements. By aggregating multiple independent accounts, the system can automatically flag contradictions and predict false claims, achieving a 0.71 correlation with human evaluators while remaining immune to the cognitive biases that plague human judges.
Background: The High Cost of Human Bias
In the justice system, the volume of data can be overwhelming. For instance, a 2016 Commission of Enquiry in Jamaica produced over 15,000 pages of transcripts. Humans task with analyzing this data fall victim to Availability Bias (remembering only recent info) and Anchoring (failing to update beliefs when new evidence emerges).
The authors argue that whereas human judgment is subjective and order-dependent, a computational approach is deterministic—it processes the first statement and the thousandth with the same level of scrutiny.
Methodology: How PHCI Works
The core innovation is treating a set of text documents as a "crowd" that has already voted. Instead of asking people to vote on a claim, the system extracts claims from their existing statements and counts "implicit votes."
1. The PHCI Framework
The process follows five key steps:
- Goal: Find contradictions and representative statements.
- Knowledge Retrieval: Collecting independent witness accounts.
- Knowledge Representation: Using the Stanford Parser to break sentences into Subject-Verb-Object (SVO) assertions.
- Knowledge Aggregation: Comparing every assertion against every other assertion to find matches or oppositions.
- Output: A net agreement score that identifies how much a statement aligns with the "consensus."

2. The Arithmetic of Truth
The "Assertion Score" is calculated by subtracting opposing counts from similar counts. If a claim is contradicted by many but supported by few, it receives a high negative weight. This allows the system to Mathematically "isolate" outliers.

Experimental Results
The researchers tested the system using a video of a police incident. They collected 32 genuine witness statements and added one deliberate outlier containing false claims (e.g., "no weapon was present" when there clearly was a machete).
- Deception Detection: The PHCI algorithm successfully assigned the lowest "StatementScore" to the outlier.
- Human Alignment: The algorithm’s scoring of statement "goodness" correlated strongly with human ratings (r = 0.71).
- Visualization: The system generated a heat-map of statements, where green indicates corroborated claims and red indicates contradictions.

Critical Insight: Why This Matters
The most profound takeaway is that truth often emerges from the overlap of independent perspectives. While one witness might be mistaken or lying, it is statistically unlikely that a dozen independent witnesses would align on the same lie.
However, the system has limits. It relies on a "majority view." In cases of systemic bias (where an entire crowd shares the same prejudice), the algorithm might support a "consensual lie."
Conclusion & Future Work
PHCI offers a powerful "second opinion" for legal professionals. It doesn't replace human judges but acts as a spotlight, pointing them toward exactly where statements conflict. Future versions aim to incorporate Machine Learning to better handle context and word-sense disambiguation—ensuring that if one person says "cutlass" and another says "machete," the system knows they are describing the same threat.
Takeaway: In an era of information overload, the most reliable "intelligence" might already be sitting in our archives, waiting for the right aggregation algorithm to unlock it.
