From Search Logs to Silver Data: Automating Professional QA Systems

Exploiting Search Logs to Aid in Training and Automating Infrastructure for estion Answering in Professional Domains

Filippo Pompili, Thomson Reuters, Jack Conrad, Carter Kolbeck
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a method to build Question Answering (QA) systems in professional legal and regulatory domains by leveraging "silver data"—Implicit Relevance Feedback (IRF) derived from search logs. It demonstrates that user interactions like saving or printing documents can substitute for expensive expert annotations, effectively addressing the cold start problem and boosting SOTA ranking performance.

TL;DR

In specialized domains like law, training a Question Answering (QA) system is notoriously expensive due to the need for expert-level "Gold Data." This paper from Thomson Reuters reveals a breakthrough: by mining "Silver Data" from user activity logs (actions like printing or saving), we can bypass the "Cold Start" problem. Their approach achieves performance comparable to having 25% of a full expert-labeled dataset, without the cost.

The Professional Data Gap: Why "Clicks" Aren't Enough

In general web search, a "click" is a noisy signal. In professional legal research, however, the stakes are higher. A lawyer doesn't just click—they save, print, or export a document when it solves a specific legal query.

The authors argue that these high-intent actions are a goldmine of implicit feedback. Before this study, the industry relied on Subject Matter Experts (SMEs) to manually grade thousands of document pairs—a process that is slow, unscalable, and creates a massive barrier for new entries (the "Cold Start" problem).

Methodology: The "Silver" and the "Imputed"

The research utilizes a two-stage retrieval architecture:

  1. Recall Stage: An internal legacy engine retrieves a broad set of candidates.
  2. Precision Stage (The Re-ranker): A learning-to-rank algorithm re-sorts these candidates based on features like n-gram overlap, embeddings, and parse tree correspondence.

Data Innovation

  • Silver Data: QA pairs where users printed/saved/emailed the result. A validation study confirmed that 90% of these "Silver" documents were indeed relevant.
  • Imputed Negatives: To teach the model what a "bad" answer looks like without asking an expert, the team sampled documents from deep within the search results (e.g., beyond rank 10), assuming they are highly unlikely to be relevant.

Verification of Silver Data Quality Table: SME validation confirming that Silver Data is 64% 'A' grade and 90% 'A' or 'C' grade.

Experimental Insights

The authors tested their hypothesis across two jurisdictions: Federal and State.

1. Solving the Cold Start

In scenarios where 0% gold data was available, the silver data provided an immediate, functional baseline. In the State jurisdiction, the performance leap was staggering—starting at nearly 50% accuracy at Rank 1, whereas sparse gold data only managed 16%.

2. Silver vs. Gold Efficiency

How much expert work does silver data replace?

  • For Federal data, it took roughly 25% of the total gold data to match the performance of automated silver data.
  • For State data, which is more structurally varied, silver data was even more dominant, outperforming small-to-medium gold datasets across the board.

Federal Performance Comparison Figure 1: Baseline performance using only expert-graded Gold Data. Note the low starting point and high variance.

Federal Performance with Silver Data Figure 3: Performance with Silver Data added. The "Cold Start" (0% Gold) begins at a much higher accuracy level (~0.47).

Critical Analysis & Conclusion

This work proves that professional AI doesn't always need a human-in-the-loop for every training sample. By treating user behavior as a proxy for expertise, companies can build "automated infrastructure"—systems that learn and improve every time a user saves a document.

Key Limitations:

  • The "Imputed Negatives" strategy is simple; it doesn't specifically target "hard negatives" (documents that look relevant but aren't).
  • Performance saturation occurs quickly as silver data grows, suggesting that while it's great for starting, high-end "Gold" data is still needed for peak optimization.

Future Outlook: This paves the way for Continuous Integration/Continuous Deployment (CI/CD) for AI models, where the system undergoes "Continuous Training" fed by real-time user logs, drastically reducing the time-to-market for legal and financial tech.

Find Similar Papers

Try Our Examples

  • Search for recent papers on "Learning to Rank" (LTR) that specifically use clickstream data or Implicit Relevance Feedback in the legal or medical domains.
  • Who first formally defined the "imputed negatives" or "pseudo-relevance feedback" sampling strategy used in this paper, and how have modern "Hard Negative Mining" techniques improved upon it?
  • Find research exploring the application of User Activity Log mining for cold-start problems in transformer-based (BERT/Longformer) retrieval systems.
Contents
From Search Logs to Silver Data: Automating Professional QA Systems
1. TL;DR
2. The Professional Data Gap: Why "Clicks" Aren't Enough
3. Methodology: The "Silver" and the "Imputed"
3.1. Data Innovation
4. Experimental Insights
4.1. 1. Solving the Cold Start
4.2. 2. Silver vs. Gold Efficiency
5. Critical Analysis & Conclusion