August 20, 2026

Screening calibration exercises for research teams

Screening calibration is the practice of having reviewers apply the same inclusion and exclusion criteria to a small set of records before full screening begins. It helps teams find disagreements early, clarify ambiguous criteria, and.

Written byWisPaper TeamAI Research Workflow Team
Editorial cover for Screening calibration exercises for research teams

Screening calibration is the practice of having reviewers apply the same inclusion and exclusion criteria to a small set of records before full screening begins. It helps teams find disagreements early, clarify ambiguous criteria, and reduce inconsistent decisions.

AI-assisted workflows make calibration even more useful. If AI helps prioritize records or suggest screening reasons, reviewers still need a shared understanding of what counts as included, excluded, uncertain, or out of scope.

This guide explains how to run screening calibration exercises for literature review teams.

What is screening calibration?

Screening calibration is a short practice round before full screening. Multiple reviewers screen the same records, compare decisions, and revise criteria if needed.

It helps answer:

  • Do reviewers understand the criteria the same way?
  • Which exclusion reasons are unclear?
  • Which records create disagreement?
  • Do criteria need examples?
  • Is the screening form usable?
  • Should the protocol be revised before full work begins?

Calibration is not a test of reviewer intelligence. It is a test of whether the workflow is clear.

Why do research teams need calibration?

Research teams need calibration because written criteria often look clearer than they are. Reviewers may interpret the same phrase differently.

Calibration can reveal:

  • Broad inclusion criteria.
  • Overlapping exclusion reasons.
  • Ambiguous source types.
  • Unclear population boundaries.
  • Confusion about study design.
  • Different tolerance for uncertainty.
  • Missing conflict-resolution rules.

Finding these issues after thousands of records have been screened is painful. Finding them during calibration is useful.

For criteria design, see inclusion and exclusion criteria for literature reviews.

When should calibration happen?

Calibration should happen before full title and abstract screening begins. It should also happen again if the review question, criteria, or source set changes.

Run calibration:

  • After criteria are drafted.
  • Before main screening.
  • Before full-text screening if criteria change.
  • When adding a new reviewer.
  • When disagreement rates are high.
  • When AI-assisted ranking changes the workflow.

Calibration is not one meeting. It is a check whenever the decision rules may have shifted.

How many records should be used?

Use enough records to expose ambiguity without slowing the project too much. The exact number depends on review size, but the set should include easy, hard, and borderline examples.

Include:

  • Clearly relevant records.
  • Clearly irrelevant records.
  • Borderline records.
  • Different study designs.
  • Different source types.
  • Records likely to cause disagreement.

The point is not a representative sample in a statistical sense. The point is to stress-test the criteria.

How do you choose records for calibration?

Choose records from the actual search results when possible. Artificial examples may miss the messiness of the real source set.

Select:

  • Records from different databases or sources.
  • Records with similar terminology but different scope.
  • Records using unexpected methods.
  • Records with unclear abstracts.
  • Records near the inclusion boundary.
  • Records suggested by AI as relevant and not relevant, if applicable.

If all calibration records are easy, the exercise will not reveal much.

One useful tactic is to include records that look relevant for different reasons. Add one record that matches the topic but not the method, one that matches the method but not the population, one that uses the right keywords in a different context, and one that lacks enough abstract detail. These records make reviewers practice the exact judgment calls they will face later.

What should reviewers record during calibration?

Reviewers should record their decision and reason. A decision without a reason is hard to discuss.

Record:

  • Include, exclude, or unsure.
  • Exclusion reason.
  • Criteria used.
  • Confidence level.
  • Notes on ambiguity.
  • Suggested criteria change.
  • Whether full text is needed.

This creates a conversation about rules, not personal preference.

For screening stages, see title and abstract screening vs full-text screening.

How should disagreements be handled?

Disagreements should be discussed as criteria problems first. Ask why reviewers made different decisions before deciding who is right.

For each disagreement, ask:

  • Which criterion was applied?
  • Was the abstract unclear?
  • Did reviewers use different definitions?
  • Was the exclusion reason missing?
  • Should uncertain records move forward?
  • Does the protocol need an example?

Then revise criteria or add examples. Do not rely only on verbal agreement if the issue will recur.

How does AI affect screening calibration?

AI can affect calibration if it ranks records, suggests decisions, summarizes abstracts, or proposes exclusion reasons. Reviewers need to know how to use those outputs.

Define:

  • Whether reviewers can see AI suggestions.
  • Whether AI labels are advisory.
  • How to handle disagreement with AI.
  • Whether AI-prioritized records are included in calibration.
  • How to record tool-assisted decisions.
  • Whether audit samples are required.

AI should make the workflow more inspectable, not less.

For AI screening concepts, see what is active learning screening in systematic reviews.

What should happen after calibration?

After calibration, update the screening rules before main work begins.

Actions may include:

  • Revise inclusion criteria.
  • Split exclusion reasons.
  • Add examples.
  • Clarify uncertain cases.
  • Update the screening form.
  • Define conflict resolution.
  • Decide whether another calibration round is needed.

Do not treat calibration as complete until reviewers know how to handle the most common borderline cases.

How do you turn screening calibration exercises for research teams into a repeatable workflow?

Turn the advice into a repeatable workflow by defining the decision you need to make, the evidence required for that decision, and the record that will prove how the decision was made. In screening, the problem is rarely one missing tool. The problem is usually that search, reading, checking, and writing happen in separate places without a shared rule.

Use a short operating routine:

  • Name the review question or subquestion.
  • Define the source set you are working from.
  • Decide what counts as enough evidence for the next step.
  • Apply the same criteria to every paper in that step.
  • Mark uncertain cases instead of forcing a clean answer.
  • Keep source locations for claims that may enter the final review.
  • Review the workflow after each major search, screening, or writing session.

This routine keeps the work moving without making the review careless. It also gives supervisors, collaborators, and future you a way to understand why the source set changed.

What should you record while using this workflow?

Record the pieces that would be hard to reconstruct later. You do not need a diary of every click, but you do need enough detail to explain the path from question to source to claim.

For this topic, the most useful record usually includes criteria, reviewer decisions, exclusion reasons, audit checks, and source flow. Add the date, tool or source used, reviewer status, and next action. If AI assisted the step, write down what it helped with and what a human checked.

The record should distinguish discovery from evidence. A tool may help find a paper, but the paper itself must support the claim. A summary may help triage a source, but the original source should support any statement that appears in the literature review.

What should you check before writing from this work?

Before writing, check whether the workflow has produced usable evidence or only useful notes. Notes help you think. Evidence supports a sentence.

Ask:

  • Which claim will this source support?
  • Is the claim narrower than the evidence?
  • Have methods, sample, outcome, or concept details been checked?
  • Are limitations visible?
  • Are conflicting papers handled rather than ignored?
  • Is the citation real, current, and relevant?
  • Can another reader understand how this source entered the review?

If the answer is unclear, keep the point in notes rather than moving it into the draft. This is the small pause that prevents AI-assisted research from becoming polished but weak writing.

How can WisPaper support calibration preparation?

WisPaper can help teams build and inspect the candidate paper set used before calibration. Deep Search, Scholar Agent, and Inspiration Discovery support natural-language academic search, while paper cards show source labels, summaries, authors, publication details, and preview images.

Researchers can save or upload candidate papers into My Library, creating a working source set for discussion. Library QA can answer questions based on the user's own library, which may help reviewers inspect whether selected papers share methods, populations, or concepts.

WisPaper supports discovery and source inspection. Calibration decisions, criteria changes, and final screening rules remain the team's responsibility.

Try WisPaper

FAQs

It is strongly useful for team reviews and formal reviews. It helps prevent inconsistent screening before the main workload begins.