August 10, 2026

AI in Systematic Reviews: A PRISMA-Aware Workflow

AI can support systematic reviews, but it does not remove the need for transparent reporting, protocol discipline, and human responsibility. The question is not simply "Can I use AI?" The better question is: where did AI affect the review, and how will that be reported?

Written byWisPaper TeamAI Research Workflow Team
WisPaper search results for AI-assisted paper screening

AI can support systematic reviews, but it does not remove the need for transparent reporting, protocol discipline, and human responsibility. The question is not simply "Can I use AI?" The better question is: where did AI affect the review, and how will that be reported?

PRISMA is a reporting guideline, not an approval stamp for any software. A review can use AI and still be poorly reported. A review can avoid AI and still have weak methods. PRISMA helps readers understand what was done, what was found, what was excluded, and how the review moved from records to included studies.

This guide gives a PRISMA-aware workflow for AI-assisted systematic reviews. It is written for researchers who want to use AI without making the methods section vague. If you are still deciding whether AI belongs in your workflow at all, start with is using AI for a literature review cheating.

What PRISMA Does And Does Not Do

PRISMA helps authors report systematic reviews and meta-analyses clearly. The PRISMA site says the PRISMA 2020 statement includes a checklist of items and sub-items plus an expanded checklist with reporting recommendations for each item or sub-item. The PRISMA statement paper presents a 27-item checklist, an expanded checklist, an abstract checklist, and template flow diagrams.

PRISMA does not tell you that a tool is valid. It does not certify AI accuracy. It does not decide whether a review is methodologically sound. It asks you to report the work clearly enough that readers can understand and evaluate it.

For AI use, that means you should be able to report:

  • Which tool was used.
  • Which stage it affected.
  • What input it received.
  • What output it produced.
  • Whether humans checked or overrode the output.
  • How the AI-assisted step changed the review process.

If that information is missing, the review may look faster, but it will be harder to trust.

The AI Use Must Fit The Protocol

AI use should be planned before it affects the review. If a protocol already exists, check whether AI-assisted searching, screening, extraction, or synthesis is allowed. If AI is added after the protocol, document the change and why it was made.

Do not let the tool quietly redefine the method. A model's relevance ranking is not an inclusion criterion. A generated summary is not extracted data until it has been checked. A suggested theme is not synthesis until the authors have interpreted the evidence.

Write down:

  • The review question.
  • Eligibility criteria.
  • Search sources and strategies.
  • Screening process.
  • Extraction fields.
  • Synthesis method.
  • AI tools and the stages where they may be used.
  • Human review and conflict-resolution process.

This record matters later. It lets you describe AI use in methods, answer reviewer questions, and distinguish planned assistance from improvised shortcuts.

RAISE And Responsible AI Use

Evidence-synthesis organizations are moving toward explicit responsible-AI expectations. Cochrane describes RAISE as a framework to guide ethical and transparent AI use across evidence synthesis and notes expectations around transparent reporting, responsibility, and preserving methodological rigor in AI-supported evidence synthesis.

The Stockholm Environment Institute page for the position statement says the statement advocates responsible AI and automation in evidence synthesis, emphasizing human oversight, methodological rigor, integrity of findings, and fully transparent reporting across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence.

For working researchers, the message is practical:

  • AI use should be visible.
  • Humans remain responsible.
  • Methods should not become weaker because the tool is convenient.
  • AI output needs review, especially when it affects inclusion, extraction, or conclusions.

That is the standard to carry into the workflow.

AI can help with search planning. It can suggest synonyms, related terms, adjacent disciplines, seed papers, and alternate phrasing for a concept. This is useful when the topic crosses fields or when the first keyword search returns noisy results.

But the final search still needs reporting. A reader should know which databases or sources were searched, which queries were used, and when the searches were run. If AI contributed to query development, say how.

Report:

  • The AI tool used for search planning, if any.
  • The prompts or instructions at a high enough level for readers to understand the task.
  • Which AI-suggested terms were retained or rejected.
  • The final search sources and queries.
  • Any date, language, or publication-type limits.

Use AI to discover better vocabulary, not to hide the search strategy. The workflow in AI academic search beyond Google Scholar is useful at this stage because it separates concept discovery from controlled searching.

Screening: Keep Human Decisions Visible

Screening is one of the most sensitive AI-assisted stages. A tool can rank papers or suggest relevance, but inclusion decisions affect the evidence base.

Cochrane's MECIR standards say to use at least two people working independently to decide whether each study meets eligibility criteria, with a predefined process for resolving disagreements. The same page explains that duplication reduces mistakes and reduces the chance that selection is influenced by a single person's bias.

AI-assisted screening should therefore be documented carefully. Record whether AI ranked records, suggested exclusions, summarized abstracts, or flagged uncertain cases. Record who made the final decision.

Report:

  • Screening criteria.
  • Whether AI was used for prioritization, classification, or summarization.
  • Whether AI suggestions were checked by human reviewers.
  • How disagreements or uncertain cases were resolved.
  • Whether excluded records and reasons were preserved.

Do not describe AI ranking as eligibility. Ranking changes the order of review. Eligibility decides whether a study belongs in the review.

Flow Diagram: Track Records, Not Just Papers

The PRISMA flow diagram is where many weak review processes become visible. The PRISMA site says the flow diagram depicts information flow through review phases and maps records identified, included, excluded, and reasons for exclusions through the review process.

If AI affected search or screening, your record trail needs to survive. Do not lose excluded records just because an AI tool sorted them quickly. Do not merge search results, AI candidates, and manually added studies without tracking their origin.

Track:

  • Records identified through databases or registers.
  • Records identified through other methods.
  • Duplicates removed.
  • Records screened.
  • Reports sought for retrieval.
  • Reports assessed for eligibility.
  • Studies included.
  • Exclusions and reasons.

The exact labels depend on review type, but the principle is the same. If a reader cannot see how the final studies emerged from the candidate set, the review becomes harder to evaluate.

Data Extraction: AI Can Pre-Fill, Humans Must Verify

AI can help extract study characteristics, intervention details, outcomes, methods, limitations, or notes from PDFs. This can save time, especially when included studies have consistent structure.

Extraction is still high risk. A model can miss a subgroup, misread a table, flatten uncertainty, or confuse reported outcomes. If extracted data enters analysis, the authors need to check it against the original source.

Cochrane's Handbook chapter on collecting data notes evidence supporting duplicate data extraction and says one study observed that independent data extraction by two authors resulted in fewer errors than extraction by one author followed by verification by a second. It also notes a high prevalence of data extraction errors in prior reviews in the same discussion.

Report:

  • Which extraction fields AI assisted with.
  • Whether AI used PDFs, abstracts, tables, or uploaded text.
  • How extracted fields were checked.
  • Whether humans corrected AI output.
  • Whether uncertain fields were flagged.

For practical table design, use extracting data from research papers so each field is traceable to the source.

Synthesis: AI Can Suggest Structure, Not Decide Meaning

Synthesis is where the review's argument is built. AI can suggest themes, group studies, create draft tables, or summarize patterns. That can be useful. It can also be misleading if it turns complexity into smooth prose too early.

Use AI cautiously for:

  • Grouping studies by method, population, or outcome.
  • Suggesting possible evidence tables.
  • Identifying apparent similarities and differences.
  • Drafting questions for human review.
  • Checking whether a synthesis section has clear source support.

Do not let AI decide:

  • Which evidence is most trustworthy.
  • Whether findings are consistent enough to combine.
  • Whether a gap is real.
  • How limitations affect confidence.
  • What conclusion the review should draw.

Human judgment is not a decorative layer added at the end. It is the method.

If AI helped organize themes, disclose that role. If AI drafted synthesis prose, the manuscript may also need an AI-use statement. Use how to disclose AI use to a journal before submission.

Citation Verification Is Part Of The Method

AI-assisted reviews can accumulate citation risk. A tool may suggest papers, summaries, or references that look correct but need verification.

Before final submission:

  • Confirm each included paper exists.
  • Match title, authors, year, venue, and DOI.
  • Verify that cited claims are supported by the source.
  • Check retractions or major corrections for central studies.
  • Rebuild final references from trusted source records.

This is not only a writing issue. In a systematic review, citations define the evidence base. A fake, duplicated, or unsupported citation can distort the review.

The guide on how to verify AI-generated citations gives a step-by-step verification process for AI-assisted references.

What To Put In The Methods Section

A PRISMA-aware AI methods note should be specific enough for readers to understand the role of AI.

A useful methods paragraph can include:

AI-assisted tools were used during [stage]. The tool was used to [task]. Human reviewers applied the predefined eligibility criteria, checked AI-assisted outputs against source records or full texts, and made all final inclusion, extraction, and interpretation decisions. Disagreements or uncertain outputs were resolved by [process].

Adapt that template to the review. Do not make it broader than the actual use. If AI was used only for query expansion, say that. If AI ranked abstracts, say that. If AI pre-filled extraction fields, describe the verification step.

For manuscript-level AI disclosure, the methods note may not be enough. Some journals also require a separate AI declaration. The two statements should match.

WisPaper search results for AI-assisted paper screening

Where WisPaper Fits

WisPaper helps researchers search and screen academic papers with AI. Its search workspace supports Deep Search, Scholar Agent, and Inspiration Discovery, while paper cards show source labels, summaries, and preview images so users can triage results before deciding what to read.

WisPaper also lets users build a paper library and ask questions against that library. Papers can be uploaded or added from search results, then used as the basis for library-specific QA.

Try WisPaper

FAQs

AI can be used when it fits the protocol, is appropriate for the task, and is transparently reported. The authors remain responsible for final decisions, verification, and interpretation.