What can peer-review quality studies fail to reveal about AI-assisted peer review?

AI peer review studies show high-quality text but poor agreement with human decisions, risking false acceptance and hidden biases.

Direct answer

Peer-review quality studies can reveal that AI-generated reviews often read as polished and thorough—scoring higher on quality checklists than human reviews—but they can fail to expose that AI decisions frequently disagree with human editorial judgments. For example, in a 2025 study of 11 surgery manuscripts, ChatGPT's acceptance rate was 95–98%, yet its decisions matched the higher-impact journal only 32% of the time, meaning it would have accepted nearly all papers that a top journal rejected [1]. Across the studies, AI reviews showed only moderate similarity to human reviews (around 3.6–3.8 out of 5) [4], and AI can hallucinate or introduce biases that quality metrics don't catch [2][3]. So, while AI can assist, these studies fail to reveal the real-world risk of over-acceptance and the need for human oversight.

5sources cited

This article was generated with WisPaper-powered search and paper analysis.

Why do AI reviews score high on quality but fail to match human decisions?

The cleanest evidence comes from a 2025 study that compared AI-generated peer reviews to human reviews for 11 hand-surgery manuscripts [1]. When an experienced surgeon scored the reviews using a 20-item quality checklist (ARCADIA), ChatGPT's reviews averaged 4.8–4.9 out of 5, while human reviews averaged only 2.8–3.2. That sounds like AI is better—but the same study found that ChatGPT's acceptance decisions matched the higher-impact journal only 32% of the time (for GPT-4o) and 29% (for GPT-o1). In plain terms, the AI would have accepted nearly every paper that a top journal rejected, revealing a disconnect between how polished a review looks and whether it reaches the right editorial call.

This pattern is not isolated. A 2024 study of 21 articles found that human and AI review similarity was only moderate—around 3.6–3.8 out of 5 on a Likert scale—and the correlation between human and AI acceptance decisions was statistically significant for ChatGPT 3.5 but not for ChatGPT 4.0 [4]. That means even when AI reviews read well, they don't reliably align with human judgment on whether a paper should be published. So, quality studies that only measure the text's fluency or completeness can miss the fundamental issue: AI may be good at writing reviews but not at making the right call.

What hidden risks do quality studies overlook?

Quality studies often focus on surface-level attributes like specificity and tone, but they can fail to reveal deeper problems such as AI hallucinations—making up claims or references—and inherent biases. A 2024 opinion piece in the Journal of the Association for Information Systems explicitly warns that AI tools introduce biases and can hallucinate, issues that standard quality metrics may not capture [2]. Similarly, a 2025 perspective in Analytical Chemistry highlights concerns about originality and the homogenization of scientific voice, noting that AI-assisted reviews might lack the critical thinking and creativity that define genuine scientific authorship [3].

The 2025 surgery study itself acknowledged that ChatGPT's high acceptance rate (95–98%) was partly due to its tendency to accept papers, and the authors stressed that precise instructions are needed to avoid hallucinations [1]. This suggests that quality scores can be inflated by AI's overly positive or generic responses, masking the risk that it might rubber-stamp flawed research. A 2026 study on LLM-based review auditing found that while LLMs can assess review quality, they are not always reliable—hybrid approaches with interpretable machine learning and retrieval performed more reliably in verifying whether reviewer claims are supported by the paper [5]. So, quality studies that don't test for factual grounding or bias can give a false sense of AI's readiness.

Can AI replace human reviewers?

The evidence across these studies points to a clear conclusion: AI can complement human peer review, but it cannot replace it. The 2024 study of 21 articles concluded that a fully automated AI review process is 'currently not advisable' and that ChatGPT's role should be 'highly constrained' [4]. The 2025 surgery study, despite showing high-quality AI reviews, still emphasized that managing AI's limitations rigorously is essential to improving publication quality [1].

The 2026 study on LLM-based auditing found that LLMs succeed in some tasks but that hybrid approaches—combining LLMs with interpretable machine learning and retrieval—are more reliable for verifying factual grounding [5]. This suggests that AI is best used as a tool to assist human reviewers, not as a standalone judge. The 2024 opinion piece also argues for using AI to augment the review process, not automate it entirely, given the pitfalls of bias and hallucination [2]. So, while quality studies might make AI look impressive, they fail to reveal that the safest path is human oversight with AI as a copilot, not the pilot.

About These Sources

This answer is built on 5 peer-reviewed studies — published from 2024 to 2026, 5 from 2024 or later, 2 in Q1 journals, collectively cited 85 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 47 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Comparing AI-generated and human peer reviews: A study on 11 articles

In a study of 11 surgery manuscripts, ChatGPT's reviews scored higher on a 20-item quality checklist (4.8–4.9/5) than human reviews (2.8–3.2/5), but its acceptance decisions matched the higher-impact journal only 32% (GPT-4o) and 29% (GPT-o1) of the time, with an overall acceptance rate of 95–98%.

2

Peer Review in the Age of Generative AI

An opinion piece argues that AI can augment the peer review process but warns of pitfalls including biases, hallucinations, and ethical concerns, and calls for understanding how to use AI responsibly.

3

Artificial Intelligence as a Scientific Copilot in Analytical Chemistry: Transforming How We Write, Review, and Publish

A perspective in analytical chemistry highlights concerns about AI-assisted peer review, including authorship transparency, originality, and homogenization of scientific voice, and advocates for ethical guidelines and AI training.

4

Exploring the potential of ChatGPT in the peer review process: An observational study

In an observational study of 21 articles, human-AI review similarity was moderate (3.6–3.8/5), and the correlation between human and AI acceptance decisions was significant for ChatGPT 3.5 but not for ChatGPT 4.0, leading to the conclusion that fully automated AI review is not advisable.

5

Can LLMs Uphold Research Integrity? Evaluating the Role of LLMs in Peer Review Quality

A 2026 study on LLM-based review auditing found that LLMs can assess review quality but are not always reliable; hybrid approaches with interpretable machine learning and retrieval performed more reliably in verifying factual grounding.