Why do AI reviews score high on quality but fail to match human decisions?
The cleanest evidence comes from a 2025 study that compared AI-generated peer reviews to human reviews for 11 hand-surgery manuscripts [1]. When an experienced surgeon scored the reviews using a 20-item quality checklist (ARCADIA), ChatGPT's reviews averaged 4.8–4.9 out of 5, while human reviews averaged only 2.8–3.2. That sounds like AI is better—but the same study found that ChatGPT's acceptance decisions matched the higher-impact journal only 32% of the time (for GPT-4o) and 29% (for GPT-o1). In plain terms, the AI would have accepted nearly every paper that a top journal rejected, revealing a disconnect between how polished a review looks and whether it reaches the right editorial call.
This pattern is not isolated. A 2024 study of 21 articles found that human and AI review similarity was only moderate—around 3.6–3.8 out of 5 on a Likert scale—and the correlation between human and AI acceptance decisions was statistically significant for ChatGPT 3.5 but not for ChatGPT 4.0 [4]. That means even when AI reviews read well, they don't reliably align with human judgment on whether a paper should be published. So, quality studies that only measure the text's fluency or completeness can miss the fundamental issue: AI may be good at writing reviews but not at making the right call.
Can AI replace human reviewers?
The evidence across these studies points to a clear conclusion: AI can complement human peer review, but it cannot replace it. The 2024 study of 21 articles concluded that a fully automated AI review process is 'currently not advisable' and that ChatGPT's role should be 'highly constrained' [4]. The 2025 surgery study, despite showing high-quality AI reviews, still emphasized that managing AI's limitations rigorously is essential to improving publication quality [1].
The 2026 study on LLM-based auditing found that LLMs succeed in some tasks but that hybrid approaches—combining LLMs with interpretable machine learning and retrieval—are more reliable for verifying factual grounding [5]. This suggests that AI is best used as a tool to assist human reviewers, not as a standalone judge. The 2024 opinion piece also argues for using AI to augment the review process, not automate it entirely, given the pitfalls of bias and hallucination [2]. So, while quality studies might make AI look impressive, they fail to reveal that the safest path is human oversight with AI as a copilot, not the pilot.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2024 to 2026, 5 from 2024 or later, 2 in Q1 journals, collectively cited 85 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 47 papers retrieved from a database of over 500 million.
Sources used in this answer
Comparing AI-generated and human peer reviews: A study on 11 articles
In a study of 11 surgery manuscripts, ChatGPT's reviews scored higher on a 20-item quality checklist (4.8–4.9/5) than human reviews (2.8–3.2/5), but its acceptance decisions matched the higher-impact journal only 32% (GPT-4o) and 29% (GPT-o1) of the time, with an overall acceptance rate of 95–98%.
Peer Review in the Age of Generative AI
An opinion piece argues that AI can augment the peer review process but warns of pitfalls including biases, hallucinations, and ethical concerns, and calls for understanding how to use AI responsibly.
Artificial Intelligence as a Scientific Copilot in Analytical Chemistry: Transforming How We Write, Review, and Publish
A perspective in analytical chemistry highlights concerns about AI-assisted peer review, including authorship transparency, originality, and homogenization of scientific voice, and advocates for ethical guidelines and AI training.
Exploring the potential of ChatGPT in the peer review process: An observational study
In an observational study of 21 articles, human-AI review similarity was moderate (3.6–3.8/5), and the correlation between human and AI acceptance decisions was significant for ChatGPT 3.5 but not for ChatGPT 4.0, leading to the conclusion that fully automated AI review is not advisable.
Can LLMs Uphold Research Integrity? Evaluating the Role of LLMs in Peer Review Quality
A 2026 study on LLM-based review auditing found that LLMs can assess review quality but are not always reliable; hybrid approaches with interpretable machine learning and retrieval performed more reliably in verifying factual grounding.
