Can authors game AI reviewers without changing their science?
Yes—and this is the most policy-relevant failure mode. A 2026 study introduced 'adversarial repackaging,' where authors use AI-reviewer feedback to iteratively revise only presentation-level content (abstract, framing, related work, discussion) while keeping all scientific evidence fixed. Across three mainstream AI reviewers, this achieved a 75.1% attack success rate and a mean score gain of +1.21 out of 10 [2]. That means a paper could be nudged from borderline to acceptable simply by rephrasing how it is presented.
The same study found that AI reviewers are 'easier to impress than to convince': highlighting strengths reliably increased perceived merit, while attempts to dissolve weaknesses often backfired [2]. This suggests that AI reviewers may reward confident framing over substantive engagement with limitations—a serious concern for production use.
Even without attacks, do AI reviewers have blind spots?
Yes. A 2026 benchmarking study, PRISM, evaluated five leading automated reviewer systems against human reviewers on four dimensions: depth of analysis, novelty assessment, flaw identification, and constructiveness. While LLMs matched or beat humans on individual dimensions (e.g., stronger novelty verification), no single system matched the balanced performance of the human baseline across all dimensions at once [5]. Each system had a distinct specialization profile with characteristic blind spots—failure modes that aggregate metrics miss entirely [5].
This means that in production, relying on a single AI reviewer is risky; the evidence suggests they are best used as targeted supplements to human review, not standalone replacements [5].
About These Sources
This answer is built on 5 studies (all preprints) — published in 2026, 5 from 2024 or later — selected as the most relevant from 11 studies that passed quality screening, drawn from 54 papers retrieved from a database of over 500 million.
Sources used in this answer
ZFCρ Paper LXIII: Increment Second-Order Orthogonality and the SAE Axiom Intersection — From Deterministic Pairing to Zero-Mode Pinning / ZFCρ 第LXIII篇:余项增量二阶正交性与SAE公理的交汇——从确定性配对到零模Pinning
A mathematical study on integer complexity and zero-mode pinning, not directly about AI peer review, but it used multi-AI peer review (Claude, ChatGPT, Grok, Gemini) for its own experiments, illustrating the informal adoption of AI review.
No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
In a 2026 study, adversarial repackaging—presentation-only revisions—achieved a 75.1% attack success rate and a mean score gain of +1.21/10 across three AI reviewers, showing that AI reviewers are easier to impress than to convince.
When AI reviews science: Can we trust the referee?
A 2026 security analysis mapped attacks across the review lifecycle and demonstrated that hidden prompt injections can steer LLM reviews toward unjustifiably positive judgments, along with brittleness to adversarial phrasing and authority/length biases.
Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review
A 2026 benchmark, PaperGuard, showed that AI reviewers are pervasively vulnerable to cross-modal attacks targeting figures and text, and proposed a chunk-based embedding search defense to localize and mitigate harmful instructions.
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
A 2026 benchmarking study, PRISM, found that while LLM reviewers can match or beat humans on individual dimensions (e.g., novelty verification), no single system matches the balanced performance of human reviewers across all dimensions, indicating they are best as supplements, not replacements.
