Who Audits the Reviewers? A Multi-Model Consensus Framework for Characterizing Failures in LLM-Assisted Peer Review
Olalekan Joseph Akintande
Abstract
Large Language Models (LLMs) are increasingly used to assist with academic peer review, yet their outputs remain prone to systematic errors often presented with high confidence. This paper presents a comprehensive empirical characterisation of failure modes in LLM-assisted academic review through a controlled experiment involving 15 papers, 5 LLMs, 3 prompt conditions, and 672 independent judge-review evaluations across 11 error types. We find that Unsupported Claims is the most prevalent error (mean frequency 3.02 per review), followed by Plausible Reasoning Gaps (2.86) and Confirmation Bias (2.19). Llama 3.1 70B consistently outperforms commercial alternatives, including GPT-4o (overall error mean 3.23 vs. 3.53), challenging assumptions regarding the superiority of proprietary architectures. Structured prompting reduces error frequency by approximately 35\%, while self-correction reduces errors by 44\% on average; however, self-correction efficacy is highly architecture-dependent, ranging from a 50.73\% reduction in GPT-4o to negligible improvements in Gemini 3.5 Flash. Inter-judge agreement is poor (mean Fleiss' Kappa: -0.013), with a 2.6-fold stringency differential between judges, mirroring human peer review variability. A counterintuitive "Self-Correction Paradox" emerges wherein higher-quality reviews generate both increased consensus and specific points of heightened disagreement. Logistic regression confirms that model choice, prompt condition, and domain, particularly in Medicine, are significant predictors of error prevalence. Novice reviewers miss 2.5\% of errors on average. We propose a robust failure taxonomy and offer practical recommendations for journals, emphasising that while LLMs offer promising structural support, rigorous human oversight remains indispensable.