AI Review Update: Up to 5 Times More Issues Detected Than Before
We rebuilt ReviewerZero's AI review around a much deeper, broader and more consistent reading of the manuscript, with evidence on every finding and calibrated severity. On manuscripts where we planted the errors ourselves, it finds up to five times as many of them as the previous version did.
A year ago, we described an AI review that read the whole manuscript, looked for contradictions between sections, and returned a structured report of major and minor issues tied to the passages they concerned.
This summer we rebuilt it to go further, drawing on dozens of experiences shared by real users and on new evaluation sets and methods of our own. The result is the deepest and most precise reviewer we have built.
The question we set out from is what a careful referee actually does: hold the abstract against the results, the methods against the tables, and the sample-size plan against the number of participants who appear, and notice when the pieces do not fit. The rebuilt review is built to read that way, and to show its work.
How much better
We test the review on manuscripts with errors introduced by hand, across fields and formats, spanning several categories of issues.
The rebuilt reviewer finds up to five times as many of the errors as the previous version, and every one of those is a true finding: an automated judge matches each finding to a planted error we know is there, and only a match counts. It now catches nearly all of the errors that show up as a number contradicting the text around it, most of those that take methodological judgment to see, and a growing share of those that are omissions rather than mistakes.
Finding distant contradictions and issues
One of the most important new capabilities is finding long-range contradictions and inconsistencies. Many of them are statements in the text that do not match the tables, equations or figures.
Neither sentence is wrong on its own. Together they say the replication was presented as covariate-adjusted and analyzed unadjusted. The review found it by holding the abstract against the degrees of freedom in the results, and it cites all three places an editor would need to see.
The review did the subtraction the authors did not, noticed that completion also varies across outcomes within the same group, and asked for the participant flow, the reasons, and a sensitivity analysis.
Built on what we learned
This is not our first reviewer. At the 2025 Peer Review Congress our team presented an AI peer review panel that assessed manuscripts for clarity, novelty and impact. In head-to-head comparisons judged by readers, its reviews scored above the human reviews of the same manuscripts, and they were more detailed about limitations and possible improvements. One lesson from that work is that a reviewer's praise is far less useful to an editor than its findings, so the rebuilt review keeps the depth and keeps its issue list to problems only.
Interested in trustworthy science? Book a demo to see the AI review on your own manuscripts.
