How We Catch Partially Duplicated Tables
Copy-pasted rows and columns are one of the clearest signatures of a fabricated or mill-produced table. ReviewerZero checks every table against the rest of the paper and against a background corpus of more than 17 million tables from open-access publications.
The (synthetic) table above looks ordinary. It reports six secondary outcomes from a clinical trial, each with a treatment effect, a comparison, and a P value. Look at the two highlighted columns, though, and the values for Drug A and Drug B are identical on every row, down to the last decimal. Two independent treatment groups do not produce the same mean and standard deviation by chance. One column was copied from the other.
We check every table against the rest of the paper and against the literature
ReviewerZero screens the tables in a submission automatically, in two directions at once.
First, against the rest of the paper. Every row, column, and block of cells is compared with the others, so a column duplicated across treatment arms (like the one above) or a row reused between two tables is surfaced.
Second, against the literature. We maintain a background corpus of more than 17 million tables drawn from open-access publications, and we are expanding it quickly. A submitted table whose values match one already in the published record is a different but equally important signal. It can point to self-plagiarism, a recycled cohort, or a table lifted from another paper.
Small variations still count
Copy-paste is not always pixel-perfect. Someone might round a cell to one fewer decimal place, swap in a different minus sign, or add a significance star to a single row. We normalize numbers before comparing them, so these superficial edits do not hide a duplicated block. Values like 1.23 and 1.230 are treated as the same number, as are comma-separated thousands, alternate dash characters, and other common formatting differences.
What we found
A limited internal large-scale auditing revealed that most of the raw duplicates are standard datasets in a field (e.g., income distributions, demographics, etc.) and machine learning benchmarks. That's why we have multiple layers of checks to determine whether a duplication is legitimately problematic and you only get to see the truly suspicious ones.
What you see
When the system finds a match, it does not just say "duplicate". It shows the two matched regions side by side, with their row labels, column headers, and the exact cells that line up.
Each finding carries a severity and a short, plain explanation of why it was flagged, so an editor can act in seconds instead of re-deriving the match by hand.
Try it
Book a demo or join the beta to run table duplication, image duplication, statistical, and reference checks on your own papers.
