How the test worked
- I selected open-access papers whose retraction notices, found through Crossref's retraction records (which include the Retraction Watch data), give a specific reason that could in principle be seen in the manuscript. Fabrication and misconduct cases were left out. Each paper carries a licence that allows reuse.
- I took the full text of each paper from Europe PMC's open XML service and typeset it. Its PDF service sits behind a bot check, which I did not bypass. The typeset versions have no figures, and some tables and equations were lost in conversion.
- Each paper went through the production review (Standard tier, general field) with no hint that it was retracted.
- A separate model read only the notice's reason and the review's findings, not the paper, and answered yes, partly or no. I read every yes, partly and no against the report.
Result
Of 23 papers, the review pointed at the documented problem in 14 and partly in 4. It did not in 5. I count 11 or 12 of the 14 as clear matches. Two are generous: one review matched only half of its notice, and one notice is too vague to confirm. Where the notice itself said the error was visible in the manuscript, the review pointed at it in 9 of 11 papers. Where that was uncertain, it did in 5 of 12.
Some of the clear matches: too few events per variable for the number of predictors, a median below the minimum in a table, electron-volt values converted to kcal/mol with the wrong factor, odds ratios that were the reciprocal of what the counts imply, and a confidence interval that crossed zero beside a claim of significance.
What it missed
Two of the five misses are partly the test's fault: my conversion dropped a table in one paper and rendered another paper's equations unreadable (the review said so). Two notices describe clerical or record-handling mistakes, such as data from the wrong wave merged into a file, which the published text does not show. One is a plain miss: a paper excluded pauses and unvoiced segments from an acoustic analysis, and the review did not question it.
Is it just flagging everything?
A fair worry, because a report lists about seven high-confidence concerns on almost any paper. I ran six papers that were not retracted. Their reports averaged 6.5 such concerns, against 6.4 for the retracted set, so the count tells you nothing about whether a paper has a problem. The content does. I gave the same judge each retraction reason together with a report written for an unrelated paper: it said yes in 2 of 46 cases, against 14 of 23 for the matching reports. Both of those two were real overlaps; one unrelated paper really does have too few events per variable.
I did not check whether the concerns raised on the non-retracted papers are correct. They are real published papers with real flaws, and six is too few to give a false-alarm rate.
Limits
- The 23 papers were chosen for retractions with a specific, checkable cause, about half from PLOS ONE. Subtle flaws and misconduct are not represented, and those are harder to find.
- One model scored the results, and each paper was run once, so the percentages are not tight.
- Figures were absent, so the figure scan was not exercised.
- The review reads only what the paper shows. It cannot catch a mistake in data the paper does not include.
- This says nothing about matching a human reviewer, and I do not claim it does.
Cost
Model fees were about $3.90 per paper, up to $5.40 for papers of up to 58 pages. The Standard review costs $9. The whole test, including the non-retracted papers and the baseline, cost about $113.
The numbers in more detail
The papers were published between 2015 and 2024. Eleven are from PLOS ONE, three from Cureus, and the other nine from nine different journals. The retractions fall into five causes, taken from the notices. "Duplicated or inconsistent data" means numbers in the paper that contradict each other or cannot be real. "Data error" means a mistake in the data file itself, such as the wrong records merged.
| Cause | Papers | Yes | Partly | No |
|---|---|---|---|---|
| Duplicated or inconsistent data | 5 | 5 | 0 | 0 |
| Statistical error | 5 | 2 | 2 | 1 |
| Calculation error | 5 | 3 | 1 | 1 |
| Analysis error | 4 | 3 | 0 | 1 |
| Data error | 4 | 1 | 1 | 2 |
The pattern is what you would expect from a tool that reads the paper and not the data. Contradictions that show on the page were found in all five cases. Mistakes that live in a data file mostly were not.
Before running anything, I rated each retraction on whether its error could be seen from the manuscript alone. I rated 11 as visible and 12 as uncertain.
| Expected visibility | Papers | Yes | Partly | No |
|---|---|---|---|---|
| Visible in the manuscript | 11 | 9 | 1 | 1 |
| Uncertain | 12 | 5 | 3 | 4 |
The list of papers is available on request. Write to ben@purplelink.llc and I will send the DOIs, the retraction reasons and the verdicts. I left the per-paper list off the page because each row is a named author's retracted paper.
If you want to see what a report looks like, there is a sample report. To run one on your own manuscript, go to Paper Review. The free Retraction Checker checks your reference list for retracted papers with no upload.