Claude and I planted 100 known errors into 10 open-access psychology papers and then ran them through frontier models and two commercial AI review tools. In brief:
- The best single system caught 71 of 100 errors, while the worst caught 30.
- Pooling every system’s output caught 93 of 100. Models are only partly correlated in the errors they find, making ensembling a big lever for finding issues in papers. Check your papers against multiple models!
- Seven errors could not be caught by any system. All were omissions — information deleted from a paper rather than mistakes inserted into it.
- Refine.ink contributes more unique catches than any other single system, though it’s expensive.
- I didn’t measure false positives and I don’t know how this error distribution compares to the distribution of errors in real papers.
- I’ve made the papers, errors, model outputs, and the full experiment log public. I hope people can build on this work to create a comprehensive eval benchmark across disciplines.
That is from Paul Litvak, here is more. Note that is not even using the very latest generation of models.





