Paper Review · Machine learning
A pre-submission review for machine-learning papers.
Choose the machine-learning profile at upload and the four reviewers go looking for the problems ML referees raise most: contamination, unfair baselines and results that sit inside the noise.
The machine-learning profile
What the reviewers look for
These are the instructions the profile adds to the review, in plain words. Each finding in the report quotes the passage it is about and says what would fix it.
- Contamination. Test data seen during training or model selection, and pre-training contamination on common benchmarks, which matters most for LLM evaluations.
- Seeds and error bars. Single-run results, no error bars, and claims resting on three or five runs with no significance test.
- Fair baselines. More compute, data or tuning for the proposed method than for the baselines, including a hyperparameter search budget the baselines never had.
- Ablations. Architectural choices with no ablation showing they matter.
- Bold-the-winner tables. Tables that bold the best number when the differences are within noise.
- Compute claims. Cost comparisons that leave out the hyperparameter search.
- Artifacts. Code and data that are "available upon request" rather than released.
- LLM-as-judge. Evaluation by a judge from the same model family as the system being tested.
What a finding looks like
An example from the report
Illustrative, not from a real manuscript: this is the shape each finding takes. The full sample report shows every section.
Baselines tuned less than the proposed method.
Quoted: "We tuned learning rate and batch size on the validation set; baselines used reported defaults."
Surfaced by: Methodology Critic, Statistical Skeptic (consensus).
Fix: Give each baseline the same search budget, or report results with defaults for every method, and state the protocol in the methods section.
In every review
The rest of the report
- Four reviewers: a Methodology Critic, a Statistical Skeptic, a Data Integrity Officer and an Editor-in-Chief, with findings more than one of them raise listed first.
- Every reference checked against CrossRef, and for the most important claims, the cited paper's abstract read to see whether it supports the sentence citing it.
- A figure and table scan for integrity and presentation problems.
- An anonymity scan for double-blind submission.
- A prioritized checklist, as a Markdown report and an annotated PDF with comments on the pages they refer to.
- The Journal Pack ($11) adds a check against one venue's submission rules. In machine learning it covers NeurIPS, ICML, ICLR, AAAI, TMLR and JMLR. All 25 venues.
Questions
Does it check anonymity for a double-blind ML venue?
How do I choose this profile?
What happens to the manuscript?
How long does it take and what does it cost?
Read it before Reviewer 2 does.
From $9, usually back in 4 to 8 minutes. No account, no subscription.
Other profiles: BiomedicinePsychology and social scienceChemistry and materials