Each section gives a one-sentence definition, why reviewers care, a sentence of the kind that draws the comment, and the fix. The example sentences are invented but typical. Primary literature is listed in the references, and each section links to a shorter explainer you can send to a co-author.
None of these problems means a paper is wrong. They mean it has not yet shown that it is right. Most can be fixed with an extra table, a clearer methods paragraph, or narrower wording in the abstract.
Statistics and inference
The first four problems share a root. A p-value means what it claims only if the analysis was fixed before the data could influence it. Each is a different way the data ends up choosing the analysis.
p-hacking
p-hacking is running several analyses of the same data and reporting the ones that reach statistical significance while leaving the others out.
Why reviewers care. A 5% significance threshold assumes one planned test. Simmons, Nelson and Simonsohn (2011) showed by simulation that a handful of ordinary, individually defensible choices (picking between two correlated outcome measures, adding participants after an early look, controlling for a covariate, dropping one condition) could together push the false-positive rate from 5% to more than 60%. Head et al. (2015) later found distributions of reported p-values across many fields that are consistent with p-hacking. Reviewers look for the traces: several p-values just under .05, measures in the methods that never reach the results, and covariates that appear in one model but not the next.
How it shows up. A results paragraph such as: Participants in the treatment condition reported higher satisfaction (p = .047); no significant differences emerged on the remaining measures, which are not discussed further.
"Not discussed further" tells the reviewer other tests were run and set aside.
How to fix it. Decide the primary outcome and analysis before looking at the data, and write the decision down somewhere with a timestamp, such as a preregistration (Nosek et al., 2018). Report every measure, every condition and every exclusion, and if you excluded observations or added a covariate, also report the result without that choice, as Simmons and colleagues recommend. State how the sample size was decided, and label any unplanned analysis as exploratory.
Short explainer: p-hacking
The garden of forking paths
The garden of forking paths is the set of reasonable, data-dependent analysis choices a researcher could have made, which inflates false positives even when only one analysis was actually run.
Why reviewers care. Gelman and Loken (2013) introduced the term to describe a problem that needs no fishing at all. A researcher runs one analysis in good faith, but the exclusion rule, the transformation, the subgroup and the covariate set were each chosen after seeing the data. Had the data come out differently, a different set of choices would have looked just as reasonable, so the single reported test carries the same inflated error rate as an explicit search (Gelman & Loken, 2014).
How it shows up. A methods sentence such as: Because reaction times were positively skewed, we log-transformed them and excluded trials more than 2.5 standard deviations from each participant's mean.
Every clause is standard practice, but "because" signals that the choices followed the data.
How to fix it. Preregister the exclusion rules, transformations and covariates, or state that they follow an established lab or field convention and cite it. Where choices were made after seeing the data, show that they do not drive the result. A multiverse analysis reports the estimate under every reasonable combination of choices (Steegen et al., 2016), and a specification curve plots those estimates in order so the reader can see how many support the claim (Simonsohn, Simmons & Nelson, 2020).
Short explainer: garden of forking paths
HARKing
HARKing (hypothesizing after the results are known) is presenting a hypothesis that was formed after seeing the data as if it had been stated before the study began.
Why reviewers care. Kerr (1998) named the practice and set out its costs. A hypothesis suggested by a pattern in the data cannot fail a test on that same data, so the "confirmation" carries no evidential weight. It also hides the original hypothesis, which may have failed. Reviewers become suspicious when every prediction is confirmed, or when predictions are oddly specific, such as an effect in exactly one subgroup.
How it shows up. A results sentence such as: As predicted, the effect emerged only among participants high in need for cognition.
If nothing earlier in the paper predicts that moderation, the reviewer will assume the prediction came after the finding.
How to fix it. Keep a dated record of your hypotheses before analysis; a preregistration is the most credible version. In the paper, separate confirmatory from exploratory results under their own headings. An unexpected finding is worth reporting, as long as it is reported as unexpected and framed as a hypothesis for future work. Some journals accept Registered Reports, in which the introduction and methods are reviewed and accepted in principle before data collection (Chambers, 2013).
Short explainer: HARKing
Multiple comparisons
The multiple comparisons problem is that the chance of at least one false positive rises with every additional test run at the same significance threshold.
Why reviewers care. At α = .05, each test of a true null hypothesis has a 5% chance of a false positive. Across 20 independent tests, the chance of at least one is 1 − 0.9520, about 64%, and the expected number of false positives is one. Reviewers count the tests, including those implied by tables, and compare that number to the number of starred cells.
How it shows up. A sentence such as: We correlated each of the 12 personality facets with each of the 8 outcome measures; significant correlations are marked with an asterisk in Table 3.
That is 96 tests; about five would reach p < .05 by chance.
How to fix it. Say how many tests you ran and apply a correction that fits your goal. To control the chance of any false positive (the familywise error rate), use Bonferroni or Holm's step-down procedure, which never rejects fewer hypotheses than Bonferroni (Holm, 1979). When you are screening many hypotheses and can tolerate a known proportion of false discoveries, control the false discovery rate with the Benjamini-Hochberg procedure (Benjamini & Hochberg, 1995). Better still, designate a few primary comparisons in advance, so the correction applies to a short list.
Short explainer: multiple comparisons
Samples and generalization
One of these limits what a study can detect; the other limits whom the findings apply to.
Underpowered samples
An underpowered sample is one too small to reliably detect an effect of the size the authors consider meaningful.
Why reviewers care. Button et al. (2013) estimated the median statistical power of neuroscience studies at about 21% and explained why this matters even for positive results: when power is low, a significant finding is less likely to reflect a true effect, and the effect sizes that do reach significance tend to be overestimates. This follows from Ioannidis (2005), where the probability that a significant result is true depends on power, prior odds and bias. A reviewer who sees a small sample and a large effect reads the effect as a warning.
How it shows up. A pairing such as: The sample size was based on previous studies in this area.
followed later by With 18 participants per group, the difference was significant (p = .03, d = 0.74), indicating a large effect of the intervention.
The first is not a power analysis; the second is the pattern low power produces.
How to fix it. Justify the sample size before collecting data. A power analysis should state the smallest effect size you care about, where that number comes from, the planned test, alpha and target power, and the resulting N. Lakens (2022) covers this and other defensible justifications. If the data are already collected, report a sensitivity analysis: the smallest effect your design could detect with 80% power. Avoid "observed power" computed from your own effect estimate; it is a transformation of the p-value and adds no information (Hoenig & Heisey, 2001).
Short explainer: underpowered sample
WEIRD samples
A WEIRD sample is drawn from Western, Educated, Industrialized, Rich and Democratic societies; the problem arises when findings from such a sample are presented as true of people in general.
Why reviewers care. Henrich, Heine and Norenzayan (2010) reviewed comparative evidence across domains including visual perception, fairness, cooperation, spatial reasoning and self-concept, and found that WEIRD participants were often outliers rather than typical humans. Susceptibility to the Müller-Lyer illusion, for example, varies widely across societies. Arnett (2008) found that the great majority of participants in leading psychology journals came from the United States and a few other Western countries. A reviewer will check whether the abstract's claim is about the population sampled or about "people".
How it shows up. An abstract sentence such as: These results show that people rely on intuition rather than deliberation when making moral judgments.
where the methods describe 240 undergraduates at one US university.
How to fix it. Report the sample's demographics in enough detail to judge its range. Match the wording of the conclusion to the population studied. Add a short statement of the constraints on generality, naming the populations, settings and materials to which you expect the result to extend and those to which you do not (Simons, Shoda & Lindsay, 2017). Where the claim needs to be universal, it needs data from more than one kind of society.
Short explainer: WEIRD samples
Machine learning evaluation
Machine learning papers usually make a comparative claim: method A beats method B on benchmark C. Each of the next five problems weakens one part of that comparison.
Test-train contamination
Test-train contamination occurs when examples from the evaluation set, or close copies of them, are present in the data a model was trained or tuned on, so the test measures memory rather than generalization.
Why reviewers care. Kapoor and Narayanan (2023) documented data leakage across many scientific fields that use machine learning, including preprocessing fitted on the full dataset, duplicates across splits, and splits that ignore dependence between samples. For large language models the risk is larger, because pretraining corpora are huge and benchmarks are public. Brown et al. (2020) ran an n-gram overlap analysis between their pretraining data and each benchmark and reported the results; reviewers now expect something similar.
How it shows up. A methods sentence such as: We randomly split the 12,000 images into training (80%) and test (20%) sets.
With several images per patient, the same patients land in both sets, and the model can score well by recognizing patients rather than disease.
How to fix it. Split by the unit you want to generalize to (patient, user, document, site or time period), not by row. Fit normalization, feature selection and any other preprocessing inside the training split only. Deduplicate across splits with a near-duplicate check, not just exact matching. For pretrained models, prefer benchmarks released after the training data cutoff, report an overlap check against whatever training data you can inspect, and compare scores on items flagged as possibly contaminated with scores on the rest.
Short explainer: test-train contamination
Cherry-picked seeds
Cherry-picked seeds describes reporting a single training run, or the best of several, so that random run-to-run variation is presented as a difference between methods.
Why reviewers care. Random initialization, data order, data splits and nondeterministic hardware all move the final score. Henderson et al. (2018) showed in deep reinforcement learning that two groups of runs of the same algorithm, differing only in random seed, could produce results that looked like two different algorithms. Bouthillier et al. (2021) catalogued the sources of variance in machine learning benchmarks and recommended randomizing as many of them as possible across runs. If the gap between methods is smaller than the spread across seeds, a single number per method does not support a ranking.
How it shows up. A results sentence such as: Our method achieves 84.7% accuracy, a 0.6-point improvement over the strongest baseline (84.1%).
with no mention of how many runs, no error bars, and the best number in each column in bold.
How to fix it. Run several seeds for every method, including baselines, and say how many. Report the mean with a standard deviation or confidence interval, and report all runs rather than the best. For comparisons across many tasks or with few runs, Agarwal et al. (2021) recommend interval estimates such as stratified bootstrap confidence intervals and robust aggregates such as the interquartile mean. If a single run is all you can afford, say so plainly and soften the claim.
Short explainer: cherry-picked seeds
Hyperparameter asymmetry
Hyperparameter asymmetry is tuning the proposed method more thoroughly than the baselines it is compared against, so part of the reported gain comes from the extra tuning.
Why reviewers care. Re-evaluations have found that reported gains shrink or disappear once baselines get the same tuning effort. Melis, Dyer and Blunsom (2018) found that standard LSTMs, tuned with large-scale automatic search, outperformed more recent architectures on language modeling benchmarks. Lucic et al. (2018) found that with enough hyperparameter search and random restarts, most GAN variants reached similar scores, and no tested variant consistently beat the original. Dodge et al. (2019) showed that which model looks best can depend on the tuning budget, and proposed reporting expected validation performance as a function of that budget.
How it shows up. A methods sentence such as: Baselines were trained with the hyperparameters reported in their original papers; for our method, we selected the learning rate and dropout rate by grid search on the validation set.
How to fix it. Give every method the same search procedure and the same number of trials, over ranges that are reasonable for each. Use the same preprocessing, augmentation, training length and early-stopping rule. Report the search space, the budget and the selected values for every method in an appendix. If you copy baseline numbers from earlier papers, say so and explain why the setups are comparable.
Short explainer: hyperparameter asymmetry
Missing ablations
A missing ablation is a proposed component whose individual contribution to the result was never tested by removing it and measuring the change.
Why reviewers care. If a paper introduces several changes at once, the combined result cannot show which one mattered. Lipton and Steinhardt (2019) listed "failure to identify the sources of empirical gains" among the troubling trends in machine learning scholarship, noting cases where gains attributed to a new architecture actually came from tuning or from a single small change.
How it shows up. A sentence such as: Combining the gated attention module, the curriculum schedule and the auxiliary contrastive loss, our model improves F1 by 3.2 points over the baseline.
with no table showing what each of the three adds.
How to fix it. Add an ablation table. Each row removes one component, or adds one to the baseline, with everything else held fixed, including the tuning budget. Run ablation rows with the same number of seeds as the main results; otherwise small differences between rows are noise. Include the row a skeptic would ask for, usually the baseline plus only the simplest change. If a component does not help, drop it or say so.
Short explainer: missing ablations
LLM self-judge bias
LLM self-judge bias is the tendency of a language model used as an evaluator to favor outputs from itself or its own model family, which inflates scores when the judge and the system under test are related.
Why reviewers care. Zheng et al. (2023) found that GPT-4's judgments agreed with human preferences at a rate similar to agreement between humans. They also documented position bias (favoring the first or second answer), verbosity bias (favoring longer answers) and self-enhancement bias (favoring its own answers), while noting that their evidence on the last was limited. Panickssery, Bowman and Feng (2024) found that models can recognize their own outputs and that this ability is linked to preferring them. When the judge is related to a system being compared, the reviewer cannot tell how much of the gap is the judge.
How it shows up. A methods sentence such as: Responses from our system and three baselines were rated for helpfulness on a 1 to 10 scale by the same model our system was fine-tuned from.
How to fix it. Use a judge from a different model family than any system under test, or several judges,. Collect human ratings on a random subset and report how well the judge agrees with them. In pairwise comparisons, present each pair in both orders. Report response length. Publish the judge prompt and the exact model version so the evaluation can be rerun.
Short explainer: LLM self-judge bias
Reporting, citations and figures
These problems are not about the design of the study, but reviewers treat them as evidence of how carefully the rest was done.
Hallucinated citations
A hallucinated citation is a reference that looks complete and plausible but does not correspond to a real publication, usually because it was produced by a generative AI tool.
Why reviewers care. Language models generate references by predicting plausible text, not by looking papers up. Walters and Wilder (2023) asked ChatGPT to write short literature reviews and checked every citation: 55% of the GPT-3.5 citations and 18% of the GPT-4 citations were fabricated, and many of the real ones contained errors. Some fabrications pair real authors with a title they never wrote, or attach a real DOI to the wrong paper. A reviewer who cannot find a key reference will wonder what else was not checked.
How it shows up. A sentence such as: Prior work has established that trust calibration improves human-AI team accuracy (Nguyen & Park, 2021).
where the reference list gives a complete journal citation with volume, pages and DOI, and no such paper exists. (The citation here is invented to illustrate the pattern.)
How to fix it. Check every reference against a bibliographic database before submission: resolve each DOI and confirm that it lands on the paper you meant, and search each title in Crossref or Google Scholar. For the references that carry your main claims, open the source and confirm that it supports the sentence you attached it to. Remove anything you cannot find.
Short explainer: hallucinated citation
Figure presentation
A figure presentation problem is a design choice in a chart, such as a truncated axis, missing or unlabeled error bars, or color-only encoding, that changes what a reader concludes from the data.
Why reviewers care. Weissgerber et al. (2015) showed that bar graphs of continuous data hide the distribution: very different datasets, including ones with outliers or bimodal groups, can produce identical bars. Crameri, Shephard and Heron (2020) showed that rainbow color maps distort data by creating false boundaries, and that they are hard to read for people with color-vision deficiency.
How it shows up. A text sentence such as: As Figure 3 shows, accuracy improved substantially in the treatment group.
where Figure 3 is a bar chart with a y-axis running from 71% to 74%, six observations per group, and a caption reading only "Error bars show SE."
How to fix it. For small samples, plot individual observations, alone or over a box or violin plot. Start bar charts at zero; for small differences, use a dot plot instead of cropped bars. In every caption, state what the error bars represent (standard deviation, standard error or a confidence interval) and the n behind them. Use perceptually uniform color maps, and pair color with shape, line style or direct labels. Check that text is readable at the journal's column width.
Short explainer: figure presentation
Study design and registration
The last two problems apply mostly to experiments with human or animal participants. In both cases a safeguard may well have been in place, but the manuscript does not give the reader enough to check it.
Blinding not evidenced
Blinding not evidenced means a manuscript says a study was blinded without describing who was blinded, to what, and how the blinding was maintained.
Why reviewers care. Blinding protects against expectation effects in participants, in those delivering an intervention, and in those measuring outcomes. The word alone does not say which were covered. Schulz and Grimes (2002) pointed out that labels such as "double-blind" are used inconsistently and argued that authors should state exactly who was blinded. The CONSORT statement for randomized trials asks for who was blinded after assignment and how (Schulz, Altman & Moher, 2010), and the ARRIVE 2.0 guidelines for animal research include blinding among their ten essential items (Percie du Sert et al., 2020).
How it shows up. A single methods sentence such as: All behavioral scoring was performed blind to genotype.
with no detail about who held the key or when it was opened.
How to fix it. Name each group that was blinded: participants, experimenters, outcome assessors, data analysts. Say how allocation was concealed (who generated the codes, who held them, how samples or cages were labeled) and when the blind was broken relative to the analysis. If you tested whether blinding held, report the result. If blinding was not possible, say so, explain why, and describe what you did instead, such as using an objective outcome or an independent assessor.
Short explainer: blinding not evidenced
Trial not registered
Trial not registered describes a clinical trial reported without a public registry entry made before the first participant enrolled, which leaves readers unable to check whether the reported outcomes were the planned ones.
Why reviewers care. In 2004 the International Committee of Medical Journal Editors announced that its member journals would consider a clinical trial only if it had been registered in a public registry, and for trials beginning enrollment after 1 July 2005, registered at or before the start of enrollment (De Angelis et al., 2004). The current ICMJE Recommendations keep the requirement and define a clinical trial broadly enough to include many behavioral interventions: any study that prospectively assigns people to an intervention to study its effect on a health outcome (ICMJE). Chan et al. (2004) compared trial protocols with the resulting publications and found that 62% of trials had at least one primary outcome that was changed, introduced or omitted.
How it shows up. A methods section that begins: Participants were randomly assigned to the intervention or to usual care, and the primary outcome was change in HbA1c at 12 weeks.
with no registration number in the paper, or a registry entry that names 24 weeks, or one dated after enrollment began.
How to fix it. Register before the first participant enrolls and name the primary outcome and its time point precisely. Put the registry name and number in the abstract and methods. Make the reported primary outcome match the registered one. If it changed, report the change, when it was made and why. If a trial was registered late, say so rather than leaving the reviewer to find the date.
Short explainer: trial not registered
Checking your own manuscript
The most reliable check is a colleague who knows your field and reads the manuscript as a skeptic. A useful way to brief them is to send this list and ask which of these a reviewer would raise.
If you want a first pass before that, Purplelink's Paper Review ($9) reads a manuscript PDF and reports methodology concerns in the terms used on this page. Several of these problems are flagged automatically, including hyperparameter asymmetry, missing sample-size justification, single-seed results, figure presentation problems and references that do not resolve, and findings link to the matching short explainer. It is advisory, it can be wrong, and it does not replace a person who knows your field. Its field profiles list what each one checks: machine learning, biomedicine, psychology and social science, and chemistry and materials.
For citations alone, the free BibTeX Validator checks each entry in a .bib file against CrossRef and Semantic Scholar and flags entries that do not match a real paper. The file is checked on our server and not stored.
References
- Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., & Bellemare, M. G. (2021). Deep reinforcement learning at the edge of the statistical precipice. Advances in Neural Information Processing Systems 34 (NeurIPS 2021).
- Arnett, J. J. (2008). The neglected 95%: Why American psychology needs to become less American. American Psychologist, 63(7), 602–614.
- Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1), 289–300.
- Bouthillier, X., Delaunay, P., Bronzi, M., Trofimov, A., Nichyporuk, B., Szeto, J., et al. (2021). Accounting for variance in machine learning benchmarks. Proceedings of Machine Learning and Systems 3 (MLSys 2021).
- Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems 33 (NeurIPS 2020).
- Chambers, C. D. (2013). Registered Reports: A new publishing initiative at Cortex. Cortex, 49(3), 609–610.
- Chan, A.-W., Hróbjartsson, A., Haahr, M. T., Gøtzsche, P. C., & Altman, D. G. (2004). Empirical evidence for selective reporting of outcomes in randomized trials: Comparison of protocols to published articles. JAMA, 291(20), 2457–2465.
- Crameri, F., Shephard, G. E., & Heron, P. J. (2020). The misuse of colour in science communication. Nature Communications, 11.
- De Angelis, C., Drazen, J. M., Frizelle, F. A., Haug, C., Hoey, J., Horton, R., et al. (2004). Clinical trial registration: A statement from the International Committee of Medical Journal Editors. New England Journal of Medicine, 351(12), 1250–1251.
- Dodge, J., Gururangan, S., Card, D., Schwartz, R., & Smith, N. A. (2019). Show your work: Improved reporting of experimental results. Proceedings of EMNLP-IJCNLP 2019.
- Gelman, A., & Loken, E. (2013). The garden of forking paths: Why multiple comparisons can be a problem, even when there is no "fishing expedition" or "p-hacking" and the research hypothesis was posited ahead of time. Unpublished manuscript, Department of Statistics, Columbia University.
- Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460–465.
- Head, M. L., Holman, L., Lanfear, R., Kahn, A. T., & Jennions, M. D. (2015). The extent and consequences of p-hacking in science. PLoS Biology, 13(3).
- Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018). Deep reinforcement learning that matters. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI-18).
- Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83.
- Hoenig, J. M., & Heisey, D. M. (2001). The abuse of power: The pervasive fallacy of power calculations for data analysis. The American Statistician, 55(1), 19–24.
- Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70.
- International Committee of Medical Journal Editors. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals, section on clinical trial registration. icmje.org.
- Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124.
- Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9).
- Kerr, N. L. (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review, 2(3), 196–217.
- Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1).
- Lipton, Z. C., & Steinhardt, J. (2019). Troubling trends in machine learning scholarship. ACM Queue, 17(1).
- Lucic, M., Kurach, K., Michalski, M., Gelly, S., & Bousquet, O. (2018). Are GANs created equal? A large-scale study. Advances in Neural Information Processing Systems 31 (NeurIPS 2018).
- Melis, G., Dyer, C., & Blunsom, P. (2018). On the state of the art of evaluation in neural language models. International Conference on Learning Representations (ICLR 2018).
- Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
- Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems 37 (NeurIPS 2024).
- Percie du Sert, N., Hurst, V., Ahluwalia, A., Alam, S., Avey, M. T., Baker, M., et al. (2020). The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research. PLoS Biology, 18(7).
- Schulz, K. F., & Grimes, D. A. (2002). Blinding in randomised trials: Hiding who got what. The Lancet, 359(9307), 696–700.
- Schulz, K. F., Altman, D. G., & Moher, D., for the CONSORT Group. (2010). CONSORT 2010 Statement: Updated guidelines for reporting parallel group randomised trials. BMJ, 340, c332.
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.
- Simons, D. J., Shoda, Y., & Lindsay, D. S. (2017). Constraints on generality (COG): A proposed addition to all empirical papers. Perspectives on Psychological Science, 12(6), 1123–1128.
- Simonsohn, U., Simmons, J. P., & Nelson, L. D. (2020). Specification curve analysis. Nature Human Behaviour, 4(11), 1208–1214.
- Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11(5), 702–712.
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13.
- Weissgerber, T. L., Milic, N. M., Winham, S. J., & Garovic, V. D. (2015). Beyond bar and line graphs: Time for a new data presentation paradigm. PLoS Biology, 13(4), e1002128.
- Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track.