Purplelink
← All tools

Free web tool

Replication Package Checker

Drop in the .zip of your replication package: code, data and README. You get a reproducibility check with ten results, each with file paths and line numbers: README sections, license, dependencies, entry point, absolute paths, random seeds, data files, outputs and stray files. The zip never leaves your browser, and none of your code is run.

Up to 500 MB. Files over 5 MB are listed but not read.

What it checks

Each check is a pattern match on file names and text. The source beside it is where the expectation comes from.

  1. README. Looks for a README at the top level of the package, then matches the headings of a text README against the seven sections of the template: Data Availability and Provenance Statements, Dataset list, Computational requirements, Description of programs/code, Instructions to replicators, List of tables and programs, and References. It also looks for a runtime and a hardware statement. The standard asks for a README with a Data Availability Statement, the software and hardware requirements including the expected run time, and instructions for reproducing the results, following the template. The AEA guidance asks for the README in the package, ideally in the root directory. Source: Social Science Data Editors' template README; Data and Code Availability Standard, rule 13; AEA Data Editor, Preparing your files for verification
  2. License. Looks for a LICENSE file or a license statement in the README. The standard asks for a license that sets the terms of use of code and data. The template says the license should be in a LICENSE.txt file, separate from the README. Source: Data and Code Availability Standard, rule 15; Social Science Data Editors' template README
  3. Environment. For each language found, looks for a dependency file (requirements.txt, environment.yml, pyproject.toml, Pipfile.lock, renv.lock, DESCRIPTION, Project.toml and Manifest.toml), checks whether versions are exact, and looks for the software version in the README. Stata has no manifest, so it reports where ado packages are installed. The template says to list all software requirements and, in all cases, the version you used. The AEA guidance says not to assume the replicator has your packages installed and to provide a setup program. Source: Social Science Data Editors' template README; AEA Data Editor, Preparing your files for verification
  4. Entry point. Looks for a main, master or run script or a Makefile, or for numbered scripts. The template says that when replication takes more than four or five manual steps, a main program or Makefile should wrap them. Source: Social Science Data Editors' template README
  5. Paths. Lists every line of code with an absolute path: a Windows drive letter, /Users/, /home/, ~/, setwd() or cd with an absolute path. The AEA guidance says code and data should run as downloaded, without further manual changes, and that the only exception is a single change to set a small number of program and data directory paths. Source: AEA Data Editor, Preparing your files for verification
  6. Randomness. Finds code that draws random numbers and checks for a seed in the same file or in the entry script. The template says code that uses a pseudorandom number generator should be given a deterministic seed, set once and not from a time stamp, and asks the README to name the line and program where it is set. Source: Social Science Data Editors' template README
  7. Data. Compares the data files named in code with the data files in the package, in both directions, and reports size by folder and any file over 100 MB. The standard asks for analysis data to be in the package unless it can be fully reproduced from accessible data in reasonable time. The template asks for a list of datasets and whether each is provided. The AEA guidance advises keeping data and code in separate folders. The 100 MB line is this tool's own cut-off, not a rule from these sources. Source: Data and Code Availability Standard, rule 3; Social Science Data Editors' template README; AEA Data Editor, Preparing your files for verification
  8. Outputs. Looks for code that writes tables or figures, for an output folder, and for README lines that tie a table or figure number to a program. The template asks for a list of tables and figures that identifies the program, and possibly the line number, where each is created. Source: Social Science Data Editors' template README
  9. Hygiene. Lists system and editor debris (.DS_Store, Thumbs.db, .Rhistory, __pycache__, .ipynb_checkpoints), lines that look like a stored API key, token or password (the value is hidden in the report), and notebooks with more than 1 MB of stored outputs. This check is not taken from the three sources. It is here because a deposit is public and these files are not needed to reproduce anything.
  10. Summary. Lists the languages found with file counts and says what a person still has to do. The AEA guidance says the package should reproduce the tables, figures and in-text numbers by running code without manual intervention. Only running it shows that. Source: AEA Data Editor, Preparing your files for verification

Want it checked against your journal's policy?

The checks above are the same for every journal. A paid AI read of the package against a named journal's data and code policy is being considered. Tell me which journal or data editor you are submitting to, and I will build it if enough people ask for the same one.

Questions

Is my replication package uploaded?
No. The zip is opened in your browser with a copy of JSZip hosted on this site, and nothing is sent anywhere. You can disconnect from the network after the page loads and it still works.
Does it run my code?
No. It reads file names and the text of code files and matches patterns. It cannot tell whether the code runs or whether the output matches the paper. Someone still has to run the package from start to finish on a clean machine.
Which standards does it follow?
The Social Science Data Editors' template README, the Data and Code Availability Standard, and the AEA Data Editor's guidance on preparing a deposit. Each check on this page links to the source it comes from. The hygiene check is this tool's own and is marked as such.
Why does it say Review instead of Fail?
A pattern match can be right and the package still be fine. A data file may be absent because the README tells the replicator to download it, and a path may sit in the one configuration file a replicator is expected to edit. Review means a person should look at those lines.
Will the data editor accept my package if every check passes?
Not necessarily. Journals set their own policies, and a data editor runs the code, which this tool does not. Read the policy of the journal you are submitting to.
What does it read?
A zip up to 500 MB. It reads R, R Markdown, Quarto, Python, Jupyter, Stata, Julia, MATLAB, SAS and shell files, Makefiles, and Markdown, text, YAML, TOML and lock files. Files over 5 MB are listed but not read, and only the header row of a CSV is kept.