Reproducible Research Code: Best Practices for Researchers
Reproducible research code, step by step: lockfiles, containers, seeds, data versioning, experiment tracking, testing and releasing code with your paper.

Reproducible research code means someone else (including you, six months from now) can take your repository, set up the same environment, run one command, and get the same figures and tables. You get there with a few habits: lock your dependencies, pin your data, control randomness, track experiments, write a few tests, and release everything with a DOI. None of this needs a software engineering degree, and this guide shows the minimum that works.
Why research code stops working
Research code usually works on the day of the deadline. It breaks later because:
- A dependency released a new version with different defaults.
- The data file was edited in place and nobody knows which version produced Table 2.
- Results depended on a random seed, a GPU type, or a notebook cell run out of order.
- The code only ran on one person’s laptop, with paths like
/Users/anna/Desktop/final_v3/.
Each fix below targets one of these failure modes.
1. Start with a predictable project layout
A consistent structure makes it obvious where things live:
project/
├── README.md # what this is, how to run it
├── LICENSE
├── CITATION.cff
├── pyproject.toml # dependencies
├── uv.lock # exact resolved versions
├── configs/ # one YAML file per experiment
├── data/
│ ├── raw/ # never edited by hand
│ └── processed/ # produced by scripts
├── src/project_name/ # importable code
├── scripts/ # entry points: prepare, train, evaluate, plot
├── notebooks/ # exploration only
├── tests/
└── results/ # generated outputs (often not committed)
Use relative paths from the project root, and read machine-specific settings (data directory, API keys) from environment variables or a local config file that is not committed.
2. Lock your environment
A requirements.txt with unpinned names is not enough. You need a lockfile that records the exact version of every package, including indirect dependencies.
As of September 2026, good options are:
- uv for pure Python projects. It writes a
uv.lockfile, anduv sync --lockedfails if the lockfile is out of date instead of silently changing it. - pixi when you need conda packages (CUDA libraries, GDAL, R, compilers). It supports both conda-forge and PyPI packages and writes a
pixi.lockfile. - conda with conda-lock, or
renvfor R projects andManifest.tomlin Julia.
# Create a locked Python project with uv
uv init my-study
cd my-study
uv add numpy pandas scikit-learn matplotlib
uv lock # writes uv.lock
# Collaborators and CI reproduce the exact environment:
uv sync --locked
uv run python scripts/train.py --config configs/baseline.yaml
Commit both pyproject.toml and the lockfile. Record the Python version too.
3. Use containers for the full system
Lockfiles pin Python packages, but not the operating system, system libraries or CUDA drivers. A container captures those.
- Docker is the standard for building images and works well on laptops and cloud machines.
- Apptainer (formerly Singularity) is common on university HPC clusters, where Docker is often not allowed. It can run images built from Docker.
FROM python:3.12-slim
COPY --from=ghcr.io/astral-sh/uv:latest /uv /usr/local/bin/uv
WORKDIR /app
COPY pyproject.toml uv.lock ./
RUN uv sync --locked --no-dev
COPY . .
CMD ["uv", "run", "python", "scripts/reproduce_all.py"]
Pin the base image by version (or by digest for stronger guarantees), and publish the built image with your release.
4. Control randomness
Set seeds for every source of randomness, and save the seed with the results.
import os
import random
import numpy as np
import torch
def set_seed(seed: int) -> None:
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed)
torch.cuda.manual_seed_all(seed)
os.environ["PYTHONHASHSEED"] = str(seed)
# Slower, but makes many GPU operations deterministic
torch.backends.cudnn.benchmark = False
torch.use_deterministic_algorithms(True, warn_only=True)
Two important caveats:
- Some GPU operations are not deterministic even with seeds, and results can differ across GPU models and library versions. PyTorch’s reproducibility notes list the details.
- The scientific answer is not “one lucky seed”. Run several seeds and report the mean and spread. A result that only holds for one seed is not a result.
5. Version your data
Code in git is only half of the story. You also need to know exactly which data produced which result.
- Never edit raw data by hand. Keep it read-only, and produce cleaned data with scripts.
- Record a checksum (e.g. SHA-256) of every input file in your results metadata.
- Use a data versioning tool for larger datasets. DVC stores small pointer files in git and the data itself in cloud or network storage, so a git commit identifies both code and data. git-lfs also works for moderate file sizes.
- Archive the final dataset in a repository with a persistent identifier, such as Zenodo or a domain or institutional repository, when licences and ethics approvals allow.
dvc init
dvc add data/raw/survey_2026.csv
git add data/raw/survey_2026.csv.dvc .gitignore
git commit -m "Track raw survey data with DVC"
dvc remote add -d storage s3://my-bucket/dvc-store
dvc push
6. Drive experiments from config files
Hard-coded parameters scattered through scripts are a common reason results can’t be reproduced. Put every parameter in a config file, and save a copy of the config next to each run’s output.
# configs/baseline.yaml
seed: 42
data:
path: data/processed/train.parquet
sha256: "3a7bd3e2..."
model:
type: logistic_regression
C: 1.0
output_dir: results/baseline
Tools such as Hydra help when you have many configurations and sweeps, but plain YAML plus argparse is enough for many projects.
7. Track experiments
When you run dozens or hundreds of experiments, a spreadsheet stops working. Experiment trackers record parameters, metrics, code version and output files for each run automatically. As of September 2026:
- MLflow is open source and can run locally or on a lab server.
- Weights & Biases is a hosted service that is free for students, educators and academic researchers, according to its research page.
- DVC experiments work well if you already use DVC for data.
Whatever you use, make sure each run records the git commit hash, and refuse to run from a dirty working tree for final results.
8. Notebooks vs scripts
Notebooks are excellent for exploration and for explaining results. They are risky as the only record of an analysis, because cells can run out of order and hidden state is invisible.
A practical split:
- Explore in notebooks. Keep them in
notebooks/, numbered and dated. - Move anything you depend on into functions in
src/, and import them in the notebook. - Generate final results with scripts that run top to bottom from a clean state.
- Before you share a notebook, restart the kernel and run all cells. Tools like
nbconvert --executeor papermill can run notebooks in CI.
9. Test the parts that matter
You don’t need full test coverage. Test the code where a silent bug would change your conclusions:
- Data loading and cleaning (row counts, missing values, units)
- Metric and statistics functions (compare against a hand-calculated example)
- A small end-to-end “smoke test” that runs the pipeline on a tiny dataset
# tests/test_metrics.py
from project_name.metrics import accuracy
def test_accuracy_simple_case():
assert accuracy([1, 0, 1, 1], [1, 0, 0, 1]) == 0.75
Run the tests automatically on every push with a simple CI workflow (for example GitHub Actions or GitLab CI). A CI job that installs from the lockfile and runs the smoke test also proves your environment can be recreated from scratch.
10. One command to reproduce everything
Give readers a single entry point, such as a Makefile or a reproduce_all.py script, that goes from raw data to every figure and table in the paper.
all: results/figures
data/processed/train.parquet: data/raw/survey_2026.csv
uv run python scripts/prepare.py
results/metrics.json: data/processed/train.parquet
uv run python scripts/train.py --config configs/baseline.yaml
results/figures: results/metrics.json
uv run python scripts/plot.py
Say in the README how long it takes and what hardware it needs. If full training takes weeks on a cluster, also provide trained checkpoints and a quick path that reproduces the tables from saved outputs.
11. Release code with the paper
- Licence: choose one explicitly. Without a licence, others cannot legally reuse your code. Permissive licences (MIT, BSD-3-Clause, Apache-2.0) are common for research code; GPL-family licences require derived works to stay open. Check your institution’s and funder’s policies, and use choosealicense.com to compare. Data often needs a different licence (e.g. Creative Commons).
- Citation: add a
CITATION.cfffile so people know how to cite the software. - DOI: connect the repository to Zenodo; each GitHub release then gets its own DOI (GitHub docs). Cite the exact release in the paper.
- README: installation, a quick start, how to reproduce each figure, expected runtime, and contact details.
- Tag the release that matches the paper version, so later changes don’t confuse readers.
Many venues and journals now include reproducibility or code-availability checklists. Following the steps above covers most of what they ask.
Reproducible research code checklist
- Clear project layout; no absolute paths
- Dependencies locked (uv, pixi, conda-lock or equivalent) and lockfile committed
- Container definition for the full system, with a pinned base image
- Seeds set and saved; results reported across several seeds
- Raw data read-only, checksummed and versioned
- All parameters in config files, saved with each run
- Runs tracked with git commit hashes
- Final results produced by scripts, not only notebooks
- Tests for data processing and metrics, run in CI
- One command reproduces all figures and tables
- Licence, CITATION.cff, DOI and a tagged release
Need a hand with the engineering?
Many researchers are experts in their field but find this setup work slow and frustrating. That’s normal. It’s software engineering, not science. We help researchers turn working-but-fragile code into reproducible pipelines, build research tools, and run benchmarks at scale. If you’re evaluating models, see also our guide to benchmarking LLMs for research. To talk about your project, see our research engineering services or contact us.