Reproducibility here means something narrow and checkable: that the record a run leaves behind is sufficient to say what was computed, with which software, from which inputs. That record is written whether or not anyone reads it, which is the point — its value appears months later, when the question arrives and the run is long finished.
What a run records#
The run summary names the application and workflow versions that actually executed, which are not necessarily the versions bundled with the application you launched. It records the execution-tree digest separately from the bundled-workflow identity, the installed environment specification, and the tools and reference files belonging to the route that ran. Any setting that differs from the bundled default is recorded with both its default and the value used, so a customization is never left implicit. A control set by hand to the value it already had is indistinguishable from one never touched, because the two produce the same configuration.
For a linux-64 full installation, the exact lock specifies Conda package names, versions and builds plus four pip distribution versions. After tool and R load probes, setup checks those installed identities before recording a lock success marker. A mismatch enters one bounded reinstall and recheck; an unresolved mismatch fails setup. The validator reports extra packages without deleting them. When the exact lock is unavailable and the floating full specification is installed, provenance says fallback and does not claim a lock match. This record does not by itself establish that a biological analysis was rerun.
Workflow synchronization stops if a recorded project copy has local edits, rather than overwriting them. An edited workflow is a legitimate thing to have; an edited workflow silently replaced mid-study is not.
Two versions, two jobs#
The application version and the workflow version answer different questions. The application version identifies the interface you used. The workflow version governs whether an existing project re-copies the bundled workflow, but it is not the only trigger: a project also re-syncs when the bundled workflow's file content has changed since its last copy, so a same-version revision of the analysis code still reaches projects that already exist. When you compare two runs, the versions in their summaries are the ones that matter, not the version of the application currently installed.
Why results can move between releases#
Some releases change a scientific output: a mapping rule, a selection boundary, a ranking input. When that happens the change is marked in the changelog with a re-run notice, and the site's version notices list which release changed which output. This is deliberate. A silent correction would leave two incompatible sets of numbers circulating with no way to tell them apart.
If you have reported results from an earlier release, a notice affecting your route identifies what may change on rerun; assess the claim against its recorded method. For pooled-effect meta adjusted values produced before 0.34.0, recompute before interpreting that adjusted family. Preserve the original result and provenance while reviewing the corrected output.
Cite the version that produced your numbers#
Cite the exact application version recorded in your run summary, not the latest release. Troubleshooting, versions and citation carries the current citation guidance.
What provenance does not give you#
A complete record of what ran is not a claim that what ran was appropriate. Provenance makes an analysis auditable; it does not make it correct, and it is not a substitute for the judgements described in What the workflow checks for you.