Before you start, save Samples, Reference, Analysis settings and Resources. A check of an earlier configuration does not certify a later edit.
1. Validate current inputs#
On Pre-run checks, click Validate current run inputs. Read each message and resolve failures involving file paths, schema, contrast levels, reference consistency or design rank.
Understand the status#
| Status | Meaning | What to do |
|---|---|---|
| FAIL | A condition prevents the requested workflow from proceeding correctly. | Fix the named cause and revalidate. Start is disabled until you do. Do not rename an invalid file to hide its format. |
| REVIEW_REQUIRED | A result or configuration needs interpretation before reliance. | Investigate the evidence, then tick the acknowledgement on Pre-run checks. It is neither a fatal error nor an acceptable result by default. |
| WARNING | An advisory condition may limit interpretation. | Read and record it; a completed run can still have warnings. |
| PASS / not applicable | A specific check passed, or the route does not produce the relevant quantity. | Neither establishes universal validity. A results-only route has no alignment-QC measurement. |
| STALE | Interface only. The saved input validation no longer matches the current configuration, sample sheet or inputs. | Validate again. Start is disabled until the checks match what you saved. |
The overall status is the most severe finding: FAIL, then REVIEW_REQUIRED, then WARNING, then PASS, with STALE above FAIL in the interface. A workflow already running stops for a FAIL in check 00, project setup; anything but PASS in check 05, reference validation, on FASTQ and SRA routes; or a failed microarray import recorded in checks 11 and 12. Meta-analysis also has a strict input gate: check 01 must be valid and report PASS or WARNING before the per-study fit. A REVIEW_REQUIRED, FAIL, missing or malformed check 01 stops that route. Other checks record their status for review after the run. Exit codes and statuses puts these beside the command-line and setup codes.
2. Inspect the workflow plan#
Use Dry Run before Start Run to validate the configured graph and see which steps would execute without running the full analysis.
What each check tests#
| Check | What it tests | Statuses | Stops a running workflow |
|---|---|---|---|
| 00 · project setup | Configuration paths and sections, the sample sheet, the design factors against its columns, imported-results provenance, and whether the R packages load. | PASS, WARNING, FAIL | Yes, on FAIL |
| 01 · input validation | Required sample-sheet columns, duplicate or unsafe sample names, replicates per condition, and study confounding in a multi-study contrast. | PASS, WARNING, FAIL | No; the interface will not start a run on its FAIL |
| 05 · reference validation | Genome and annotation integrity, the counting feature and attribute, and at least 95% agreement between FASTA and annotation sequence names. FASTQ and SRA only. | PASS, FAIL | Yes, on anything but PASS |
| 06 · alignment | Each sample’s uniquely mapped percentage from STAR or HISAT2. | PASS, WARNING, REVIEW_REQUIRED | No |
| 07 · quantification | Each sample’s assigned-read fraction, or its library size for an imported count matrix. | PASS, WARNING, REVIEW_REQUIRED, FAIL | No |
| 08 · design | Whether the design matrix is full rank, and interaction terms or numeric covariates that need review. | PASS, REVIEW_REQUIRED, FAIL | No |
| 09 · differential expression | Whether any gene reaches the significance threshold; none is REVIEW_REQUIRED. | PASS, REVIEW_REQUIRED | No |
| 10 · enrichment | Whether enrichment ran, its annotation coverage, how well identifiers mapped, and whether the KEGG code names the configured organism. | PASS, WARNING, REVIEW_REQUIRED | No |
| 11 · normalization | Microarray only: probe count, the normalization applied and the log2 decision. | PASS, WARNING, FAIL | Yes, when the import fails |
| 12 · probe mapping | Microarray only: the fraction of probes that resolve to exactly one gene; below half is REVIEW_REQUIRED, none is FAIL. | PASS, REVIEW_REQUIRED, FAIL | Yes, on FAIL |
| 13 · equivalence | DESeq2 only: genes statistically equivalent to no change at the configured threshold. Informational. | PASS | No |
| 14 · Wilcoxon sensitivity | A rank-sum cross-check of the primary calls, with a warning when the smallest group has fewer than five samples. | PASS, WARNING | No |
| 15 · set overlap | Overlap of the differential-expression genes with MSigDB Hallmark sets for the organism. | PASS, REVIEW_REQUIRED | No |
| 16 · protein network | Whether a STRING network was built; a STRING or layout failure is a WARNING. | PASS, WARNING | No |
| 17 · meta-analysis | Multi-study only: the studies admitted, shared genes and the combined result; FAIL with fewer than two admissible studies or no shared genes. | PASS, REVIEW_REQUIRED, FAIL | No |
| 18 · meta-enrichment | Multi-study only: cross-study enrichment and how identifiers mapped to Entrez IDs. | PASS, REVIEW_REQUIRED | No |
| 19 · orientation | A contrast whose numerator and denominator look inverted, or an unconfirmed imported-results direction. | PASS, REVIEW_REQUIRED, FAIL | No |
| 20 · duplicate study | Multi-study only: available per-sample counts and fold changes are screened for duplicate or near-duplicate studies. Coverage is assessed, partial or unassessed; missing evidence requires review and cannot prove independence. | PASS, REVIEW_REQUIRED | No |
| 21 · strandedness | Samples in one study that disagree on inferred strandedness, and studies with divergent assignment fractions. | PASS, REVIEW_REQUIRED | No |
| 22 · sample structure | Whether replicates cluster more tightly than samples of other conditions. | PASS, WARNING | No |
| 23 · covariate structure | Sample-sheet columns outside the design that track PC1 or PC2 beyond chance, or alias the contrast. | PASS, WARNING, REVIEW_REQUIRED | No |
| 24 · custom enrichment | Only with custom gene sets: whether they overlap the differential-expression genes. | PASS, REVIEW_REQUIRED | No |
| 25 · annotation transfer | Only when annotation-transfer enrichment runs: the share of tested genes mapped to STRING proteins or to an imported annotation, and whether the tested genes are as well annotated as the organism’s proteome. | PASS, WARNING, REVIEW_REQUIRED | No |
Checks 22 and 23 use WARNING differently. In check 22 the WARNING message says which of two things happened: not assessable when replicate structure could not be assessed, and advisory finding when replicates cluster no more tightly than other conditions. In check 23 a WARNING always means the screen could not run, and a finding is REVIEW_REQUIRED. Check 08 can list a WARNING message, such as a condition with fewer than two replicates, without that becoming its overall status. Check 01 is written twice, by the interface before the run and by the workflow when the run starts, from one replicate rule: a condition with fewer than two biological replicates, or fewer than the recommended three, is a WARNING, and a sample with an empty or unknown condition is REVIEW_REQUIRED. A design with two replicates per group therefore reads WARNING before and after the run. When meta-analysis is enabled, check 01 must be valid and report PASS or WARNING for its strict input gate; REVIEW_REQUIRED, FAIL or missing evidence stops per-study fitting.
3. Review checks produced during and after the run#
Read checks/sanity_checks.txt together with the individual machine-readable checks. Applicability depends on route. Alignment and assignment checks require their corresponding measurements; imported counts receive library-size checks instead of read-assignment fractions.
Review low mapping, empty/outlying libraries, sample structure, within-study strandedness disagreements and enrichment/network availability. A lack of adjusted-p hits is reported for review; that is not proof of a failed experiment or a reason to relax thresholds automatically.
Advanced: checks with narrow meanings
Check 23 is an output-stage screen on every route that fits a local model and writes results/deseq2/pca_coordinates.csv: DESeq2, edgeR, limma-voom and microarray limma. Only imported results are excluded, because they carry no expression matrix. Columns that name a sample rather than describe its biology are left out of the comparison: the sample identifier, the file paths, the ingest accessions, and the free-text labels sample_title, title and library_name. A label often groups the same way the contrast does, because that is how people write labels, and that resemblance is not evidence of unmodelled structure. A recorded technical factor such as batch, library_prep or sequencing_run is still compared. It compares each sample-sheet column outside the design formula with PC1 and PC2 and reports review-required when that column’s adjusted R² exceeds the value only 5% of random label assignments reach at the study’s number of samples and groups (0.85 for four samples in two groups, 0.42 for eight, 0.17 for eighteen) and the F-test p-value is below 0.05. Adjusted R² rises monotonically with the F statistic, so those two conditions are one calibrated test reported twice; the fixed floor of 0.5 it replaces bound only at small sample sizes. A column whose grouping coincides exactly with the contrast factor is reported as aliased with it instead of scored, and each engine screens its own coordinates, so a verdict is engine-relative. The screen is advisory and does not change the design.
Check 13 and results/deseq2/unchanged_genes.csv are produced on the count-based DESeq2 route only; edgeR, limma-voom, microarray limma and imported results do not write them. Check 13 records completion of a TOST-style calculation; its pass is not an independent proof of equivalence. Check 19 uses configuration names to flag likely orientation mistakes; it cannot know your experimental intent. Checks 02–04 are unused in the published baseline. Read exact rule and route evidence rather than interpreting a check count as a quality score.
When a check and your expectation disagree#
Inspect the underlying inputs and report. A batch associated with a PCA axis can be a useful warning, while a real biological grouping can also explain variation. Keep the original run record before deliberately revising the model. See design guidance.