BulkSeq Studiov0.34.0
Download

Reference

Pre-run and result checks

Separate input readiness, execution success and scientific interpretation.

On this page

Before you start, save Samples, Reference, Analysis settings and Resources. A check of an earlier configuration does not certify a later edit.

1. Validate current inputs#

On Pre-run checks, click Validate current run inputs. Read each message and resolve failures involving file paths, schema, contrast levels, reference consistency or design rank.

Understand the status#

StatusMeaningWhat to do
FAILA condition prevents the requested workflow from proceeding correctly.Fix the named cause and revalidate. Start is disabled until you do. Do not rename an invalid file to hide its format.
REVIEW_REQUIREDA result or configuration needs interpretation before reliance.Investigate the evidence, then tick the acknowledgement on Pre-run checks. It is neither a fatal error nor an acceptable result by default.
WARNINGAn advisory condition may limit interpretation.Read and record it; a completed run can still have warnings.
PASS / not applicableA specific check passed, or the route does not produce the relevant quantity.Neither establishes universal validity. A results-only route has no alignment-QC measurement.
STALEInterface only. The saved input validation no longer matches the current configuration, sample sheet or inputs.Validate again. Start is disabled until the checks match what you saved.

The overall status is the most severe finding: FAIL, then REVIEW_REQUIRED, then WARNING, then PASS, with STALE above FAIL in the interface. A workflow already running stops for a FAIL in check 00, project setup; anything but PASS in check 05, reference validation, on FASTQ and SRA routes; or a failed microarray import recorded in checks 11 and 12. Meta-analysis also has a strict input gate: check 01 must be valid and report PASS or WARNING before the per-study fit. A REVIEW_REQUIRED, FAIL, missing or malformed check 01 stops that route. Other checks record their status for review after the run. Exit codes and statuses puts these beside the command-line and setup codes.

2. Inspect the workflow plan#

Use Dry Run before Start Run to validate the configured graph and see which steps would execute without running the full analysis.

What each check tests#

CheckWhat it testsStatusesStops a running workflow
00 · project setupConfiguration paths and sections, the sample sheet, the design factors against its columns, imported-results provenance, and whether the R packages load.PASS, WARNING, FAILYes, on FAIL
01 · input validationRequired sample-sheet columns, duplicate or unsafe sample names, replicates per condition, and study confounding in a multi-study contrast.PASS, WARNING, FAILNo; the interface will not start a run on its FAIL
05 · reference validationGenome and annotation integrity, the counting feature and attribute, and at least 95% agreement between FASTA and annotation sequence names. FASTQ and SRA only.PASS, FAILYes, on anything but PASS
06 · alignmentEach sample’s uniquely mapped percentage from STAR or HISAT2.PASS, WARNING, REVIEW_REQUIREDNo
07 · quantificationEach sample’s assigned-read fraction, or its library size for an imported count matrix.PASS, WARNING, REVIEW_REQUIRED, FAILNo
08 · designWhether the design matrix is full rank, and interaction terms or numeric covariates that need review.PASS, REVIEW_REQUIRED, FAILNo
09 · differential expressionWhether any gene reaches the significance threshold; none is REVIEW_REQUIRED.PASS, REVIEW_REQUIREDNo
10 · enrichmentWhether enrichment ran, its annotation coverage, how well identifiers mapped, and whether the KEGG code names the configured organism.PASS, WARNING, REVIEW_REQUIREDNo
11 · normalizationMicroarray only: probe count, the normalization applied and the log2 decision.PASS, WARNING, FAILYes, when the import fails
12 · probe mappingMicroarray only: the fraction of probes that resolve to exactly one gene; below half is REVIEW_REQUIRED, none is FAIL.PASS, REVIEW_REQUIRED, FAILYes, on FAIL
13 · equivalenceDESeq2 only: genes statistically equivalent to no change at the configured threshold. Informational.PASSNo
14 · Wilcoxon sensitivityA rank-sum cross-check of the primary calls, with a warning when the smallest group has fewer than five samples.PASS, WARNINGNo
15 · set overlapOverlap of the differential-expression genes with MSigDB Hallmark sets for the organism.PASS, REVIEW_REQUIREDNo
16 · protein networkWhether a STRING network was built; a STRING or layout failure is a WARNING.PASS, WARNINGNo
17 · meta-analysisMulti-study only: the studies admitted, shared genes and the combined result; FAIL with fewer than two admissible studies or no shared genes.PASS, REVIEW_REQUIRED, FAILNo
18 · meta-enrichmentMulti-study only: cross-study enrichment and how identifiers mapped to Entrez IDs.PASS, REVIEW_REQUIREDNo
19 · orientationA contrast whose numerator and denominator look inverted, or an unconfirmed imported-results direction.PASS, REVIEW_REQUIRED, FAILNo
20 · duplicate studyMulti-study only: available per-sample counts and fold changes are screened for duplicate or near-duplicate studies. Coverage is assessed, partial or unassessed; missing evidence requires review and cannot prove independence.PASS, REVIEW_REQUIREDNo
21 · strandednessSamples in one study that disagree on inferred strandedness, and studies with divergent assignment fractions.PASS, REVIEW_REQUIREDNo
22 · sample structureWhether replicates cluster more tightly than samples of other conditions.PASS, WARNINGNo
23 · covariate structureSample-sheet columns outside the design that track PC1 or PC2 beyond chance, or alias the contrast.PASS, WARNING, REVIEW_REQUIREDNo
24 · custom enrichmentOnly with custom gene sets: whether they overlap the differential-expression genes.PASS, REVIEW_REQUIREDNo
25 · annotation transferOnly when annotation-transfer enrichment runs: the share of tested genes mapped to STRING proteins or to an imported annotation, and whether the tested genes are as well annotated as the organism’s proteome.PASS, WARNING, REVIEW_REQUIREDNo

Checks 22 and 23 use WARNING differently. In check 22 the WARNING message says which of two things happened: not assessable when replicate structure could not be assessed, and advisory finding when replicates cluster no more tightly than other conditions. In check 23 a WARNING always means the screen could not run, and a finding is REVIEW_REQUIRED. Check 08 can list a WARNING message, such as a condition with fewer than two replicates, without that becoming its overall status. Check 01 is written twice, by the interface before the run and by the workflow when the run starts, from one replicate rule: a condition with fewer than two biological replicates, or fewer than the recommended three, is a WARNING, and a sample with an empty or unknown condition is REVIEW_REQUIRED. A design with two replicates per group therefore reads WARNING before and after the run. When meta-analysis is enabled, check 01 must be valid and report PASS or WARNING for its strict input gate; REVIEW_REQUIRED, FAIL or missing evidence stops per-study fitting.

3. Review checks produced during and after the run#

Read checks/sanity_checks.txt together with the individual machine-readable checks. Applicability depends on route. Alignment and assignment checks require their corresponding measurements; imported counts receive library-size checks instead of read-assignment fractions.

Review low mapping, empty/outlying libraries, sample structure, within-study strandedness disagreements and enrichment/network availability. A lack of adjusted-p hits is reported for review; that is not proof of a failed experiment or a reason to relax thresholds automatically.

Advanced: checks with narrow meanings

Check 23 is an output-stage screen on every route that fits a local model and writes results/deseq2/pca_coordinates.csv: DESeq2, edgeR, limma-voom and microarray limma. Only imported results are excluded, because they carry no expression matrix. Columns that name a sample rather than describe its biology are left out of the comparison: the sample identifier, the file paths, the ingest accessions, and the free-text labels sample_title, title and library_name. A label often groups the same way the contrast does, because that is how people write labels, and that resemblance is not evidence of unmodelled structure. A recorded technical factor such as batch, library_prep or sequencing_run is still compared. It compares each sample-sheet column outside the design formula with PC1 and PC2 and reports review-required when that column’s adjusted R² exceeds the value only 5% of random label assignments reach at the study’s number of samples and groups (0.85 for four samples in two groups, 0.42 for eight, 0.17 for eighteen) and the F-test p-value is below 0.05. Adjusted R² rises monotonically with the F statistic, so those two conditions are one calibrated test reported twice; the fixed floor of 0.5 it replaces bound only at small sample sizes. A column whose grouping coincides exactly with the contrast factor is reported as aliased with it instead of scored, and each engine screens its own coordinates, so a verdict is engine-relative. The screen is advisory and does not change the design.

Check 13 and results/deseq2/unchanged_genes.csv are produced on the count-based DESeq2 route only; edgeR, limma-voom, microarray limma and imported results do not write them. Check 13 records completion of a TOST-style calculation; its pass is not an independent proof of equivalence. Check 19 uses configuration names to flag likely orientation mistakes; it cannot know your experimental intent. Checks 02–04 are unused in the published baseline. Read exact rule and route evidence rather than interpreting a check count as a quality score.

When a check and your expectation disagree#

Inspect the underlying inputs and report. A batch associated with a PCA axis can be a useful warning, while a real biological grouping can also explain variation. Keep the original run record before deliberately revising the model. See design guidance.

Search the documentation

Type to search every page.

Figure viewer

100%