BulkSeq Studiov0.34.0
Download

Explanation

What the workflow checks for you

The verification the application performs on your behalf, and the judgements it deliberately leaves to you.

On this page

A pipeline can complete without producing a defensible result. The safeguards described here exist because the failures that matter in expression analysis are quiet ones: an input replaced between two runs, a library orientation assumed rather than measured, a probe standing for two genes. Each is checked because it can be checked mechanically. What remains is the part that cannot be, and that part is yours.

Completion is not validity#

A successful run means the computation finished. It does not mean the design was adequate, the samples were comparable, the contrast answered your question, or the result is biologically meaningful. The application is built so that the record of what it did survives the run and can be examined afterwards; it is not built to certify the study.

Validation is revalidated#

Pre-run checks store content fingerprints for the configuration, the sample sheet, the local inputs, the reference locks and the index files. Starting or resuming revalidates them. The reason is specific: a validation result that is remembered rather than recomputed will happily certify a file that has since been replaced, and the analyst has no way to see that from the output. An edited input cannot inherit an earlier pass.

Strandedness is decided per sample#

Each sample's library type is inferred from its own data and applied with its own counting orientation. A project that mixes library types is named as such rather than forced onto one code. Guessing a single orientation for a heterogeneous set of libraries produces counts that look ordinary and are wrong for a subset of the samples, which is exactly the kind of error that survives review.

Unmodelled structure is screened, not corrected#

For every sample-sheet column outside the design formula that varies between samples without being unique to each one, the workflow tests that column against the run's principal components and reports the result. A column holding one value throughout, or a different value in every row, carries no structure the test could detect and is skipped without a verdict of its own. This is a warning system, not an adjustment: nothing is added to your model automatically. Two engines can reach different verdicts on the same samples, because each screens its own coordinates, so the screen tells you where to look rather than what to conclude.

Descriptive columns are excluded from the screen. A library name or a sample title labels a sample; testing it against the principal components asks a question with no answer, and reporting the result would invite a reader to act on noise.

Imported results are not relabelled#

A differential-expression table you supply is validated in full, bound to its hash and schema, and reported as what it is. The report never claims a model the application did not fit. The alternative — presenting an imported table under the application's own method description — would make an analysis look reproducible by this workflow when it is not.

Count matrices are validated before conversion#

Missing, nonnumeric, nonfinite and negative cells are refused before anything is written. A non-integer value requires you to declare explicitly that the matrix holds RSEM or tximport estimated counts, and is then rounded by a stated rule. When at least half the sample columns have totals within one per cent of a million, a warning says the data may be normalized; that warning does not by itself reclassify valid integer counts as TPM. The distinction matters because count models assume counts, and a normalized matrix passed to one produces confident nonsense.

Ambiguous microarray probes are excluded#

Platform annotation is parsed across every listed candidate and every row for a probe. Only probes resolving to exactly one gene enter the collapse step. Ambiguous, unknown and missing mappings are retained in the probe-map evidence and counted, rather than silently assigned to whichever gene happens to be listed first.

This applies to the route that has platform annotation to parse. A gene-level matrix you supply yourself carries no probe-to-gene table, so its row identifiers are taken as gene identifiers directly and no ambiguity screening happens at all. The identifiers you supply are the identifiers the analysis uses.

What none of this checks#

No safeguard here evaluates whether your replicates are biological or technical, whether the contrast you configured is the comparison your question needs, whether a batch you did not record is driving the result, or whether the effect you found is large enough to matter. Those judgements are not mechanical, and presenting them as passed checks would be worse than leaving them open. Read Pre-run and result checks for what is tested, and What the numbers mean for how far a result reaches.

Search the documentation

Type to search every page.

Figure viewer

100%