The outputs of this workflow are easy to read and easy to over-read. This page is about the second problem. It describes what each quantity establishes, and the claims that do not follow from it.
An adjusted p-value is a selection rule#
For locally fitted results, Benjamini–Hochberg (BH) adjustment targets the expected proportion of false discoveries in the p-value rejection set, under valid testing and dependence assumptions. BulkSeq Studio uses the strict rule padj < alpha. This error-rate statement concerns that rejection set across repeated experiments; it is not the probability that an individual gene is false, nor a guarantee about the realised fraction in one list. See Benjamini and Hochberg (1995).
The final up/down lists also apply an inclusive screen on the absolute raw log2 fold-change estimate. Such a post hoc effect-size screen does not automatically preserve the same FDR guarantee for the remaining genes. The informational padj_lfc_ge_threshold column instead comes from a hypothesis test against an effect-size threshold; it is not used for the main up/down calls. These are different questions, as explained by Love et al. (2014) and McCarthy and Smyth (2009). See calling thresholds for the implemented selection rule.
DESeq2's independent filtering is tuned to the alpha supplied to the fitted analysis, so rerunning it with a different alpha can change which weakly expressed genes receive adjusted p-values. Moving a cutoff on a fixed table does not refit that analysis. Imported adjusted values retain their declared source method and testing family; a column named padj alone does not establish BH control.
Not significant is not evidence of no change#
A gene that fails the threshold may be unchanged, or may be changed by an amount this experiment could not resolve. The two are not distinguishable from a p-value above a cut-off. If the absence of an effect is your claim, an equivalence test asks that question directly; a non-significant result does not answer it, however many genes share the outcome.
Fold change, shrunken and raw#
A log2 fold change on a base-two scale is an effect size: one unit is a doubling. Shrinkage moves the estimates of poorly measured genes towards zero, which makes ranking and plotting more stable but changes the numbers. This workflow computes shrunken estimates and records them, but it applies the fold-change threshold to the unshrunken value, so the gene list is selected on the raw estimate. Raw and shrunken estimates answer slightly different questions, so a figure and a table that disagree may simply be reporting different ones. Check which you are looking at before reconciling them.
Enrichment answers a narrower question than it appears to#
Over-representation asks whether your selected genes contain more members of a gene set than a background would give. That background — the tested universe — determines the answer as much as the selection does, and genes excluded from the universe cannot contribute to any term. Ranked enrichment asks a different question again, using the whole ranked list rather than a selection. Neither establishes that a pathway is active, that it is causal, or that it is regulated; both establish that annotated membership is distributed unevenly with respect to your statistic.
An empty result can be the correct result#
A small study with few selected genes often has no enrichment to report. An empty dot plot in that situation is an accurate summary, not a broken figure, and the enrichment summary file records why nothing was drawn. Reading it before assuming a fault will save more time than rerunning the step.
A protein-association network is a hypothesis#
STRING edges represent predicted functional associations, which can include physical binding, shared function or regulatory relationships. The combined association score integrates database evidence channels, including experiments, curated databases, coexpression and text mining; it is not a measured interaction strength or a probability of direct binding in your samples. A dense cluster suggests genes to investigate but does not independently confirm your differential-expression result. The display filter changes which associations are drawn, not their evidence, and layout distance has no biological unit. See the STRING interpretation guide.
What to report#
Report the engine, the thresholds, the contrast and its direction, the number of genes tested rather than only the number selected, and the software version that produced the numbers. Each changes what a reader can conclude. The settings are kept with the run in its configuration and parameter records, described on Read figures, tables and reports; the versions and environment are covered on Reproducibility and provenance. The output files themselves are described on Read figures, tables and reports and Enrichment and protein networks.