BulkSeq Studiov0.34.0
Download

Explanation

Choosing a differential-expression engine

Why DESeq2 is the default, when another engine is a better fit, and what changes when you switch.

On this page

Four differential-expression engines are available: DESeq2, edgeR quasi-likelihood and limma-voom for count data, and limma for microarray intensities. They are not interchangeable spellings of one method. Each makes different assumptions about how variance behaves, and on a given dataset they will agree about the strong effects and disagree at the margin.

What they have in common#

All four fit the design you configure and write the same results schema, so the downstream figures, tables, enrichment and reports do not change shape when you change engine. They do not all get there the same way. DESeq2 receives your formula directly and names its contrast by level. The three alternatives are fitted through a group-means matrix and form the numerator-minus-denominator comparison by factor position, which is what keeps condition levels that differ only in spacing or punctuation distinct and stops a sample-sheet column from displacing the contrast factor you configured.

The shared schema has documented gaps rather than identical contents. Microarray limma assigns no NCBI GeneID and no biotype; edgeR quasi-likelihood reports no per-gene standard error for the log fold change, so that column is empty rather than different; and the equivalence test is produced on the DESeq2 route only.

DESeq2, the default#

DESeq2 is the default because it behaves well on the designs this application is most often used for: few replicates per group, count data, and a single contrast of interest. It estimates dispersion by pooling information across genes, which is what makes a two-versus-two comparison tractable at all, and it applies independent filtering and outlier handling by rules that are documented and reproducible.

edgeR quasi-likelihood#

The quasi-likelihood F-test is designed to account for the uncertainty in each gene's dispersion estimate, and the edgeR literature reports error-rate control closer to nominal than a likelihood-ratio test at small sample sizes. That is a property of the method, not a prediction about your gene list: which engine returns more genes on a given dataset depends on the dispersion structure, the effect sizes and the depth, and nothing in this application measures that comparison for you. If your concern is that a list is too permissive, running edgeR as a second engine is more informative than tightening a threshold, because the two methods disagree for reasons you can inspect rather than by construction.

limma-voom#

limma-voom transforms counts and models the mean-variance relationship with observation weights, then uses the linear-model machinery that limma has long applied to arrays. It scales well to larger sample numbers, and a reviewer used to array analysis may ask for it. On very small designs it has less to work with than the count models: the mean-variance trend it fits pools information across all genes, but each gene still has few residual degrees of freedom, so the resulting weights rest on a thin per-gene base.

limma for microarrays#

Microarray intensities are not counts, and no count model applies to them. limma is not an alternative here but the method for the route. The comparison it forms follows the same factor-position rule as the count engines, so a contrast means the same thing across routes even though the underlying model does not.

What switching changes#

Changing the engine is a change of analysis, not of presentation. The gene list, the effect sizes, the standard errors and every downstream output derived from them will differ. If you report results from one engine, report which one; a figure regenerated after an engine change is a different result, not a restyled one.

What the alternative engines will not do#

edgeR, limma-voom and microarray limma fit the documented additive group-means design only. Interaction and nesting operators are rejected by them rather than silently reinterpreted, because a formula that is accepted and then fitted as something simpler is the worst available outcome. DESeq2 is the exception: it accepts a typed interaction design, subject to the contrast-extraction caveat described on Design, contrast and thresholds. If your question requires an interaction term, that is a design question to resolve before choosing an engine. Practical settings for all of this are on Design, contrast and thresholds.

Search the documentation

Type to search every page.

Figure viewer

100%