BulkSeq Studiov0.34.0
Download

Reference

Enrichment and protein networks

Connect gene-level results to annotations while retaining mapping and service limits.

On this page

Before you start, verify the organism, gene identifiers and selected-gene thresholds. Enrichment summarizes annotated sets; it does not demonstrate that a pathway is activated or that an interaction occurred in your samples.

1. Enable and review enrichment#

Enrichment is enabled by default. Direction-split over-representation analyses use selected genes; GSEA uses a ranked list. GO uses an available OrgDb/clusterProfiler route or g:Profiler fallback, and KEGG requires a supported organism code.

Interpretation settings#

SettingWhat it changesWhen to change it / example
Enrichment on/offControls whether enrichment outputs are computed.Disable when outside the analysis plan or when required services cannot be used.
Custom GMT / annotation tableDefines your own gene sets, alongside built-in routes.Use a curated collection matching the run’s identifiers. A set of symbols cannot directly match unrelated accession IDs.
Background gene listChanges the eligible comparison universe for custom over-representation.Use a justified tested/eligible background; changing it can change enrichment without changing DE results.
GSVA: offProduces descriptive per-sample scores from supplied custom sets.Enable with an expression matrix and suitable sets. The heatmap is not a significance test.
Genes of interestCreates focused expression views/count tables where a matrix exists.Paste IDs that match the run. Low matching coverage must be resolved before interpreting an empty panel.
Advanced: ranked lists and annotation limits

Built-in over-representation uses tested genes with an adjusted p-value as its universe. That universe and its selected foregrounds do not expand when an independently filtered row is added to GSEA. Local GSEA ranks by the model’s signed statistic; imported results rank by confirmed effect. A table without raw-p evidence retains the legacy adjusted-p eligibility because the workflow cannot infer why an adjusted value is missing. On fallback KEGG routes, every eligible source score is reduced once after its final GeneID is known; multiple source aliases can therefore change an existing canonical rank even when no adjusted-p-missing row is added. KEGG ORA and GSEA are audited separately, so an ORA with few selected genes does not automatically invalidate ranked analysis.

KEGG keys genes either on bare NCBI GeneIDs or on locus tags, and which form an organism uses decides whether the gene ids must be bridged before the query. The form is taken from the bundled organism catalogue, which records a measured value for every preset KEGG organism; for an organism outside the catalogue the workflow probes KEGG live, falling back through four endpoints, and classifies the keys by majority. The evidence line in enrichment_summary.txt names the form, whether it came from the catalogue or a live probe, and which probe answered; it carries a fraction of numeric keys only where a live probe measured them, so a catalogued organism records probe=not attempted instead. When the form cannot be established and the gene ids are not bare GeneIDs, check 10 reports review-required naming the organism, rather than leaving an empty KEGG table unexplained; the same escalation names the engine or input route when a GeneID-keyed organism is reached from a route that carries no usable ncbi_geneid values. The bridge itself is derived from the annotation’s db_xref field and is written by DESeq2, edgeR and limma-voom; the microarray and imported-results routes have no annotation to derive it from. Do not assume identical annotation coverage across engines or organisms.

Custom ORA populations#

The configured background and selected gene list are the supplied populations. Only identifiers represented in the supplied gene sets can enter the effective annotated ORA populations, so those denominators may be smaller. The custom-enrichment summary names both, and the report shows GeneRatio as term overlap over effective selected genes and BgRatio as term members over effective annotated background. An unavailable model does not establish a zero-term result. This reporting distinction changes no enrichment input, selection or p-value.

OrgDb-backed routes now retain a directly supplied NCBI Entrez ID during mapping. If an older run failed at that conversion, revalidate its inputs and rerun the affected enrichment; saved output is not retroactively corrected.

Annotation transfer for organisms without a curated package#

Only a few organisms in the catalogue have a Bioconductor annotation package. For the others, most fungi and crops among them, g:Profiler and KEGG cover only part of the genes, and for some reference annotations g:Profiler recognises none of the gene identifiers. For these organisms the workflow also runs annotation-transfer enrichment: over-representation and GSEA against the Gene Ontology, Reactome and InterPro term sets that STRING 12.0 distributes for each of its genomes. STRING assigns these terms by orthology, carrying annotation from studied relatives to proteins that have none of their own (Szklarczyk et al. 2023). The terms are therefore transferred, not curated, and the outputs say so.

The route runs by itself when the organism has a STRING taxon and no annotation package. Its results are written to results/enrichment/transfer/ with a dot plot, a provenance record and check 25; the curated routes and their files are unchanged. The two STRING files are downloaded once into the project with their retrieval date and checksum, and the analysis then runs offline.

SettingWhat it changesWhen to change it / example
Annotation transfer: AutomaticRuns the route for organisms without an annotation package that have a STRING taxon, and for any organism once an eggNOG-mapper or KO file is supplied. On runs it for every organism, Off never.Leave on Automatic. Switch On only to compare transferred with curated GO on an organism that has both.
eggNOG-mapper annotationsAdds GO terms from an eggNOG-mapper v2 .emapper.annotations file, propagated to their ancestors, and KEGG reference pathways from its KO column.Use it for an organism STRING does not cover. Protein ids in the file are matched to genes through the project annotation (Cantalapiedra et al. 2021).
KO assignments (KofamScan or KofamKOALA)Adds KEGG reference pathways from a KofamScan or KofamKOALA table; only hits above the KO-specific threshold count.Pathway links come from the live KEGG REST service, as on every KEGG route. Maps under Organismal Systems, Human Diseases and Drug Development are excluded (Aramaki et al. 2020).

Gene identifiers are matched to STRING proteins through STRING’s alias file, never through its sequence-similarity aliases. An identifier that names more than one protein is excluded, genes that share a protein are collapsed to one ranked value, and a protein whose genes change in opposite directions is left out of every set and of the universe. The universe is the mapped tested genes, those with an adjusted p-value, and each category and direction is corrected separately. Check 25 records the share of tested genes that mapped: below 80% it reads WARNING and below 50% REVIEW_REQUIRED, the bands the curated routes use, and a download or parsing failure is REVIEW_REQUIRED. How many genes carry a term is reported without a threshold, because partial coverage is normal for transferred annotation. The one coverage judgement is a WARNING when the tested genes are less annotated for GO biological process than 0.8 times the organism’s whole proteome, which points to an identifier problem rather than to biology.

How the route was checked#

On the three bundled model-organism datasets, where curated annotation exists, STRING’s transferred GO biological-process annotation matched the Bioconductor packages’ curated annotation with a precision of 0.80 to 0.87 and a recall of 0.77 to 0.82 over gene–term pairs. Every gene set significant by GSEA in both routes changed in the same direction, and the enrichment scores of those shared sets correlated with a Spearman coefficient of 0.91 to 0.97. Which sets reached significance agreed only in part, with Jaccard indices of 0.22 to 0.63 between the two routes, so the routes agree more on direction than on the exact list of significant terms. Permuting the gene labels of the Fusarium spores dataset removed every over-represented term. A KO table built from KEGG’s own Fusarium graminearum assignments reproduced exactly every one of the 152 KEGG pathway gene sets for that organism that contain a tested gene; KEGG lists 153, and the other contains none.

Two bundled datasets exercise the route on organisms with no annotation package: the rice blast fungus Magnaporthe oryzae, wild type against a deletion of the carbon catabolite repressor MoCreA (GEO series GSE153084), and sorghum shoots under sulfur deficiency (GEO series GSE184725). The tutorial explains how to start a bundled project.

2. Explore or rebuild the STRING network#

Open Protein network. Load / refresh network displays the saved run network. Rebuild from STRING contacts STRING and can retrieve a network under a changed confidence threshold. Pointer or keyboard navigation exposes node details; exported GraphML, SIF and cytoscape.js JSON support external editing.

SettingWhat it changesWhen to change it / example
STRING score: 400; range 1–1000Sets the association-confidence cutoff used to build a network.A higher cutoff, such as 700, is more stringent and can remove associations; it is not an expression p-value.
Seed cap: 400 genesLimits genes used to seed the network. Selection splits the budget across up/down directions, ranked by adjusted p-value.Change only with a reason; a capped network is not every significant gene.
STRING version: 12.0 baselineSelects the database version used for associations.Record realized version and query date when comparing runs.
Layout, colour, node size and in-view filtersChanges the displayed saved network.Re-layout does not alter DE statistics; an in-view edge filter differs from rebuilding the query.

3. Account for online services and privacy#

Environment installation and public-read/reference downloads require network access. g:Profiler analysis submits gene identifiers to a remote service. STRING mapping/rebuilding involves remote gene/protein identifier and organism queries or resources; KEGG access retrieves online annotation resources. Review the service and institutional requirements before enabling online interpretation for restricted identifiers.

Local computation is not a promise of offline operation. STRING is contacted by the documented network route; unavailability yields empty outputs with a warning rather than a biological “no interactions” conclusion. Raw reads are not the input required for these gene-list services, but do not infer a blanket no-network/no-data-transfer guarantee.

Search the documentation

Type to search every page.

Figure viewer

100%