Before you start, identify the organism and, for read analysis, the intended genome assembly and annotation. A reference determines where reads can map and what the workflow counts as a gene.
1. Choose an organism preset#
Open Reference, select an organism and click Use Selected Preset. The preset supplies genome/annotation resources and downstream organism settings where supported. Imported tables still need the correct organism and identifier mapping even though they do not align reads.
Reference choices and their consequences#
| Setting | What it changes | When to change it / example |
|---|---|---|
| Genome FASTA | Provides the sequence against which reads are aligned or a transcriptome is generated. | Use the assembly for your experiment; a different assembly can change mappings and counts. |
| Annotation GTF/GFF3 | Defines gene/transcript features and identifiers. | Use an annotation compatible with the FASTA. Matching species names alone do not prove matching assemblies or chromosome names. |
| Organism enrichment settings | Select annotation databases, KEGG organism and STRING taxon. | Use the organism represented by your data. Incorrect species settings can produce missing or misleading mappings. |
| Identifier namespace | Determines how result rows join annotation records. | Symbols, Ensembl IDs and NCBI GeneIDs are not interchangeable strings. Confirm the namespace of imported tables and custom sets. |
2. Supply a custom reference when needed#
Use a custom genome and annotation if the catalog does not supply your required assembly. Keep the original source URL, version and integrity information.
Advanced: reuse a prebuilt index
The Reference page does not expose the prebuilt index fields in the published baseline. Set reference.star_index or reference.salmon_index to an existing directory; reference.hisat2_index is the hisat2-build prefix, not an individual index file. An optional reference.transcriptome_fasta supplies Salmon’s index input.
bulkseq config set reference.star_index PATHA missing configured path is refused rather than silently rebuilt. The genome/annotation check still applies with prebuilt indexes. Salmon still derives its transcript-to-gene map from the annotation; transcript names in an external index must agree with that map. A protein FASTA is not an active substitute.
Troubleshooting: no enrichment mappings#
Check species and identifier namespace before interpreting an empty table biologically. Missing OrgDb support does not automatically disable all enrichment: available KEGG, g:Profiler and STRING routes have their own organism coverage and connectivity requirements. See enrichment and networks.
Terms on this page: Gene identifier.