BulkSeq Studiov0.34.0
Download

How-to guides

Select a reference

Match the organism, assembly, annotation and gene identifiers.

On this page

Before you start, identify the organism and, for read analysis, the intended genome assembly and annotation. A reference determines where reads can map and what the workflow counts as a gene.

1. Choose an organism preset#

Open Reference, select an organism and click Use Selected Preset. The preset supplies genome/annotation resources and downstream organism settings where supported. Imported tables still need the correct organism and identifier mapping even though they do not align reads.

Reference choices and their consequences#

SettingWhat it changesWhen to change it / example
Genome FASTAProvides the sequence against which reads are aligned or a transcriptome is generated.Use the assembly for your experiment; a different assembly can change mappings and counts.
Annotation GTF/GFF3Defines gene/transcript features and identifiers.Use an annotation compatible with the FASTA. Matching species names alone do not prove matching assemblies or chromosome names.
Organism enrichment settingsSelect annotation databases, KEGG organism and STRING taxon.Use the organism represented by your data. Incorrect species settings can produce missing or misleading mappings.
Identifier namespaceDetermines how result rows join annotation records.Symbols, Ensembl IDs and NCBI GeneIDs are not interchangeable strings. Confirm the namespace of imported tables and custom sets.

2. Supply a custom reference when needed#

Use a custom genome and annotation if the catalog does not supply your required assembly. Keep the original source URL, version and integrity information.

Advanced: reuse a prebuilt index

The Reference page does not expose the prebuilt index fields in the published baseline. Set reference.star_index or reference.salmon_index to an existing directory; reference.hisat2_index is the hisat2-build prefix, not an individual index file. An optional reference.transcriptome_fasta supplies Salmon’s index input.

bulkseq config set reference.star_index PATH

A missing configured path is refused rather than silently rebuilt. The genome/annotation check still applies with prebuilt indexes. Salmon still derives its transcript-to-gene map from the annotation; transcript names in an external index must agree with that map. A protein FASTA is not an active substitute.

Troubleshooting: no enrichment mappings#

Check species and identifier namespace before interpreting an empty table biologically. Missing OrgDb support does not automatically disable all enrichment: available KEGG, g:Profiler and STRING routes have their own organism coverage and connectivity requirements. See enrichment and networks.

Search the documentation

Type to search every page.

Figure viewer

100%