BulkSeq Studiov0.34.0
Download

How-to guides

Read processing and analysis settings

Know which controls affect measurements, models and optional outputs.

On this page

Analysis settings contains Comparison, Read processing, Output options and Advanced tabs. Inactive controls follow the input route: a count table does not undergo trimming, and imported results do not fit a local differential-expression engine.

1. Choose read processing#

For FASTQ/public-read routes, keep the documented fastp trimming default unless your protocol requires another supported trimmer or the input is already trimmed. Raw and post-trimming FastQC reports feed MultiQC.

Core controls#

SettingWhat it changesWhen to change it / example
Trimming: on; fastp defaultRemoves adapter/low-quality sequence before alignment. Trim Galore and Trimmomatic are alternatives.Disable only when intentionally using already-trimmed reads. Keep the upstream trimming record.
rRNA filtering: off; SortMeRNA selectedRemoves reads classified as ribosomal RNA before alignment when enabled.Enable to implement a planned rRNA-removal step. SortMeRNA uses a database; RiboDetector is a CPU classifier without that reference database.
FastQ Screen: offReports matches to an existing genome panel; it is a screen, not a contamination-removal filter.Enable with a configured panel to inspect unexpected sources. No genome panel is downloaded automatically.
Aligner: STARSTAR and HISAT2 provide genome BAMs; Salmon performs transcriptome quantification without BAMs.Consider HISAT2 or Salmon when resource constraints and the analysis goal warrant them. Changing the route can change estimated counts.
Quantifier: featureCountsCounts annotated features from genome alignments. STAR gene-counts is an alternative with STAR; HISAT2 uses featureCounts and Salmon uses tximport.Use STAR gene-counts only when you intentionally choose STAR’s counting definition.
RSeQC: offAdds read-distribution and gene-body coverage diagnostics from BAMs.Enable on STAR/HISAT2 when you need those views; it cannot run on Salmon output.
Organellar genes: keepKeep includes them; discard removes them before DE; separate runs nuclear DE and writes organellar counts/fractions.For a planned nuclear-only comparison, choose discard or separate and record the choice. Separate does not promise a separate organellar DE fit.

2. Understand advanced read parameters#

Advanced: trimming settings with examples
SettingWhat it changesWhen to change it / example
fastp quality Phred 15; low-quality limit 40%Controls the quality classification and tolerated fraction of low-quality bases.Raise stringency only with a protocol/QC reason; stricter filtering can lose usable reads as well as poor reads.
Minimum length 36Rejects reads shorter than the retained-length requirement.Illustration: a 30-base read fails 36. Raising the limit to 50 also removes a 45-base read.
Poly-G off; poly-X offEnables terminal homopolymer trimming.Use when library/instrument characteristics justify it, not to make a QC panel appear cleaner.
Trimmomatic window 4, quality 15; leading/trailing 3Sets sliding-window and end-quality trimming.A stricter window criterion can shorten more reads; preserve the reason for the change.
RiboDetector ensure=norrna; chunk size 256 × 1024 readsSets classification retention mode and processing batch size.Change chunk size for memory constraints; it is not a biological significance threshold.
FastQ Screen subset 100000; panel config requiredControls how many reads are sampled and which genomes are screened.A larger screen samples more reads but costs time; an absent organism in the panel cannot be detected by that panel.
Advanced: alignment and counting settings
SettingWhat it changesWhen to change it / example
STAR two-pass offControls a second alignment pass informed by detected splice junctions.Enable for an intentional alignment protocol, then treat outputs as a changed analysis.
STAR maximum multimappers 10; mismatch/read-length ratio 1.0Constrains reported mapping multiplicity and mismatch acceptance.Tighter constraints may reduce assigned reads; do not tune solely to increase a downstream hit count.
featureCounts feature exon; attribute gene_idDefines the annotation features and grouping key counted as genes.A custom annotation must actually contain the chosen fields. Changing exon to another feature changes the quantified object.

Strandedness is inferred and applied per sample in the published baseline. Inspect results/aligned/strandedness_per_sample.tsv; a configured generic code is not evidence that every library has that orientation.

3. Save the settings#

Click Save analysis settings. Complete Comparison, then choose resources and run pre-run checks.

Search the documentation

Type to search every page.

Figure viewer

100%