BulkSeq Studiov0.34.0
Download

How-to guides

Project and sample sheet

Connect every data file to the biological sample and group it represents.

On this page

Before you start, prepare a sample inventory from your laboratory records. Each biological replicate needs a stable identifier, a condition and the correct input files or matrix column.

1. Create or open a project#

On Project, choose a new local project folder. The application writes configuration under config/, copies its workflow and records a provenance manifest. The name must identify an immediate child of the selected working directory: `.` and `..`, and links that resolve outside it, are refused. An existing non-empty directory requires confirmation before it is scaffolded; an existing file is refused. On Windows, Use WSL filesystem places the project on the Linux filesystem; ordinary network shares are refused.

2. Review the Samples table#

Open Samples after adding data. Inspect sample names, condition labels and read layout. Expand More table tools to add rows or columns, or use Import table… for TSV, CSV or XLSX metadata. Match matrix columns exactly; do not assume table order identifies the samples.

Columns that determine the analysis#

SettingWhat it changesWhen to change it / example
conditionDefines the groups used by the default comparison.Illustration: use treated and control consistently, not a mixture of Control, control and untreated for the same group.
datasetRecords the study of origin for a requested cross-study analysis; the name alone does not establish independence.Use study_A and study_B for distinct source studies, not sequencing lanes from one sample.
batch / donor / other covariatesCan enter the model to account for recorded sample structure.Use b1 and b2 for categorical batches. Purely numeric 1 and 2 are fitted as a continuous trend.
File paths / sample identifiersConnects observations to their metadata.A renamed matrix column must still match its intended sample. Do not move an input after validation without rechecking.

Naming a library#

The sheet may also carry an optional library_name column directly after sample_id: the submitter's own label for the sequencing library. An import by accession fills it from the archive's record where one exists and leaves it empty otherwise; you can type or change it like any other column. It is descriptive, so it may be left blank and two rows may share a name.

sample_id stays the identifier the analysis runs on. File names, matrix columns, results tables and every match between a sample and its data use it, and nothing is derived from the library name, which is why duplicates are safe. Where a human-readable name helps, the run summary, the study design and the report show it beside the sample identifier rather than instead of it, and per-sample figure labels use it, adding the sample identifier only for rows whose name repeats so that no two points on a plot carry the same label. Because it names a sample rather than describing a condition, it is not offered as a design covariate and the covariate screen does not test it. A project whose sheet has no such column is unaffected, and its sheet is not rewritten to add an empty one.

3. Save the sheet#

Click Save samples.tsv, then continue to the design and contrast. Changes in group membership or covariates can change the fitted comparison; they are analysis changes, not cosmetic labels.

Advanced: preserving a project

Retain config/, logs/, checks/ and the applicable result files together. The active workflow may be synchronized when a project made by an older application runs again; record the software version and preserve the original project before an update intended to reproduce an older result.

Advanced: start from a bundled example study

Create Benchmark Project on the Project page configures a project for one of the bundled public studies, as a worked example or an end-to-end check of an installation. Read-based examples download their reads on the first run, so they need network access and disk space; the microarray examples fetch processed intensities instead.

Bundled studyRoute and samplesComparison
Pasilla paired-end subset, Drosophila melanogaster (GSE18508)Public paired-end reads, 4 samplesCG8144 RNAi against untreated
Yeast rpd3-delta Ume6 truncation subset, Saccharomyces cerevisiae (PRJNA630199)Public paired-end reads, 4 samplesUme6 delta2-508 against the rpd3-delta control
Rice CY1000 salt stress, Oryza sativa Japonica Group (PRJDB38133)Public paired-end reads, 6 samples5-day NaCl stress against control
Arabidopsis hub2-3 mutant, ATH1 array (GSE30735)Microarray intensities, 6 sampleshub2-3 against Col-0 wild type
Yeast cbc2-delta, YG-S98 array (GSE6705)Microarray intensities, 6 samplescbc2-delta against wild type
Fusarium graminearum PH-1 spores against mycelium (GSE55477, PRJNA239711, SRP039087)Public paired-end reads, 6 samples, 90 bp mates, unstrandedspores against mycelium
Fusarium graminearum Z-3639 heat shock (GSE78885, PRJNA314297, SRP071140)Public paired-end reads, 6 samples, 151 bp mates, reverse-stranded37 °C against 25 °C

The two Fusarium studies were added in 0.31.0. The spores-against-mycelium study is the fungal example used as a guardrail; its dataset publication is Zhao et al. (2014) BMC Genomics 15:191, doi:10.1186/1471-2164-15-191, not the PH-1 genome paper. The heat-shock study sequenced strain Z-3639 and is analysed against the PH-1 reference preset, and its reverse-stranded library exercises a different counting orientation. The differential-expression counts the catalogue records for both studies were first obtained in the 0.2x-era runs and reproduced from FASTQ under the 0.31.0 workflow, each count matrix and DESeq2 table identical to a v0.29.1 run of the same reads.

Search the documentation

Type to search every page.

Figure viewer

100%