Before you start, prepare a sample inventory from your laboratory records. Each biological replicate needs a stable identifier, a condition and the correct input files or matrix column.
1. Create or open a project#
On Project, choose a new local project folder. The application writes configuration under config/, copies its workflow and records a provenance manifest. The name must identify an immediate child of the selected working directory: `.` and `..`, and links that resolve outside it, are refused. An existing non-empty directory requires confirmation before it is scaffolded; an existing file is refused. On Windows, Use WSL filesystem places the project on the Linux filesystem; ordinary network shares are refused.
2. Review the Samples table#
Open Samples after adding data. Inspect sample names, condition labels and read layout. Expand More table tools to add rows or columns, or use Import table… for TSV, CSV or XLSX metadata. Match matrix columns exactly; do not assume table order identifies the samples.
Columns that determine the analysis#
| Setting | What it changes | When to change it / example |
|---|---|---|
| condition | Defines the groups used by the default comparison. | Illustration: use treated and control consistently, not a mixture of Control, control and untreated for the same group. |
| dataset | Records the study of origin for a requested cross-study analysis; the name alone does not establish independence. | Use study_A and study_B for distinct source studies, not sequencing lanes from one sample. |
| batch / donor / other covariates | Can enter the model to account for recorded sample structure. | Use b1 and b2 for categorical batches. Purely numeric 1 and 2 are fitted as a continuous trend. |
| File paths / sample identifiers | Connects observations to their metadata. | A renamed matrix column must still match its intended sample. Do not move an input after validation without rechecking. |
Naming a library#
The sheet may also carry an optional library_name column directly after sample_id: the submitter's own label for the sequencing library. An import by accession fills it from the archive's record where one exists and leaves it empty otherwise; you can type or change it like any other column. It is descriptive, so it may be left blank and two rows may share a name.
sample_id stays the identifier the analysis runs on. File names, matrix columns, results tables and every match between a sample and its data use it, and nothing is derived from the library name, which is why duplicates are safe. Where a human-readable name helps, the run summary, the study design and the report show it beside the sample identifier rather than instead of it, and per-sample figure labels use it, adding the sample identifier only for rows whose name repeats so that no two points on a plot carry the same label. Because it names a sample rather than describing a condition, it is not offered as a design covariate and the covariate screen does not test it. A project whose sheet has no such column is unaffected, and its sheet is not rewritten to add an empty one.
3. Save the sheet#
Click Save samples.tsv, then continue to the design and contrast. Changes in group membership or covariates can change the fitted comparison; they are analysis changes, not cosmetic labels.
Advanced: preserving a project
Retain config/, logs/, checks/ and the applicable result files together. The active workflow may be synchronized when a project made by an older application runs again; record the software version and preserve the original project before an update intended to reproduce an older result.
Advanced: start from a bundled example study
Create Benchmark Project on the Project page configures a project for one of the bundled public studies, as a worked example or an end-to-end check of an installation. Read-based examples download their reads on the first run, so they need network access and disk space; the microarray examples fetch processed intensities instead.
| Bundled study | Route and samples | Comparison |
|---|---|---|
| Pasilla paired-end subset, Drosophila melanogaster (GSE18508) | Public paired-end reads, 4 samples | CG8144 RNAi against untreated |
| Yeast rpd3-delta Ume6 truncation subset, Saccharomyces cerevisiae (PRJNA630199) | Public paired-end reads, 4 samples | Ume6 delta2-508 against the rpd3-delta control |
| Rice CY1000 salt stress, Oryza sativa Japonica Group (PRJDB38133) | Public paired-end reads, 6 samples | 5-day NaCl stress against control |
| Arabidopsis hub2-3 mutant, ATH1 array (GSE30735) | Microarray intensities, 6 samples | hub2-3 against Col-0 wild type |
| Yeast cbc2-delta, YG-S98 array (GSE6705) | Microarray intensities, 6 samples | cbc2-delta against wild type |
| Fusarium graminearum PH-1 spores against mycelium (GSE55477, PRJNA239711, SRP039087) | Public paired-end reads, 6 samples, 90 bp mates, unstranded | spores against mycelium |
| Fusarium graminearum Z-3639 heat shock (GSE78885, PRJNA314297, SRP071140) | Public paired-end reads, 6 samples, 151 bp mates, reverse-stranded | 37 °C against 25 °C |
The two Fusarium studies were added in 0.31.0. The spores-against-mycelium study is the fungal example used as a guardrail; its dataset publication is Zhao et al. (2014) BMC Genomics 15:191, doi:10.1186/1471-2164-15-191, not the PH-1 genome paper. The heat-shock study sequenced strain Z-3639 and is analysed against the PH-1 reference preset, and its reverse-stranded library exercises a different counting orientation. The differential-expression counts the catalogue records for both studies were first obtained in the 0.2x-era runs and reproduced from FASTQ under the 0.31.0 workflow, each count matrix and DESeq2 table identical to a v0.29.1 run of the same reads.
Terms on this page: Biological replicate · Covariate.