A cookbook of runnable commands for common tasks, grouped by what you're
trying to do. All assume the repository root as the working directory; add
--sdm conda (and, inside Docker, --conda-prefix /conda-envs) to any of
these to have Snakemake manage the conda envs itself, as in
Installation.
test: true (or an integer) in the config keeps only a handful of genomes
per source, so a full sample + build run finishes in minutes instead of
hours.
snakemake --configfile config/repdb.yaml --config test=true --sdm conda -j 4The large downloads (GTDB tarball, EukProt archive, virus zip) and the QC/Krona reports still run on the full sources even in test mode.
snakemake --configfile config/repdb.yaml -j 1 --until available_proteomesUseful for browsing what's available (counts by source/domain/completeness) before deciding on a subset; see the explorer notebook in Outputs and quality control.
snakemake --configfile config/repdb.yaml --sdm conda -j <N>Runs both phases (sample then build) against config/repdb.yaml. Use
snakemake sample --configfile config/repdb.yaml --sdm conda -j <N> alone
to stop at the universe and review the selection first.
snakemake build --configfile path/to/custom.yaml --sdm conda -j <N>See Configuration for the full config reference.
snakemake build builds every database your config defines: repdb if
dbs.build.repdb is set, plus every entry under dbs.build.custom. Point at
a config that defines only a custom entry (no repdb: block) to build just
that:
snakemake build --configfile config/example.yaml --sdm conda -j <N>See Custom databases for the config schema.
snakemake --configfile config/repdb.yaml --workflow-profile workflow/profiles/slurm -j <N><N> is the max number of jobs submitted in parallel, not cores per job (set
per rule in the profile). See Choosing an executor for the
other available profiles.
snakemake build --configfile resources/releases/v1/config.yaml --sdm conda -j <N>Skips taxonomy harmonization entirely and re-derives the selection deterministically from a frozen universe. See Reproducing a release and Building & releasing RepDB.
Rscript workflow/scripts/check_custom_proteomes.R resources/custom_genomes_repdb.csvCatches schema and lineage-conflict errors up front rather than mid-run. See Custom databases for the full flag set (checking against a specific reference, verifying FASTA paths exist).
python workflow/scripts/get_fasta.py <id_file> <output_fasta> results/dbs/<db>/<db>_mmseqsSee Utilities and benchmarking.
Rscript -e 'rmarkdown::render("workflow/notebooks/harmonization_report.Rmd")'Useful for re-rendering after tweaking the notebook without re-running the whole pipeline. See Outputs and quality control for what it covers and for the explorer notebook's equivalent command.
Set keep_intermediates: false in your config, then:
snakemake cleanup --configfile <your-config>.yamlDeletes the big non-final intermediates (raw and full-decontaminated fastas,
clustering scratch, decontamination scratch) once every database's final
indices exist. Never runs as part of all/build, and is a safe no-op if
keep_intermediates is left at its default (true).
Needs its own config, config/benchmark.yaml, edited first to point at
reference databases and a query set on your own filesystem (paths in the
shipped file are institution-specific placeholders):
snakemake -s workflow/rules/benchmark.smk --configfile config/benchmark.yaml --sdm conda -j <N>