Skip to content

Repository files navigation

vdjtools

vdjtools — immune-repertoire analysis

PyPI CI docs python license

Get started · Guides · Commands · API · Examples · Glossary

Analysis of T- and B-cell receptor repertoire sequencing: read any format, measure diversity and overlap, correct batch effects, score generation probability against a native V(D)J model, and reduce a whole repertoire to one comparable feature vector.

vdjtools reads the output of every common repertoire pipeline into one canonical AIRR table backed by polars, and every analysis takes and returns that same table — so results chain together instead of needing a converter between each step.

A clean-room Python + C++ rewrite of the Groovy/Java vdjtools, built on arda (germline reference and annotation), seqtree (fuzzy search and e-values) and vdjmatch (overlap and TCRnet).

Install

pip install vdjtools

That is the whole installation — every capability below works from it, with no extras to opt into. Wheels are prebuilt for CPython 3.10–3.13 on Linux, macOS (Apple Silicon) and Windows and include the native C++ extension; a source install compiles it. The three engines above are base dependencies, imported lazily so import vdjtools stays light.

Two optional extras exist because each has a working alternative:

pip install "vdjtools[overlap]"   # scikit-learn, for cluster_samples(method="mds")
pip install "vdjtools[sc]"        # single-cell bridges: anndata, awkward, mudata, pyyaml

Quickstart

from vdjtools import io as vio, stats, overlap

sample = vio.read("clones.tsv")          # MiXcr, immunoSEQ, AIRR, Parquet, native — detected
stats.diversity_stats(sample)            # richness, Chao1, Shannon, inverse Simpson, d50
stats.segment_usage(sample, "v")         # V-gene usage

cohort = vio.read_samples(vio.read_metadata("metadata.txt"), base_dir="samples/")
overlap.overlap_matrix(cohort)           # pairwise repertoire overlap

The same thing with no Python:

vdjtools diversity     sampleA.tsv sampleB.tsv -o diversity.tsv
vdjtools segment-usage *.tsv --segment v -o usage.tsv
vdjtools overlap       *.tsv -o overlap.tsv

Every command writes to -o — TSV, or Parquet when the path ends in .parquet / .pq — or to stdout, so commands pipe. Analysis commands also take a metadata table (-m plus --base-dir) and parallelise over samples with -t/--threads.

Getting started walks through this in about ten minutes, with no data to download.

Where to start

You want to Go to
Get a first result from a file you already have Getting started
Clean, filter, downsample or batch-correct a cohort Pre-processing
Measure diversity, overlap, gene usage, spectratype Guides
Score Pgen, or fit a model to your own reads Model workshop
Reduce each sample to one feature vector for a classifier Signatures
Work with 10x data, or hand off to scirpy / dandelion Single cell
Look up a command or a function Commands · API
Check what a term means Glossary

What is in the box

Module What it does
vdjtools.io Readers for native vdjtools, AIRR Rearrangement TSV and Parquet; format-detecting converters for MiXcr, MiGec, Adaptive immunoSEQ, IMGT/HighV-QUEST, Vidjil, RTCR, TRUST4 and arda; metadata-driven batches and hive-partitioned cohorts
vdjtools.stats Diversity (richness, Chao1, Efron-Thisted, Shannon, Simpson, d50), coverage-based Hill-number rarefaction with bootstrap intervals, spectratype, V/D/J/C usage
vdjtools.model Native V(D)J recombination model: Pgen over nucleotides and amino acids, the Hamming-1 ball, sequence generation, EM inference, tandem-D support, and a 13-command model workshop
vdjtools.overlap Exact and similarity-aware repertoire overlap, TCRnet and ALICE neighbourhood enrichment, sample clustering
vdjtools.preprocess Format conversion, the three filtering axes, frequency handling, downsampling, error correction, V/J-usage batch correction, pooling and joining
vdjtools.features Junction physicochemical profiles and k-mer summaries
vdjtools.biomarker Incidence association against binary, per-HLA-allele or stratified conditions; co-occurrence pairing; metaclonotypes
vdjtools.dynamics Longitudinal clonotype tracking: paired within-donor expansion testing, the VDJtrack recapture model, an edgeR NB-exact caller
vdjtools.signature One fixed, named, already-standardised feature vector per repertoire, rotated through a published corpus
vdjtools.sc Single-cell ingestion, chain pairing and QC, paired Pgen, round-trip interop with scirpy, dandelion and scRepertoire

Recombination models, built in

Models for all seven human loci ship in the wheel — no OLGA install and nothing to download:

from vdjtools.model import load_bundled, native

model = load_bundled("TRB", source="olga")            # or source="learned"
native.pgen_aa(model, "CASSLAPGATNEKLFF")             # 3.595e-08
native.pgen_aa(model, "CASSLAPGATNEKLFF", mismatches=1)   # and its whole Hamming-1 ball

Pgen agrees with OLGA to 1e-15 on every locus while being several times faster, and adds tandem-D support that OLGA and IGoR lack. You can also fit a model to your own non-productive reads, audit it against its germline, compare two models parameter by parameter, and render its recombination Bayes net — see the model workshop.

Signatures

One repertoire in, one fixed-width named feature vector out, standardised against a published reference corpus so your matrix and a collaborator's are in the same coordinate system without either of you fitting a scaler:

vdjtools signature --corpus blood samples/*.tsv.gz -o vsig.tsv
mir       signature --corpus blood samples/*.tsv.gz -o rsig.tsv   # the geometry half, from mirpy

Nine corpora are published, all at 256 components per locus. blood (11,117 real samples, 947 study groups), tissue (21,131 / 1,934) and deep-tcr (3,936 amplicon samples, TRA+TRB) are fitted on real repertoires; the synthetic ones need no cohort at all and rebuild byte for byte anywhere. Artifacts are fetched and digest-verified on first use rather than bundled.

--corpus is required: a signature is comparable to another one only if both were rotated through the same corpus, and nothing about the numbers would say otherwise. Signatures · How they were fitted · Channels

Performance

The Pgen, generation, EM and diversity hot paths are a native C++ core reached through pybind11; everything else is polars. Single thread, Apple M3, bundled human TRB model:

Operation Throughput vs OLGA
Nucleotide Pgen, single-D VDJ 0.5 ms/seq 9×
Amino-acid Pgen 0.6–0.9 ms/seq 8.6×
Amino-acid Pgen over the Hamming-1 ball 15 ms/seq 8.7×
Sequence generation, raw draws 32,100 seq/s —
Sequence generation, productive_only=True 8,960 seq/s —

The productive filter costs 3.6× on TRB and 5.1× on IGH (19,900 raw draws/s against 3,930 productive), because out-of-frame and stop-codon draws are discarded and redrawn. Batched Pgen parallelises over sequences (11× on 16 cores, and identical to the serial result); the EM E-step parallelises over reads (6.7× on 8 threads). Memory: 63 MB resident for import vdjtools plus one model, 123 MB with all seven loaded.

Reproduce the in-repo half with RUN_BENCHMARK=1 pytest tests/python -k benchmark.

Development

Uses uv — one repo-local .venv, no conda:

uv venv && source .venv/bin/activate
uv pip install -e ".[dev,test]"       # builds the _core C++ extension
pytest tests/python -q

bash setup.sh --dev-parents --tests does the same and editable-installs sibling checkouts of seqtree, arda and vdjmatch if present. You need a C++ toolchain (Xcode Command Line Tools on macOS, build-essential on Linux). MMseqs2 is needed only for arda's annotation path.

Citing

For the v2 rewrite, cite this repository. For the original tool, cite Shugay et al., PLoS Computational Biology 2015. The VDJtrack recapture model in vdjtools.dynamics is Pavlova, Zvyagin and Shugay (2024).

License

GPL-3.0-or-later. The legacy Groovy/Java v1.x tool lives on the legacy-1.x branch, with its releases under the repository tags v0.0.1 … 1.2.1.

Releases

Used by

Contributors

Languages