vdjtools#
Immune-repertoire analysis for T- and B-cell receptor sequencing: read any format, measure diversity and overlap, correct batch effects, score generation probability against a native V(D)J model, and reduce a whole repertoire to one comparable feature vector.
vdjtools is a Python library and a command-line tool. It reads the output of every common repertoire pipeline into one canonical AIRR table backed by polars, and every analysis takes and returns that same table – so results chain together instead of needing a converter between each step.
Install, load a real sample, and get your first diversity and overlap numbers. Start here if you have not used vdjtools before.
One page per task: loading data, pre-processing, diversity, overlap, dynamics, biomarkers, single cell, and the recombination model workshop.
Every command and every flag, with the defaults. No Python needed.
Every module, class and function, with signatures and types.
Fifteen runnable notebooks, several reproducing a published result — repertoire ageing, CMV and HLA association, vaccination time courses, single-cell pairing.
Every term this documentation uses, and which convention vdjtools follows where the field disagrees — junction against CDR3, productive against functional.
Install#
pip install vdjtools
That is the whole installation. Wheels are prebuilt for CPython 3.10-3.13 on Linux, macOS (Apple Silicon) and Windows, with the native C++ extension included; a source install compiles it.
Everything the documentation describes works from that one command – there are no extras to opt
into for core functionality. The three engines vdjtools delegates to are base dependencies, imported
lazily so that import vdjtools stays light:
Engine |
Powers |
|---|---|
the germline reference (V/D/J segments and CDR3 anchors), the model engine, annotation |
|
fuzzy search and e-values: error correction, similarity overlap, TCRnet |
|
sample overlap, TCRnet, metaclonotypes |
Two optional extras exist because each has a working alternative:
pip install "vdjtools[overlap]" # scikit-learn, for cluster_samples(method="mds")
pip install "vdjtools[sc]" # single-cell bridges: anndata, awkward, mudata, pyyaml
MMseqs2 is needed only for arda’s alignment and annotation path, never for germline lookup, Pgen, generation or the analytics.
Your first result#
from vdjtools import io as vio, stats, overlap
sample = vio.read("clones.tsv") # MiXcr, immunoSEQ, AIRR, Parquet -- detected
stats.diversity_stats(sample) # richness, Chao1, Shannon, Simpson, d50
stats.segment_usage(sample, "v") # V-gene usage
cohort = vio.read_samples(vio.read_metadata("metadata.txt"), base_dir="samples/")
overlap.overlap_matrix(cohort) # pairwise repertoire overlap
vdjtools diversity sampleA.tsv sampleB.tsv -o diversity.tsv
vdjtools segment-usage *.tsv --segment v -o usage.tsv
vdjtools overlap *.tsv -o overlap.tsv
Every command writes to -o – TSV, or Parquet when the path ends in .parquet or .pq –
or to stdout, so commands pipe. Getting started walks through this with real data.
Choose your path#
You want to |
Go to |
|---|---|
Get a first result from a file you already have |
|
Clean, filter, downsample or batch-correct a cohort |
|
Measure diversity, overlap, gene usage or spectratype |
|
Score generation probability, or fit a model to your own reads |
|
Reduce each sample to one fixed feature vector for a classifier |
|
Work with 10x or single-cell data, or hand off to scirpy or dandelion |
|
Look up a command, a flag or a function |
|
Understand a term used in these pages |
What is in the box#
Module |
What it does |
Guide |
|---|---|---|
|
Readers for native vdjtools, AIRR Rearrangement TSV and Parquet; format-detecting converters for MiXcr, MiGec, Adaptive immunoSEQ, IMGT/HighV-QUEST, Vidjil, RTCR, TRUST4 and arda; metadata-driven batches and hive-partitioned cohorts |
|
|
Diversity (observed richness, Chao1, Efron-Thisted, Shannon, Simpson, d50), coverage-based Hill-number rarefaction with bootstrap intervals, spectratype, V/D/J/C usage |
|
|
Native V(D)J recombination model: Pgen over nucleotides and amino acids, the Hamming-1 ball, sequence generation, EM inference, tandem-D support, and a full model workshop |
|
|
Exact and similarity-aware repertoire overlap, TCRnet neighbourhood enrichment, sample clustering |
|
|
Format conversion, the three filtering axes, frequency handling, downsampling, error correction, V/J-usage batch correction, pooling and joining |
|
|
Junction physicochemical profiles and k-mer summaries |
|
|
Incidence-based association against binary, per-HLA-allele or stratified conditions; co-occurrence pairing; metaclonotypes |
|
|
Longitudinal clonotype tracking between timepoints: the paired within-donor expansion test, the VDJtrack recapture model, an edgeR NB-exact caller |
|
One fixed, named, already-standardised feature vector per repertoire, rotated through a published corpus |
||
|
Single-cell ingestion, chain pairing and QC, paired Pgen, and round-trip interoperability with scirpy, dandelion and scRepertoire |
Performance#
The Pgen, generation, EM and diversity hot paths are a native C++ core reached through pybind11; everything else is polars. Amino-acid Pgen agrees with OLGA to 1e-15 on all seven human loci.
Single thread, Apple M3, bundled human TRB model:
Operation |
Throughput |
Relative to OLGA |
|---|---|---|
Nucleotide Pgen, single-D VDJ |
0.5 ms per sequence |
9x |
Amino-acid Pgen |
0.6-0.9 ms per sequence |
8.6x |
Amino-acid Pgen over the Hamming-1 ball |
15 ms per sequence |
8.7x |
Sequence generation, raw draws |
32,100 per second |
not applicable |
Sequence generation, |
8,960 per second |
not applicable |
The productive filter costs 3.6x on TRB and 5.1x on IGH (19,900 raw draws per second against 3,930 productive), because out-of-frame and stop-codon draws are discarded and redrawn rather than repaired. Quote whichever of the two your pipeline uses.
Batched Pgen over many junctions parallelises over sequences
(pgen_aa_batch(), 11x on 16 cores, and identical to the serial result to the last bit);
the EM E-step parallelises over reads (6.7x on 8 threads). Memory stays modest: 63 MB resident for
import vdjtools plus one model, 123 MB with all seven bundled models loaded.
Citing#
For the v2 rewrite, cite this repository. For the original tool, cite
Shugay et al., PLoS Computational Biology 2015.
The VDJtrack recapture model in vdjtools.dynamics is Pavlova, Zvyagin and Shugay (2024).
Licensed GPL-3.0-or-later. The legacy Groovy/Java v1.x tool lives on the
legacy-1.x branch, with its releases
under the repository tags v0.0.1 through 1.2.1.