vdjtools#

Immune-repertoire analysis for T- and B-cell receptor sequencing: read any format, measure diversity and overlap, correct batch effects, score generation probability against a native V(D)J model, and reduce a whole repertoire to one comparable feature vector.

vdjtools is a Python library and a command-line tool. It reads the output of every common repertoire pipeline into one canonical AIRR table backed by polars, and every analysis takes and returns that same table – so results chain together instead of needing a converter between each step.

Get started

Install, load a real sample, and get your first diversity and overlap numbers. Start here if you have not used vdjtools before.

Getting started
Guides

One page per task: loading data, pre-processing, diversity, overlap, dynamics, biomarkers, single cell, and the recombination model workshop.

User guide
Command reference

Every command and every flag, with the defaults. No Python needed.

Command reference
Python API

Every module, class and function, with signatures and types.

API reference
Worked examples

Fifteen runnable notebooks, several reproducing a published result — repertoire ageing, CMV and HLA association, vaccination time courses, single-cell pairing.

Worked examples
Glossary

Every term this documentation uses, and which convention vdjtools follows where the field disagrees — junction against CDR3, productive against functional.

Glossary

Install#

pip install vdjtools

That is the whole installation. Wheels are prebuilt for CPython 3.10-3.13 on Linux, macOS (Apple Silicon) and Windows, with the native C++ extension included; a source install compiles it.

Everything the documentation describes works from that one command – there are no extras to opt into for core functionality. The three engines vdjtools delegates to are base dependencies, imported lazily so that import vdjtools stays light:

Engine

Powers

arda

the germline reference (V/D/J segments and CDR3 anchors), the model engine, annotation

seqtree

fuzzy search and e-values: error correction, similarity overlap, TCRnet

vdjmatch

sample overlap, TCRnet, metaclonotypes

Two optional extras exist because each has a working alternative:

pip install "vdjtools[overlap]"   # scikit-learn, for cluster_samples(method="mds")
pip install "vdjtools[sc]"        # single-cell bridges: anndata, awkward, mudata, pyyaml

MMseqs2 is needed only for arda’s alignment and annotation path, never for germline lookup, Pgen, generation or the analytics.

Your first result#

from vdjtools import io as vio, stats, overlap

sample = vio.read("clones.tsv")          # MiXcr, immunoSEQ, AIRR, Parquet -- detected
stats.diversity_stats(sample)            # richness, Chao1, Shannon, Simpson, d50
stats.segment_usage(sample, "v")         # V-gene usage

cohort = vio.read_samples(vio.read_metadata("metadata.txt"), base_dir="samples/")
overlap.overlap_matrix(cohort)           # pairwise repertoire overlap
vdjtools diversity     sampleA.tsv sampleB.tsv -o diversity.tsv
vdjtools segment-usage *.tsv --segment v -o usage.tsv
vdjtools overlap       *.tsv -o overlap.tsv

Every command writes to -o – TSV, or Parquet when the path ends in .parquet or .pq – or to stdout, so commands pipe. Getting started walks through this with real data.

Choose your path#

You want to

Go to

Get a first result from a file you already have

Getting started

Clean, filter, downsample or batch-correct a cohort

Data pre-processing

Measure diversity, overlap, gene usage or spectratype

User guide

Score generation probability, or fit a model to your own reads

Recombination model workshop

Reduce each sample to one fixed feature vector for a classifier

Repertoire signatures

Work with 10x or single-cell data, or hand off to scirpy or dandelion

Single-cell

Look up a command, a flag or a function

Command reference | API reference

Understand a term used in these pages

Glossary

What is in the box#

Module

What it does

Guide

vdjtools.io

Readers for native vdjtools, AIRR Rearrangement TSV and Parquet; format-detecting converters for MiXcr, MiGec, Adaptive immunoSEQ, IMGT/HighV-QUEST, Vidjil, RTCR, TRUST4 and arda; metadata-driven batches and hive-partitioned cohorts

Data pre-processing

vdjtools.stats

Diversity (observed richness, Chao1, Efron-Thisted, Shannon, Simpson, d50), coverage-based Hill-number rarefaction with bootstrap intervals, spectratype, V/D/J/C usage

User guide

vdjtools.model

Native V(D)J recombination model: Pgen over nucleotides and amino acids, the Hamming-1 ball, sequence generation, EM inference, tandem-D support, and a full model workshop

Recombination model workshop

vdjtools.overlap

Exact and similarity-aware repertoire overlap, TCRnet neighbourhood enrichment, sample clustering

User guide

vdjtools.preprocess

Format conversion, the three filtering axes, frequency handling, downsampling, error correction, V/J-usage batch correction, pooling and joining

Data pre-processing

vdjtools.features

Junction physicochemical profiles and k-mer summaries

User guide

vdjtools.biomarker

Incidence-based association against binary, per-HLA-allele or stratified conditions; co-occurrence pairing; metaclonotypes

User guide

vdjtools.dynamics

Longitudinal clonotype tracking between timepoints: the paired within-donor expansion test, the VDJtrack recapture model, an edgeR NB-exact caller

User guide

vdjtools.signature

One fixed, named, already-standardised feature vector per repertoire, rotated through a published corpus

Repertoire signatures

vdjtools.sc

Single-cell ingestion, chain pairing and QC, paired Pgen, and round-trip interoperability with scirpy, dandelion and scRepertoire

Single-cell

Performance#

The Pgen, generation, EM and diversity hot paths are a native C++ core reached through pybind11; everything else is polars. Amino-acid Pgen agrees with OLGA to 1e-15 on all seven human loci.

Single thread, Apple M3, bundled human TRB model:

Operation

Throughput

Relative to OLGA

Nucleotide Pgen, single-D VDJ

0.5 ms per sequence

9x

Amino-acid Pgen

0.6-0.9 ms per sequence

8.6x

Amino-acid Pgen over the Hamming-1 ball

15 ms per sequence

8.7x

Sequence generation, raw draws

32,100 per second

not applicable

Sequence generation, productive_only=True

8,960 per second

not applicable

The productive filter costs 3.6x on TRB and 5.1x on IGH (19,900 raw draws per second against 3,930 productive), because out-of-frame and stop-codon draws are discarded and redrawn rather than repaired. Quote whichever of the two your pipeline uses.

Batched Pgen over many junctions parallelises over sequences (pgen_aa_batch(), 11x on 16 cores, and identical to the serial result to the last bit); the EM E-step parallelises over reads (6.7x on 8 threads). Memory stays modest: 63 MB resident for import vdjtools plus one model, 123 MB with all seven bundled models loaded.

Citing#

For the v2 rewrite, cite this repository. For the original tool, cite Shugay et al., PLoS Computational Biology 2015. The VDJtrack recapture model in vdjtools.dynamics is Pavlova, Zvyagin and Shugay (2024).

Licensed GPL-3.0-or-later. The legacy Groovy/Java v1.x tool lives on the legacy-1.x branch, with its releases under the repository tags v0.0.1 through 1.2.1.