Command reference#
pip install vdjtools installs one vdjtools command. This page lists every command, what it
is for, and the options that change an answer rather than only a path.
vdjtools <command> --help is the authoritative list of flags and defaults – it is generated from
the code, so it cannot drift. This page is the map.
Conventions#
These hold for every command, so they are stated once here rather than repeated per command.
Output. -o/--out writes TSV, or Parquet when the path ends in .parquet or .pq. Omit
it and the result goes to stdout while progress and errors go to stderr, so commands pipe cleanly.
Input. Any supported format is detected: native vdjtools, AIRR Rearrangement TSV, Parquet, MiXcr,
MiGec, Adaptive immunoSEQ, IMGT/HighV-QUEST, Vidjil, RTCR, TRUST4, arda. --format overrides the
sniffer when you need it to.
Cohorts. Analysis commands take either a list of sample files or a metadata table with
-m/--metadata plus --base-dir, mirroring the legacy tool’s workflow. --sample-col and
--file-template control how a metadata row becomes a filename.
Parallelism. -t/--threads on analysis commands parallelises over samples; 0 means all
cores. --cohort instead streams one pass over a pre-ingested Parquet store, which is what to use
when the cohort does not fit in memory. On signature, the flag is -j/--jobs and it means
worker processes – the per-sample work is polars and numpy rather than a GIL-releasing kernel,
so processes are the layer that helps.
Models. Wherever a command takes a model, it accepts LOCUS, LOCUS:source,
LOCUS:source:organism, or a path to a model directory – for example TRB, TRB:learned,
TRB:olga:human.
Data and pre-processing#
convert – any format to the canonical table#
vdjtools convert mixcr.txt.gz -o clones.parquet
Read anything supported and write the canonical AIRR-junction table. Converting a cohort to Parquet once is usually the cheapest thing you can do to a repeated analysis.
filter – three separate axes#
vdjtools filter clones.parquet --productive --min-freq 1e-4 -o productive.tsv
vdjtools filter clones.parquet --nonproductive -o nonproductive.tsv
vdjtools filter clones.parquet --functional-genes --keep-orf -o f_orf.tsv
The three filtering axes are deliberately separate flags, because they are separate facts:
Flag |
Asks |
|---|---|
|
Does the rearrangement encode a chain? In frame, no stop codon. |
|
Is the germline gene real? IMGT F, ORF or P; |
|
Is |
A productive rearrangement can use a pseudogene V, so filtering one axis says nothing about the
other. --recompute-frequencies is on by default, renormalising frequency over the survivors;
--keep-frequencies leaves the file’s own values untouched.
--min-freq, --v, --j and --remove add frequency and segment selection.
downsample – normalise depth#
vdjtools downsample clones.parquet 100000 -o ds.tsv
vdjtools downsample clones.parquet 5000 --clones -o ds.tsv # unique clonotypes, not reads
Reduces to a common depth by resampling. --seed makes it reproducible. Prefer
coverage-standardised diversity over downsampling where the question allows it – see
Data pre-processing.
correct-vj – cross-batch usage bias#
vdjtools correct-vj s1.tsv s2.tsv s3.tsv s4.tsv -b A,A,B,B \
--transform sigmoid --usage-out usage.tsv --outdir corrected/
Corrects V/J usage differences between batches and, with --outdir, rewrites the clonotype tables
to match. -b/--batches is required and parallel to the sample list. --transform chooses
location (default) or sigmoid; --scope selects v, j or vj.
Four more options change the answer rather than a path. --winsor-q clamps the per-batch mean and
sigma at a quantile — off by default, matching the published method, with 0.025 a robustness
setting for shallow or RNA-seq repertoires. --unweighted builds usage from distinct clonotypes
instead of reads, which is the right choice when one hyperexpanded clone would otherwise define a
batch’s usage. --z-cap bounds |Z| under --transform sigmoid. --rescale rewrites
counts deterministically instead of resampling them, so the output is reproducible without a seed
but no longer integer-sampled. See Data pre-processing.
pool – combine samples#
vdjtools pool s1.tsv s2.tsv s3.tsv -o pooled.tsv
vdjtools pool s1.tsv s2.tsv s3.tsv --join --min-samples 2 -o joint.tsv
Without --join, sums counts into one deeper repertoire. With it, keeps clonotypes seen in at
least --min-samples samples – an incidence join, which is the right input to public-clonotype
work. --key selects nucleotide or amino-acid identity.
Repertoire analytics#
All four take sample files or -m, and accept -t and --cohort.
Command |
What it computes |
|---|---|
|
Observed richness, Chao1, Chao coverage, Efron-Thisted, Shannon, normalised Shannon, inverse
Simpson, d50 – one row per sample. |
|
Junction-length distribution. |
|
V, D, J, C or VJ usage. |
|
Exact-match pairwise overlap (D, F, F2, R) for every sample pair. |
vdjtools diversity sampleA.tsv sampleB.tsv -o diversity.tsv
vdjtools spectratype --cohort cohort_parquet/ -o spectra.tsv
vdjtools segment-usage -m metadata.txt --base-dir samples/ -t 0 -o usage.tsv
vdjtools overlap *.tsv -o overlap.tsv
Enrichment and longitudinal#
tcrnet and alice – neighbourhood enrichment#
Both ask whether a clonotype has more similar neighbours than expected, and they differ in what supplies the expectation:
vdjtools tcrnet sample.tsv --scope 1,0,0,1 -o tcrnet.tsv # vs a CONTROL REPERTOIRE
vdjtools alice sample.tsv --scope 1,0,0,1 -o alice.tsv # vs a GENERATION MODEL
--scope is the edit-distance ball as substitutions,insertions,deletions,total. alice adds
--q (the significance threshold), --min-degree and --min-count; both take --locus,
--species or --source, and --threads.
dynamics – what changed between two timepoints#
vdjtools dynamics day0.tsv day15.tsv -o tracked.tsv
A paired within-donor test classifying each clonotype as emergent, expanded, persistent, contracted
or vanishing. --neff overrides the effective depth used for the pair, --umi declares that
counts are molecular barcodes, --min-total and --alpha set the testing thresholds.
Signatures#
signature – one row of named features per sample#
vdjtools signature --corpus blood samples/*.tsv.gz -o vsig.tsv
vdjtools signature --corpus blood --components 32 --describe
vdjtools signature --corpus tissue -m metadata.txt --base-dir samples/ -j 8 -o vsig.tsv
--corpus is required and takes a published name or a path; a silently chosen rotation would
make two matrices look comparable when they are not. Key options:
Option |
Meaning |
|---|---|
|
|
|
An integer count ( |
|
|
|
Coverage level the Hill numbers are read at: a number, |
|
Which stored percentile to clamp at, |
|
Also emit the raw blocks the rotation is fitted on, under their own names:
|
|
|
|
Print the columns this invocation emits, and exit. Resolved against |
corpus – fetch or build an artifact#
vdjtools corpus --fetch all # pre-warm the cache
vdjtools corpus --corpus synthetic-blood -o sb.npz # build one yourself
vdjtools corpus --smoke -o /tmp/smoke.npz # minutes, not hours
Building uses no samples from anybody’s cohort: every receptor is drawn from the bundled
recombination models, so the artifact is reproducible by anyone who installs the library, and is
byte-identical across processes and thread counts. --samples, --size, --seed, --loci
and --depth-spread control the draw, --components and --winsor-p control the fit, and all
are recorded in the manifest. See Repertoire signatures.
Recombination models#
pgen, generate, models#
vdjtools models # list the bundled models
vdjtools generate -m TRB -n 1000 -o gen.tsv # cf. olga-generate_sequences
vdjtools pgen seqs.tsv -m TRB -o pgen.tsv # cf. olga-compute_pgen
vdjtools pgen seqs.tsv -m TRB --mismatches 1 # and the Hamming-1 ball
pgen takes --column for the sequence column and --v-col / --j-col to condition on the
observed V and J; --type selects nucleotide or amino acid, auto by default, and
--no-header writes the values alone for piping. generate
takes -n/--number for how many sequences to draw, --productive/--no-productive to keep
only productive draws (on by default, at roughly 3.6 times the cost per kept sequence on TRB), and
--seed for reproducibility. Both commands take -m/--model for a bundled model, or
--model-path for one on disk.
model – the workshop#
Thirteen subcommands for building, fitting, auditing and comparing models. Full walkthrough in Recombination model workshop.
Subcommand |
What it does |
|---|---|
|
List the bundled models, as |
|
Scaffold a model from a germline library – arda’s, or your own FASTA plus anchors. |
|
Fit marginals from your own sequences by EM, writing the training log alongside.
|
|
Build several chains from the full AIRR read corpus: fetch, arda-map, then EM. |
|
Audit a model against its manifest, its germline and a reference library. Exits non-zero on an error, so it belongs in CI. |
|
Show the EM training log, log-likelihood per iteration, one block per run. |
|
Add alleles from a larger germline library, seeded from what the model already knows. |
|
Replace a model’s V/J usage with your own sample’s, keeping its junction model. |
|
Export every probability as tables: a hand-editable directory, or one long frame. |
|
Information content per recombination event: entropy, mutual information, or the total. |
|
Total diversity: scenario entropy, sequence entropy, effective diversity. |
|
Render the recombination Bayes net, nodes annotated with entropy and edges with mutual information. |
|
Compare two models parameter by parameter, or score one sequence set under both and compare the Pgen distributions. |
|
Log-likelihood of a sequence set under a model, with free parameters, AIC and BIC. |
vdjtools model check TRB:learned
vdjtools model template --locus TRB -o tmpl/
vdjtools model learn clones.tsv -t tmpl/ -o fitted/
vdjtools model compare TRB:olga TRB:learned --by gene --dot diff.pdf
vdjtools model loglik seqs.tsv TRB:learned
Single cell#
sc – five subcommands#
Subcommand |
What it does |
|---|---|
|
Read a single-cell contig table into the canonical long frame, or AIRR with |
|
Resolve each cell’s chains and emit one row per paired receptor.
|
|
Chain-multiplicity quadrants: how many cells carry how many light by heavy chains. |
|
Paired generation probability per cell, Pgen(alpha) times Pgen(beta). |
|
Export for scirpy, dandelion, scRepertoire or any AIRR consumer. |
vdjtools sc convert filtered_contig_annotations.csv -o long.tsv
vdjtools sc pair filtered_contig_annotations.csv --flag-mispairing -o paired.tsv
vdjtools sc pgen filtered_contig_annotations.csv -o paired_pgen.tsv
vdjtools sc export filtered_contig_annotations.csv --to scirpy -o adata.h5ad
--locus-pair defaults to TRA_TRB; set it for BCR. Details and the interop matrix are in
Single-cell.