Glossary#
Terms as this documentation uses them. Where a word is used inconsistently across the field, the entry says which convention vdjtools follows, because several of these distinctions change results rather than only wording.
- AIRR#
The Adaptive Immune Receptor Repertoire Community’s data schema. vdjtools standardises on AIRR Rearrangement column names, so
v_call,j_call,junction_aa,duplicate_countmean what the standard says they mean. Seevdjtools.io.schema.- ALICE#
Neighbourhood enrichment measured against a V(D)J generation model rather than against a control cohort: a clonotype with more similar neighbours than its generation probability predicts is evidence of antigen-driven expansion. Contrast TCRnet. Command:
vdjtools alice.- allele#
A specific sequence variant of a germline gene, written with a star suffix:
TRBV9*01. Recombination models are keyed by allele. Real repertoire files usually carry gene-level calls instead, which is a common source of silent zeros – see gene-level call.- anchor#
The conserved residues bounding the junction: Cys104 at the V end and Phe118 (or Trp118 for IGH) at the J end. Their positions in each germline sequence are what let a V or J segment be cut to its CDR3-region contribution. vdjtools resolves anchors from arda.
- arcsine transform#
A variance-stabilising transform for a proportion:
arcsin(sqrt(p)), in Anscombe’s form which also uses the number of observations behindp. A binomial proportion’s variance isp(1-p)/m– it depends on the value itself – so an untransformed share drawn from a shallow sample is noisier than the same share from a deep one, and a model cannot tell the difference. After the transform the variance no longer depends onp. Used for the amino-acid composition block.- C-star#
Written
cstar. The sample coverage a repertoire attained, in[0, 1]– the estimated fraction of the underlying population represented by the observed reads. Used as the common level at which Hill number s are read, so that two samples sequenced to different depths can be compared. Always reported in a signature, even when the diversity columns it gates are holes.- CDR3#
The third complementarity-determining region, in the IMGT definition: the hypervariable loop that contacts antigen, excluding the Cys104 and Phe118 anchor residues. Two residues shorter than the junction for the same receptor.
- channel#
A signature column carried in its own units, never rotated and never winsorized: the coverage a sample attained, whether a locus was present, how much of the row was clamped. Channels answer “can I trust this row”, and mixing a number whose job is to qualify a measurement into the measurement itself would destroy both. Distinct from a principal component and from a named block. See Channels.
- clonality#
How unevenly reads are distributed across clonotypes, reported as
1minus normalised Shannon entropy: 0 is perfectly even, 1 is one clone holding everything. Rises with age and with antigen-driven expansion. Depth-sensitive, so compare it only at matched coverage.- clone size#
How many read s support a clonotype – the
duplicate_countcolumn. The distribution of clone sizes across a repertoire is heavily skewed, which is why unnormalised richness is not comparable between samples.- clonotype#
One distinct receptor sequence in a repertoire, and one row of the canonical frame. What counts as distinct is a choice: vdjtools keys on the nucleotide junction plus V and J calls by default, and amino-acid keying is available where the analysis calls for it.
- clr#
Centred log-ratio. The transform for a closed composition – a set of shares that sum to one, such as V-gene usage or isotype composition. Take the log of each share and subtract the mean of those logs. Necessary because the parts of a composition are not free to vary independently: if one goes up another must come down, so treating them as ordinary variables makes the block singular and any correlation between them partly an artefact of the constraint. A
clrblock shipsk-1parts for the same reason.- cohort#
A set of samples analysed together, usually described by a metadata table with one row per sample. Cohort commands either parallelise over sample files or stream a pre-ingested Parquet store.
- corpus#
A fitted reference: a large collection of repertoires from which winsorization bounds, a per-locus rotation, and per-component centre and scale were all estimated in one pass. A signature is only comparable to another one rotated through the same corpus. Nine are published; see Repertoire signatures.
- coverage#
The fraction of the underlying receptor population that the observed sample represents, estimated from the abundance spectrum. The basis for coverage-standardised diversity, and the reason rarefaction beats fixed-depth downsampling. See C-star.
- d50#
A dominance index, in the legacy
getDxxIndexdefinition vdjtools keeps: rank clonotypes by descending count, take the smallest numberkwhose cumulative read fraction reaches 50 percent, and report1 - k/Sobs– the fraction of clonotypes not needed. Higher means more dominated by few clones. A value of 0.98 means 2 percent of clonotypes carry half the reads.- dispersion#
How spread out a repertoire’s receptors are in the embedding, as opposed to where their centre sits. Two donors can have the same average position and very different spread — one concentrated in a few sequence neighbourhoods, one scattered. Reported by the
rsig:divandrsig:dispfamilies, and unmeasurable below a few clonotypes, which is when the whole family becomes a hole.- downsampling#
Randomly reducing a sample to a fixed number of reads or clonotypes so that depth-sensitive statistics can be compared. Discards data; rarefaction estimates the same quantity without doing so.
- effective dimension#
How many directions a repertoire’s receptors genuinely spread across, as a continuous number: the participation ratio of the embedding’s variance spectrum. A repertoire filling 40 directions evenly and one dominated by 3 have very different effective dimensions even at identical total spread.
- expanded clone#
A clonotype supported by more than one read.
(reads - singletons) / (richness - singletons)– reads per expanded clone – is at least 2 by construction, which is why it, and not reads per clonotype, can be drawn independently of the singleton fraction.- functional gene#
A property of the germline gene, in IMGT’s classification: F (functional), ORF (open reading frame) or P (pseudogene). Orthogonal to productive, which is why vdjtools keeps the two filters on separate flags.
- gene#
A germline V, D, J or C segment, named without an allele suffix:
TRBV9. See allele.- gene-level call#
A
v_callorj_callnaming a gene rather than an allele. Recombination models are keyed by allele, so passing a gene-level name to a Pgen function raises rather than silently marginalising over every allele – an earlier version returned a V/J-agnostic value 2.38 times too high with no error. Resolve to a representative allele explicitly.- generation model#
A probabilistic model of V(D)J recombination: which segments pair, how many nucleotides are deleted from each, and what is inserted between them. Yields Pgen. Models for all seven human loci ship in the wheel. See Recombination model workshop.
- germline#
The unrearranged V, D, J and C segment sequences a receptor is assembled from. vdjtools resolves all germline through arda, which is the single source of truth: mixing germline sources within one model produces wrong answers that look plausible.
- Hill number#
A one-parameter family of diversity measures indexed by order
q, all in units of an effective number of clonotypes.q = 0is richness,q = 1is the exponential of Shannon entropy,q = 2is inverse Simpson. Higherqweights abundant clones more.- holes#
A column that could not be computed, represented as
nanand never as0. “This locus was not sequenced”, “this sample is too shallow to estimate this”, and “this field was absent from the input” are different facts, and themaskchannels record which applies. See Channels.- incidence#
Whether a clonotype is present in a sample at all, as opposed to how abundant it is. Incidence-based association tests ask whether presence co-varies with a phenotype across donors. See
vdjtools.biomarker.- isotype#
The antibody class a B-cell receptor carries — IgM, IgD, IgG, IgA, IgE — read from the
c_callcolumn. Class-switching is a record of B-cell history, so isotype composition is a feature in its own right and is reported for IGH only. Transformed with clr, because the classes are shares of one whole.- junction#
The nucleotide or amino-acid sequence spanning the V-D-J join including the Cys104 and Phe118 anchor residues – AIRR’s
junction_ntandjunction_aa. Two residues longer than the CDR3. vdjtools uses junction columns throughout; confirm which convention any external dataset uses before matching against it. Note thatpgen_aa()andinfer_nt()name their argumentcdr3_aafor historical reasons but require a junction — anchors included.- locus#
Which receptor chain a clonotype belongs to: TRA, TRB, TRG, TRD, IGH, IGK, IGL. Note that TRA and TRD share V segments, so a locus must never be re-derived per row from a gene name in a frame that is already keyed by locus.
- MAD#
Median absolute deviation, scaled by 1.4826 to match the standard deviation of a Gaussian. A robust scale estimate: a signature’s per-component scaling uses the median and MAD of the corpus’s component scores rather than mean and standard deviation, so a few extreme samples cannot set the scale.
- metaclonotype#
A group of similar clonotypes treated as one unit – typically a ball of a given edit distance around a CDR3, optionally requiring matching V and J. Recovers signal that exact matching misses, because two donors rarely share a receptor exactly but often share a motif.
- N region#
The non-templated nucleotides inserted between V and D, and between D and J, during recombination. Together with the deletions at each segment end, this is what makes the junction hypervariable and what a generation model has to model.
- n_eff#
Effective sample size: how many independent observations a weighted sample is worth, computed from the clone-size weights. A repertoire of 10,000 reads dominated by one clone is worth far fewer than 10,000 independent draws, and every estimator’s precision scales with
n_effrather than with the raw read count.- named block#
A raw feature group emitted under its own name, beside the rotated components — the diversity, depth, clone-size, length, isotype and cross-locus numbers the rotation is fitted on. Requested with
--named; off by default. The values carry their declared transform, not a natural scale: alog10diversity comes back as alog10, sovsig:div:TRB:1D_c = 2.386is not a clone count. Which groups are reportable is declared by the library, not inferred from how many columns they have.- Pgen#
The probability that a V(D)J recombination event produces a given receptor sequence, marginalised over every recombination scenario consistent with it. Low Pgen means a sequence unlikely to arise twice independently, so convergent observations of it are informative. Available over nucleotides, amino acids, and the Hamming-1 ball.
- principal component#
One axis of the rotation: a weighted combination of raw features, chosen so that the axes are uncorrelated across the corpus and ordered by how much variation each one carries. The
:pc:columns of a signature. A component is a direction, not a quantity —PC01has no units and no interpretation as a measurement, which is why the raw numbers behind it are available separately as a named block.- productive#
A property of the rearrangement: in frame, with no stop codon, so it can encode a chain. Distinct from functional gene, and filtering on one says nothing about the other – a productive rearrangement can use a pseudogene V segment. See Data pre-processing.
- public clonotype#
A receptor observed in many unrelated donors. Usually a consequence of high Pgen rather than of shared antigen exposure, which is why association testing conditions on generation probability.
- Rao quadratic entropy#
A diversity measure that accounts for how different the clonotypes are, not just how many there are: the average dissimilarity between two clonotypes drawn at random in proportion to their abundance. A repertoire of 100 near-identical sequences and one of 100 unrelated ones have the same richness and very different Rao entropy.
- rarefaction#
Estimating what a diversity statistic would have been at a smaller sample size or a lower coverage, from the abundance spectrum, without discarding reads. The basis of coverage-standardised comparison. Extrapolation is the same machinery run the other way. See
vdjtools.stats.inext().- read#
One sequencing observation supporting a clonotype, counted in
duplicate_count. Where the protocol uses molecular barcodes the unit is a UMI rather than a raw read, which matters for depth-dependent statistics because UMI counts are not inflated by PCR.- recombination#
The somatic process assembling a receptor gene: one V, optionally one or two D, and one J segment are joined, with nucleotides deleted from each segment end and non-templated nucleotides inserted between them. See N region, generation model.
- repertoire#
The full set of receptor sequences carried by one individual, in one sample, at one time. Represented as a table of clonotype s with clone size s.
- richness#
The number of distinct clonotypes observed. Strongly depth-dependent, so raw richness is not comparable across samples; Chao1 and Efron-Thisted estimate what was missed, and rarefaction puts two samples on a common footing.
- robust z-score#
The units a signature column is on: the value minus the corpus’s median for that column, divided by its MAD. Median and MAD rather than mean and standard deviation, so that a handful of extreme corpus samples cannot set the scale. A value of 2 means “two robust deviations above the corpus’s typical sample”.
- rotation#
The per-locus linear map from raw features to principal components inside a corpus artifact. A rotation is a function of features and components only: it carries no sample data and no cohort labels. Truncating it to fewer components is exact, so a wide artifact serves every narrower request.
- SHM#
Somatic hypermutation: the antigen-driven mutation of an already-rearranged B-cell receptor during affinity maturation. Summarised as mean V-segment identity to germline, so a lower value means a more mutated repertoire. IGH only.
- signature#
One fixed-width, named, already-standardised feature vector per repertoire. Split across two packages:
vsigis the statistics half (vdjtools) andrsigthe geometry half (mirpy); they join onsample_id. See Repertoire signatures.- singleton#
A clonotype supported by exactly one read. The singleton fraction drives every coverage and richness estimator, and in a synthetic draw it stands in for the naive share of the repertoire.
- spectratype#
The distribution of junction lengths across a repertoire, optionally resolved per V gene. Skew away from the germline-determined shape indicates selection or clonal expansion.
- support#
The declared range a raw feature can take — non-negative, non-positive, unbounded, or a fixed interval. Declared per feature and used to decide which side winsorization clamps. It is deliberately not derived from the feature’s transform: three features declare no transform and three different supports, so deriving one from the other trims the wrong end.
- tandem D#
A rearrangement incorporating two D segments. Real but rare, and neither OLGA nor IGoR models it; vdjtools supports it exactly. Costs roughly 2.5 times a single-D evaluation, and no exact skip exists, because essentially no read has zero tandem-D contribution.
- TCRnet#
Neighbourhood enrichment measured against a control repertoire: a clonotype with more similar neighbours than the control predicts is a candidate for antigen-driven expansion. Contrast ALICE, which uses a generation model instead. Command:
vdjtools tcrnet.- winsorization#
Clamping a value to a percentile bound estimated on the corpus, so that one extreme sample cannot dominate a rotation. Which side is clamped follows the feature’s declared support, not its transform: a non-negative count is clamped at the top only, a log-probability at the bottom only, an unbounded score at both. Clamping edits a real measurement, so the fraction clamped is always reported in
qc:-:winsor_frac.- Zipf law#
The rank-abundance law used to give a synthetic repertoire a realistic clone-size distribution: the
i-th most abundant clone gets frequency proportional toito the power-a. Note this is the rank law, not numpy’sGenerator.zipf, which samples integers from a Zipf distribution and ata = 1.5has infinite mean — normalising such a draw collapses a repertoire to a handful of clonotypes.