Glossary ======== Terms as this documentation uses them. Where a word is used inconsistently across the field, the entry says which convention vdjtools follows, because several of these distinctions change results rather than only wording. .. glossary:: :sorted: AIRR The Adaptive Immune Receptor Repertoire Community's data `schema `_. vdjtools standardises on AIRR Rearrangement column names, so ``v_call``, ``j_call``, ``junction_aa``, ``duplicate_count`` mean what the standard says they mean. See :mod:`vdjtools.io.schema`. allele A specific sequence variant of a germline :term:`gene`, written with a star suffix: ``TRBV9*01``. Recombination models are keyed by allele. Real repertoire files usually carry gene-level calls instead, which is a common source of silent zeros -- see :term:`gene-level call`. anchor The conserved residues bounding the :term:`junction`: Cys104 at the V end and Phe118 (or Trp118 for IGH) at the J end. Their positions in each germline sequence are what let a V or J segment be cut to its CDR3-region contribution. vdjtools resolves anchors from arda. ALICE Neighbourhood enrichment measured against a V(D)J :term:`generation model` rather than against a control cohort: a clonotype with more similar neighbours than its generation probability predicts is evidence of antigen-driven expansion. Contrast :term:`TCRnet`. Command: ``vdjtools alice``. C-star Written ``cstar``. The sample :term:`coverage` a repertoire attained, in ``[0, 1]`` -- the estimated fraction of the underlying population represented by the observed reads. Used as the common level at which :term:`Hill number` s are read, so that two samples sequenced to different depths can be compared. Always reported in a :term:`signature`, even when the diversity columns it gates are holes. CDR3 The third complementarity-determining region, in the IMGT definition: the hypervariable loop that contacts antigen, **excluding** the Cys104 and Phe118 :term:`anchor` residues. Two residues shorter than the :term:`junction` for the same receptor. clonotype One distinct receptor sequence in a :term:`repertoire`, and one row of the canonical frame. What counts as distinct is a choice: vdjtools keys on the nucleotide junction plus V and J calls by default, and amino-acid keying is available where the analysis calls for it. clone size How many :term:`read` s support a clonotype -- the ``duplicate_count`` column. The distribution of clone sizes across a repertoire is heavily skewed, which is why unnormalised richness is not comparable between samples. cohort A set of samples analysed together, usually described by a metadata table with one row per sample. Cohort commands either parallelise over sample files or stream a pre-ingested Parquet store. corpus A fitted reference: a large collection of repertoires from which winsorization bounds, a per-locus rotation, and per-component centre and scale were all estimated in one pass. A :term:`signature` is only comparable to another one rotated through the same corpus. Nine are published; see :doc:`signature`. coverage The fraction of the underlying receptor population that the observed sample represents, estimated from the abundance spectrum. The basis for coverage-standardised diversity, and the reason :term:`rarefaction` beats fixed-depth downsampling. See :term:`C-star`. d50 A dominance index, in the legacy ``getDxxIndex`` definition vdjtools keeps: rank clonotypes by descending count, take the smallest number ``k`` whose cumulative read fraction reaches 50 percent, and report ``1 - k/Sobs`` -- the fraction of clonotypes **not** needed. Higher means more dominated by few clones. A value of 0.98 means 2 percent of clonotypes carry half the reads. downsampling Randomly reducing a sample to a fixed number of reads or clonotypes so that depth-sensitive statistics can be compared. Discards data; :term:`rarefaction` estimates the same quantity without doing so. expanded clone A clonotype supported by more than one read. ``(reads - singletons) / (richness - singletons)`` -- reads per expanded clone -- is at least 2 by construction, which is why it, and not reads per clonotype, can be drawn independently of the :term:`singleton` fraction. gene A germline V, D, J or C segment, named without an allele suffix: ``TRBV9``. See :term:`allele`. gene-level call A ``v_call`` or ``j_call`` naming a :term:`gene` rather than an :term:`allele`. Recombination models are keyed by allele, so passing a gene-level name to a Pgen function raises rather than silently marginalising over every allele -- an earlier version returned a V/J-agnostic value 2.38 times too high with no error. Resolve to a representative allele explicitly. generation model A probabilistic model of V(D)J :term:`recombination`: which segments pair, how many nucleotides are deleted from each, and what is inserted between them. Yields :term:`Pgen`. Models for all seven human loci ship in the wheel. See :doc:`model`. germline The unrearranged V, D, J and C segment sequences a receptor is assembled from. vdjtools resolves all germline through arda, which is the single source of truth: mixing germline sources within one model produces wrong answers that look plausible. Hill number A one-parameter family of diversity measures indexed by order ``q``, all in units of an effective number of clonotypes. ``q = 0`` is richness, ``q = 1`` is the exponential of Shannon entropy, ``q = 2`` is inverse Simpson. Higher ``q`` weights abundant clones more. holes A column that could not be computed, represented as ``nan`` and never as ``0``. "This locus was not sequenced", "this sample is too shallow to estimate this", and "this field was absent from the input" are different facts, and the ``mask`` channels record which applies. See :doc:`channels`. incidence Whether a clonotype is present in a sample at all, as opposed to how abundant it is. Incidence-based association tests ask whether presence co-varies with a phenotype across donors. See :mod:`vdjtools.biomarker`. junction The nucleotide or amino-acid sequence spanning the V-D-J join **including** the Cys104 and Phe118 :term:`anchor` residues -- AIRR's ``junction_nt`` and ``junction_aa``. Two residues longer than the :term:`CDR3`. vdjtools uses junction columns throughout; confirm which convention any external dataset uses before matching against it. Note that :func:`~vdjtools.model.native.pgen_aa` and :func:`~vdjtools.model.infer.infer_nt` name their argument ``cdr3_aa`` for historical reasons but require a **junction** --- anchors included. locus Which receptor chain a clonotype belongs to: TRA, TRB, TRG, TRD, IGH, IGK, IGL. Note that TRA and TRD share V segments, so a locus must never be re-derived per row from a gene name in a frame that is already keyed by locus. MAD Median absolute deviation, scaled by 1.4826 to match the standard deviation of a Gaussian. A robust scale estimate: a :term:`signature`'s per-component scaling uses the median and MAD of the corpus's component scores rather than mean and standard deviation, so a few extreme samples cannot set the scale. metaclonotype A group of similar clonotypes treated as one unit -- typically a ball of a given edit distance around a CDR3, optionally requiring matching V and J. Recovers signal that exact matching misses, because two donors rarely share a receptor exactly but often share a motif. N region The non-templated nucleotides inserted between V and D, and between D and J, during :term:`recombination`. Together with the deletions at each segment end, this is what makes the junction hypervariable and what a :term:`generation model` has to model. Pgen The probability that a V(D)J :term:`recombination` event produces a given receptor sequence, marginalised over every recombination scenario consistent with it. Low Pgen means a sequence unlikely to arise twice independently, so convergent observations of it are informative. Available over nucleotides, amino acids, and the Hamming-1 ball. productive A property of the **rearrangement**: in frame, with no stop codon, so it can encode a chain. Distinct from :term:`functional gene`, and filtering on one says nothing about the other -- a productive rearrangement can use a pseudogene V segment. See :doc:`preprocessing`. functional gene A property of the **germline gene**, in IMGT's classification: F (functional), ORF (open reading frame) or P (pseudogene). Orthogonal to :term:`productive`, which is why vdjtools keeps the two filters on separate flags. public clonotype A receptor observed in many unrelated donors. Usually a consequence of high :term:`Pgen` rather than of shared antigen exposure, which is why association testing conditions on generation probability. rarefaction Estimating what a diversity statistic would have been at a smaller sample size or a lower :term:`coverage`, from the abundance spectrum, without discarding reads. The basis of coverage-standardised comparison. Extrapolation is the same machinery run the other way. See :func:`vdjtools.stats.inext`. read One sequencing observation supporting a clonotype, counted in ``duplicate_count``. Where the protocol uses molecular barcodes the unit is a UMI rather than a raw read, which matters for depth-dependent statistics because UMI counts are not inflated by PCR. recombination The somatic process assembling a receptor gene: one V, optionally one or two D, and one J segment are joined, with nucleotides deleted from each segment end and non-templated nucleotides inserted between them. See :term:`N region`, :term:`generation model`. repertoire The full set of receptor sequences carried by one individual, in one sample, at one time. Represented as a table of :term:`clonotype` s with :term:`clone size` s. richness The number of distinct clonotypes observed. Strongly depth-dependent, so raw richness is not comparable across samples; Chao1 and Efron-Thisted estimate what was missed, and :term:`rarefaction` puts two samples on a common footing. rotation The per-locus linear map from raw features to principal components inside a :term:`corpus` artifact. A rotation is a function of features and components only: it carries no sample data and no cohort labels. Truncating it to fewer components is exact, so a wide artifact serves every narrower request. signature One fixed-width, named, already-standardised feature vector per repertoire. Split across two packages: ``vsig`` is the statistics half (vdjtools) and ``rsig`` the geometry half (`mirpy `_); they join on ``sample_id``. See :doc:`signature`. singleton A clonotype supported by exactly one read. The singleton fraction drives every coverage and richness estimator, and in a synthetic draw it stands in for the naive share of the repertoire. spectratype The distribution of junction lengths across a repertoire, optionally resolved per V gene. Skew away from the germline-determined shape indicates selection or clonal expansion. tandem D A rearrangement incorporating two D segments. Real but rare, and neither OLGA nor IGoR models it; vdjtools supports it exactly. Costs roughly 2.5 times a single-D evaluation, and no exact skip exists, because essentially no read has zero tandem-D contribution. TCRnet Neighbourhood enrichment measured against a **control repertoire**: a clonotype with more similar neighbours than the control predicts is a candidate for antigen-driven expansion. Contrast :term:`ALICE`, which uses a generation model instead. Command: ``vdjtools tcrnet``. winsorization Clamping a value to a percentile bound estimated on the :term:`corpus`, so that one extreme sample cannot dominate a rotation. Which side is clamped follows the feature's declared support, not its transform: a non-negative count is clamped at the top only, a log-probability at the bottom only, an unbounded score at both. Clamping edits a real measurement, so the fraction clamped is always reported in ``qc:-:winsor_frac``. arcsine transform A variance-stabilising transform for a proportion: ``arcsin(sqrt(p))``, in Anscombe's form which also uses the number of observations behind ``p``. A binomial proportion's variance is ``p(1-p)/m`` -- it depends on the value itself -- so an untransformed share drawn from a shallow sample is noisier than the same share from a deep one, and a model cannot tell the difference. After the transform the variance no longer depends on ``p``. Used for the amino-acid composition block. channel A signature column carried in its **own units**, never rotated and never winsorized: the coverage a sample attained, whether a locus was present, how much of the row was clamped. Channels answer "can I trust this row", and mixing a number whose job is to qualify a measurement into the measurement itself would destroy both. Distinct from a :term:`principal component` and from a :term:`named block`. See :doc:`channels`. clonality How unevenly reads are distributed across clonotypes, reported as ``1`` minus normalised Shannon entropy: 0 is perfectly even, 1 is one clone holding everything. Rises with age and with antigen-driven expansion. Depth-sensitive, so compare it only at matched :term:`coverage`. clr Centred log-ratio. The transform for a **closed composition** -- a set of shares that sum to one, such as V-gene usage or isotype composition. Take the log of each share and subtract the mean of those logs. Necessary because the parts of a composition are not free to vary independently: if one goes up another must come down, so treating them as ordinary variables makes the block singular and any correlation between them partly an artefact of the constraint. A ``clr`` block ships ``k-1`` parts for the same reason. dispersion How spread out a repertoire's receptors are in the embedding, as opposed to where their centre sits. Two donors can have the same average position and very different spread --- one concentrated in a few sequence neighbourhoods, one scattered. Reported by the ``rsig:div`` and ``rsig:disp`` families, and unmeasurable below a few clonotypes, which is when the whole family becomes a :term:`hole `. effective dimension How many directions a repertoire's receptors genuinely spread across, as a continuous number: the participation ratio of the embedding's variance spectrum. A repertoire filling 40 directions evenly and one dominated by 3 have very different effective dimensions even at identical total spread. isotype The antibody class a B-cell receptor carries --- IgM, IgD, IgG, IgA, IgE --- read from the ``c_call`` column. Class-switching is a record of B-cell history, so isotype composition is a feature in its own right and is reported for IGH only. Transformed with :term:`clr`, because the classes are shares of one whole. n_eff Effective sample size: how many independent observations a weighted sample is worth, computed from the clone-size weights. A repertoire of 10,000 reads dominated by one clone is worth far fewer than 10,000 independent draws, and every estimator's precision scales with ``n_eff`` rather than with the raw read count. named block A raw feature group emitted under its own name, beside the rotated components --- the diversity, depth, clone-size, length, isotype and cross-locus numbers the rotation is fitted on. Requested with ``--named``; off by default. **The values carry their declared transform, not a natural scale**: a ``log10`` diversity comes back as a ``log10``, so ``vsig:div:TRB:1D_c = 2.386`` is not a clone count. Which groups are reportable is declared by the library, not inferred from how many columns they have. principal component One axis of the :term:`rotation`: a weighted combination of raw features, chosen so that the axes are uncorrelated across the :term:`corpus` and ordered by how much variation each one carries. The ``:pc:`` columns of a :term:`signature`. A component is a **direction, not a quantity** --- ``PC01`` has no units and no interpretation as a measurement, which is why the raw numbers behind it are available separately as a :term:`named block`. Rao quadratic entropy A diversity measure that accounts for how *different* the clonotypes are, not just how many there are: the average dissimilarity between two clonotypes drawn at random in proportion to their abundance. A repertoire of 100 near-identical sequences and one of 100 unrelated ones have the same :term:`richness` and very different Rao entropy. robust z-score The units a :term:`signature` column is on: the value minus the :term:`corpus`'s median for that column, divided by its :term:`MAD`. Median and MAD rather than mean and standard deviation, so that a handful of extreme corpus samples cannot set the scale. A value of 2 means "two robust deviations above the corpus's typical sample". SHM Somatic hypermutation: the antigen-driven mutation of an already-rearranged B-cell receptor during affinity maturation. Summarised as mean V-segment identity to germline, so a lower value means a more mutated repertoire. IGH only. support The declared range a raw feature can take --- non-negative, non-positive, unbounded, or a fixed interval. Declared per feature and used to decide which side :term:`winsorization` clamps. It is deliberately **not** derived from the feature's transform: three features declare no transform and three different supports, so deriving one from the other trims the wrong end. Zipf law The rank-abundance law used to give a synthetic repertoire a realistic clone-size distribution: the ``i``-th most abundant clone gets frequency proportional to ``i`` to the power ``-a``. Note this is the **rank** law, not numpy's ``Generator.zipf``, which samples integers *from* a Zipf distribution and at ``a = 1.5`` has infinite mean --- normalising such a draw collapses a repertoire to a handful of clonotypes.