VDJdb build outputs#
Every file the build produces, what it contains, whether it ships, and who consumes it.
myst-parser renders this file on the documentation site from the Markdown source, which is why it
stays Markdown rather than becoming .rst. Its section numbers are cited from ROADMAP.md, from the
package docstrings and from docs/tuning/, so treat them as stable identifiers and do not renumber a
section.
Status key: shipped in the release zip · artifact produced and uploaded by CI but not zipped · internal produced during a build, not published.
1. Release bundles#
Three zips per release. manifest.json names them by role, so consumers select by role rather than by
asset order (ROADMAP §3.1).
Asset |
Role |
Contains |
|---|---|---|
|
|
the new VDJdb format (§3) |
|
|
byte-layout-compatible with the historical release (§2) |
|
|
AIRR Rearrangement + Receptor/Reactivity (§4) |
Alongside them, as release assets rather than zip members: manifest.json, SHA256SUMS.
2. Legacy bundle - vdjdb-legacy-<version>.zip#
Exactly ten members under a single vdjdb-<version>/ directory. The member basenames are a contract:
vdjmatch looks inside the zip for vdjdb.txt, vdjdb.slim.txt and vdjdb_full.txt by basename,
and vdjdb-web resolves the rest as <database.path>/<name>.
File |
Rows (2026-06-03) |
Cols |
Consumer |
|---|---|---|---|
|
284,546 |
22 |
vdjdb-web; its schema is built from |
|
21 + header |
8 |
vdjdb-web; must match |
|
197,729 |
17 |
vdjmatch, standalone R/Python users |
|
16 + header |
2 |
standalone users |
|
192,753 |
35 |
vdjmatch, standalone users |
|
55,636 |
19 |
vdjdb-web; parsed positionally, no header check |
|
40,061 |
27 |
vdjdb-web; parsed positionally, no header check |
|
- |
- |
vdjdb-web |
|
- |
- |
- |
|
40 lines |
1 |
legacy self-update clients; line 1 must point at this zip |
Optionally also cluster_members_tcremp.txt and motif_pwms_tcremp.txt, same positional schemas.
Column orders (positional contracts)#
vdjdb.txt (22): complex.id gene cdr3 v.segm j.segm species mhc.a mhc.b mhc.class antigen.epitope antigen.gene antigen.species reference.id vdjdb.score TCR_hash method meta cdr3fix web.method web.method.seq web.cdr3fix.nc web.cdr3fix.unmp
Production vdjdb-web additionally serves five appended evidence.* columns (§3.4).
vdjdb.slim.txt (17): gene cdr3 species antigen.epitope antigen.gene antigen.species complex.id v.segm j.segm mhc.a mhc.b mhc.class reference.id vdjdb.score TCR_hash j.start v.end
vdjdb_full.txt (35): the 31 chunk columns + cdr3fix.alpha cdr3fix.beta vdjdb.score TCR_hash
cluster_members.txt (19): species antigen.epitope antigen.gene antigen.species mhc.a mhc.b mhc.class gene cdr3aa x y cid csz v.segm j.segm v.end j.start v.segm.repr j.segm.repr
motif_pwms.txt (27): species antigen.epitope gene aa pos len v.segm.repr j.segm.repr cid csz count count.bg total.bg count.bg.i total.bg.i need.impute freq freq.bg I I.norm height.I height.I.norm antigen.gene antigen.species mhc.a mhc.b mhc.class
vdjdb-web parses cluster_members.txt and motif_pwms.txt positionally, against a fixed 19-entry
and 27-entry column-type array and with no header check. An inserted, removed or reordered column
mistypes or shifts the table without an error, so the column order of both files is a requirement,
not a convention.
Preserved legacy quirks#
vdjdb_full.txt’scdr3fix.*cells are Pythondictrepr, not JSON (single quotes). All 122,930 non-empty cells in 2026-06-03 are this form.vdjdb.txt’smethod/meta/cdr3fixuse Pythonjson.dumpsdefaults:", "separators andensure_ascii=True, so the en dash inM158–66is escaped: it is written as\u2013.struct.json_encode()is not byte-compatible.No field is ever quoted; all three tables contain zero
"characters.The
web.methodrow ofvdjdb.meta.txthas a space where a tab belongs, and all fourweb.*rows have a field shift. Reproduced in legacy, fixed in the new format.
3. New VDJdb format - vdjdb-<version>.zip#
A normalised star schema rather than one denormalised table. Three fact tables plus the joined view. Parquet, with a TSV projection of each for users without a parquet reader.
Column names here are underscore_case, not the dotted names the legacy tables use: mhc_a,
antigen_epitope, v_segm, cdr3nt_pgen. A dot is a table qualifier in SQL and blocks attribute
access in most dataframe libraries, and these tables already mixed the two conventions -
record_id and pmhc_id beside mhc.a. The rule is a plain dot substitution with case left alone,
so TCR_hash and meta.donor.MHC ship as TCR_hash and meta_donor_MHC. The legacy tables do not
move: they are a positional contract vdjdb-web parses.
vdjdb.schema.json states both names per column - name is what the legacy tables ship and what
the registry indexes, ships_as is what these tables ship - and
Columns is the generated table for each, so neither can drift from the
build. The chunk format is unaffected: a submitted chunk still uses the dotted names.
3.1 records.parquet - one row per submitted record#
Primary key record_id, unique. One chunk row is one record: a chunk is one paper, a row is that
paper’s report on one clone, and the row reports both chains. method.* and meta.* sit here
because they describe what the publication reports about the record.
The table has one row per curated line, 192,609, which is the 202,263 data lines in chunks/
less 9,636 declared within-chunk duplicates and less 18 rows where one publication was curated in two
chunk files (#390). CHUNK_DEDUP_KEY contains reference.id, so a group of it spanning two chunks
is one paper reporting one clone twice - the chunk is normally the publication, and where the two
come apart the publication is what deduplication is about. 19 such groups exist over 38 rows; 18
merge, filling the base row’s blanks from the other, and the 19th is left alone because
PDB_Database.tsv and PMID_34433824.tsv give one clone structural and tetramer-sort, which is
a solved complex and the sort that found it. out/reports/repeated-references.tsv lists all 19 with
the verdict and the reason.
The receptor is not here: a chain is an observation, so it is a row of chains, while
vdjdb_full.txt folds both chains into paired columns and leaves half of them blank.
35 columns:
Group |
Columns |
|---|---|
identity |
|
antigen |
|
provenance |
|
sample |
|
annotation |
|
method |
|
score |
|
curation |
|
submitter, comment, chunk_id, meta_subset_frequency and method_pairing are kept here; the
legacy build discards all five.
method_frequency, method_frequency_count and method_frequency_total are three independent
columns, and none is derived from another (#696). A study that reports only a float has no count
behind it, so deriving the float would blank it exactly where it is the only measurement; deriving
the pair from a float is impossible. Measured: 43,231 records carry a count and total, 17,700 a
percentage, 2,601 a float, 129,109 nothing. The count and total are chunk columns, so a submitter
with a read count writes it as a number; where they do, the submitted value wins and nothing is
parsed. Where all three are present they must agree, which vdjdb qc reports and does not repair.
meta_subset_frequency is not filled from method_frequency. It is populated on 2,412 records
and left as submitted, because the records that do carry it are using it for a different quantity -
method_frequency = 17/52 beside meta_subset_frequency = 0.70% is a clonotype’s count within a
sorted subset beside that subset’s share of the sample, and both are real. A consumer that wants
“the frequency of this clonotype in its subset” should coalesce the two:
pl.coalesce("meta_subset_frequency", "method_frequency"). Filling it in the build would have
written 61,253 cells across 117 chunks, every one a copy of the column beside it, and mixed the two
readings with nothing to tell them apart.
content_hash, the record state and the release/commit provenance are in the registry (§6), not
here: they describe the record’s history rather than the record.
3.2 chains.parquet - one row per TCR chain of a record#
Primary key (record_id, gene). This is the level vdjdb.txt is written at. Chains are a separate
table so that record fields are not duplicated per chain, as in vdjdb.txt, and not folded into
paired alpha/beta columns, as in vdjdb_full.txt.
35 columns: record_id, gene (TRA/TRB), clonotype_id, clone_id, cdr3, v_segm,
d_segm, j_segm,
v_end, j_start, cdr3nt, cdr3nt_pgen, cdr3nt_margin, v_inferred, j_inferred,
d_inferred, d_start, d_end, d_posterior, v_end_inferred, j_start_inferred,
cdr3_original, fix_needed, fix_good, v_fix_type, j_fix_type, v_canonical, j_canonical,
v_segm_submitted, j_segm_submitted, d_segm_submitted, v_segm_arda, j_segm_arda,
TCR_hash.
j_segm is the J that ships, and it is not always the J the paper reported. Three values sit side by
side: j_segm_submitted is the call as submitted, j_segm_arda is the allele arda.cdr3fix aligned
the junction against, and j_segm is what ships. A J is not used when it misses the junction’s last
3 residues, anchor excluded and one mismatch beside the anchor tolerated: it is replaced by the one
other gene of the chain’s locus that matches 3 or more, or by the one gene of its own family that
matches 2 or more when nothing reaches 3. Ties are not broken, a call no other gene explains stands,
and the junction itself is never altered. j_start and j_canonical are then about the gene that
ships (#681; vdjdb.curate.jcalls.recall).
cdr3nt is inferred, not observed (#461): it is the most plausible nucleotide junction behind the
amino-acid one, from the recombination model. 261,097 of 286,047 chains have one and each
back-translates to its junction, but two models agree on only 7.2 % of the sequences, so the column is
a representative history rather than evidence. cdr3nt.pgen is its generation probability and
cdr3nt.margin its margin over the runner-up; 9.4 % of those chains have a margin below 1.1, where
the choice was near-arbitrary. Filter on the margin rather than treating the sequence as observed
(ROADMAP §19).
cdr3fix is not a column here. Every member of the legacy JSON blob is its own variable:
cdr3.original is the sequence as submitted, the four fix.* / *.fix.type columns say what was
done to it, and v.canonical / j.canonical say whether the anchors are the expected ones. In the
legacy release the same field is a JSON number on one row and a string on the next.
emit/legacy.py reassembles the blob on the way out, and nothing else may.
clonotype_id is CT plus 16 hex digits of a sha256 over (species, gene, cdr3, v.segm, j.segm).
Records reporting the same receptor chain share it, and it is the level motif evidence and the
independent-study support count attach at. clone_id is CX plus 16 hex digits over the record’s two
sorted clonotype_ids, and is the empty string on the 99,459 records reporting a single chain.
Both are hashes of their own keys rather than counters, so adding or removing a chunk cannot renumber
anything, and sha256 rather than a library hash function so that a dependency upgrade cannot renumber
the database either. records.pmhc_id and records.epitope_id are the antigen-side equivalents.
Identifiers is the reference: the five levels, the two mechanisms, the
lifecycle fields and the seven invariants vdjdb identity check asserts.
d.segm is the curated D call, as the publication reported it. d.inferred, d.start and d.end
describe the D of the recombination scenario that produced cdr3nt, so the coordinates index that
sequence (0-based, half-open). They agree with the curated call at gene level on 78.4 % of the 40,892
beta chains that have one.
v.inferred and j.inferred hold a call proposed from the junction, and only where the publication
named no segment (#462, #658): 706 of the 745 chains with no V, and 3,272 chains with no submitted J.
They never sit beside a curated call.
The two do not have the same standing, and the split is deliberate. j.inferred also ships, in
j.segm, as the equivalent has in every release; v.inferred ships nowhere, and v.segm stays blank
on all 745 chains whose publication left it blank. The reason is measured: recovering a hidden J from
the junction alone works on 93.6–97.5 % of chains at gene level, because a J germline templates a
distinctive 3′ motif, while a V works on 23.8–50.1 %, because the junction contains little V sequence
and most of what it does contain templates the same CAS. So the J proposal is good enough to be the
record’s J and the V proposal is not; it is offered beside the blank instead.
Two sources answer, in order: the recombination model (vdjtools.model.infer_nt_batch, with the blank
side marginalised), then arda’s germline anchor table where the model declines - which is also the only
source for a species no model covers, and is what fills the 74 macaque j.alpha blanks that nothing
filled before. Until #658 this was a k-mer scan over res/segments.txt, a by-product of a 2023 IMGT
import whose V half never worked at all: 3 non-empty guesses in 4,000 sequences.
v.end.inferred and j.start.inferred are the two sources of a V/J boundary, kept apart on
purpose (#631). v.end and j.start are arda.cdr3fix’s answer, read off its protein alignment
against the germline of the segment the record names, and -1 where it declined - 4,163 chains for
v.end and 1,164 for j.start. The .inferred pair is the boundary a second germline alignment
supports, from vdjtools.model.germline_boundary, which decides how far into the boundary codon the
germline reaches rather than rounding to a residue. It is filled only where the first declined:
2,442 and 488 chains. Everywhere else it is -1, so the two are distinguishable by column and
coalescing them cannot overwrite the shipped answer.
Two boundaries rather than one filled column, because they are answers to the same question with
different precision and the shipped column is what vdjdb-web reads. Against the external nucleotide
truth in tests/release/test_cdr3fix_accuracy.py the codon decision is exact on 92.9 % of v.end
and 98.0 % of j.start, where a protein alignment rounding to whole residues - VDJdb’s k-mer
scanner and arda.cdr3fix alike - is 71.8 % and 97.1 %. It is not promoted into v.end anyway: that
is a change to what the database says about every record, which belongs to a curation decision and to
its own declared rules, not to a build improvement. What it does do is replace -1, which carries
nothing at all.
The .inferred pair used to come from the argmax recombination history behind cdr3nt instead. That
was 11 points worse on v.end, because maximising P(sequence) explains N-region nucleotides as
templated whenever it can; antigenomics/vdjtools#182 was opened from this measurement and fixed it.
The legacy tables and the cdr3fix JSON vdjdb-web parses keep the alignment’s -1 untouched
throughout.
d.posterior is the probability of the gene d.inferred names, from the same recombination
scenario weights that named it - so the number beside the call is the probability of that call.
Naming the D and placing it are separate questions and one estimator answers each: the model names
the gene, the aligner places it. Measured on 4,000 real human TRB rearrangements whose D and
coordinates come from the nucleotide sequence, the gene is right on 74.35 % of all rows and 99.80 %
carry coordinates. TRBD1 and TRBD2 are short, heavily trimmed and similar, so the junction often
cannot choose between them: filter on d.posterior; do not read d.inferred alone (ROADMAP §20).
There is no d.entropy. It came from a separate estimator (arda.dpost) that named a different D
gene from the one it annotated on 21.5 % of chains, and that estimator is retired: the posterior now
comes from the scenario weights and there is no second distribution to take an entropy over.
3.3 evidence.parquet - long format, one row per piece of evidence#
Primary key (record_id, evidence_id). The table is long rather than wide because a record may have
any number of pieces of evidence of any number of kinds: a wide table would be mostly empty and would
gain a column per producer.
record_id, gene (empty when the evidence is record-level), evidence_id, evidence_type,
evidence_source, evidence_value, evidence_score, first_seen_release.
|
|
|
|
|---|---|---|---|
|
- |
the other |
count of distinct references |
|
release tag |
cluster id |
cluster size |
|
release tag |
cluster id |
cluster size |
|
PDB |
PDB id |
- |
|
model set id |
structure hash |
model confidence |
independent_study is the only producer today: 53,913 rows over 48,893 records, scores 2 to 41. It is
the same computation as the ROADMAP §11.1 tuning objective, in one implementation, so the shipped
column and the objective cannot disagree.
Structure evidence is keyed on the legacy TCR_hash today and moves to record_id when the structure
store is re-keyed.
No held-out validation data is ever an evidence row. See §7.
3.3a epitopes.parquet and restriction.parquet - the antigen catalogue#
VDJdb’s own list of epitopes and the MHCs that present them.
Table |
Key |
Rows |
|---|---|---|
|
|
2,131 |
|
|
2,343 |
epitopes has antigen.gene, epitope.length, mhc.class, and the support counts records,
chains, clonotypes and references. 379 epitopes are reported by two or more publications.
proteome_peptide and proteome_substitution link an epitope to the host-proteome peptide it is
one substitution from, where the epitope is not itself in the proteome (#632). 211 of the 2,131
epitopes carry the link, and 67 of those have the proteome form curated in VDJdb as a separate
row - 35,346 records, 18.3% of the database, on rows nothing else says are two forms of one
peptide:
epitope |
records |
|
records |
|
|
|---|---|---|---|---|---|
|
29,729 |
|
13 |
|
NY-ESO-1 |
|
2,401 |
|
140 |
|
MLANA |
|
2,495 |
|
5,048 |
|
INS |
|
100 |
|
19 |
|
PMEL |
A query for NY-ESO-1 responses returns the 29,729 and never learns the 13 exist. That is what the columns are for.
Neither row is wrong, and this is not a defect report. The epitope sequence is the ground truth -
it is the peptide the experiment used - and the difference from the proteome is almost always
deliberate: an anchor-optimised vaccine peptide, a designed altered-peptide ligand, a heteroclitic
variant, or a structure solved with a modified peptide. Of the 36,496 records on these 211 epitopes,
58 carry a meta_structure_id and 54 come from PDB_Database.tsv, so the crystallography case is
real and small. The columns are named for what was measured and not for why, because the sequence
tells those causes apart from none of the others.
Empty on the other 1,920 rows, and on every viral or bacterial epitope, because only the human and
mouse proteomes are read - a pathogen epitope’s source is the pathogen’s proteome, which is a
per-pathogen fetch and a different question. Both columns are read from
proofreading/epitope_proteome.tsv, a committed reviewed input refreshed by vdjdb antigens in its
own pull request, never resolved during a build: mhcmatch fetches the proteome from HuggingFace,
so doing it here would put a network call in the critical path (hard rule 9).
The key is the epitope and the species, not the peptide. A reader who assumes one row per peptide
which the table’s name invites - joins the records of 12 peptides twice.
patches/antigen_epitope_species_gene.dictis keyed on the peptide alone and cannot express them, so this is a table rather than a view over the patch.
Those 12 are three different things, and out/reports/epitope-sources.tsv separates them (#633):
6 are conserved peptides - the same sequence in two organisms’ proteomes, so two publications are two independent reports.
VEALYLVCGis in humanINSand mouseIns2,RPIIRPATLin influenza A and BNP,LRVMMLAPFin E. coli and S. TyphiyeiH. The report marks themconserved, which is the column to sort on, so a new collision is the only thing to look at.6 are vocabulary gaps - one organism written two ways (
AdVbesideHAdV5), orSyntheticinantigen.species, which is a provenance and not a species. These want an alias table and which way each folds is a curator’s call (#632, #637).a mis-curation, now repaired:
RGPGRAFVTIwasHomoSapiens/P18-I10on one row againstHIV-1/GP160on 85, whereP18-I10is the laboratory name of the HIV-1 V3-loop peptide itself.
The same report carries the other half of the ambiguity, which nothing reported before. epitopes keeps
the modal antigen.gene for each (peptide, species), and 30 of them have more than one label - so
the report counts them in a genes column beside the label the catalogue kept.
Its first run found FVVKAYLPVNESFAFTADLRSNTGGQA carrying 187 labels, Eef2 to Eef188, one per
record and consecutive: a spreadsheet autofill had incremented one gene name down the column. That is
repaired (#397, #694) - mhcmatch’s mouse proteome assigns the peptide to Eef2 and to nothing else -
so the largest remaining is 2.
restriction checks each allele against IPD-IMGT/HLA (https://www.ebi.ac.uk/ipd/imgt/hla/) by
prefix, since a VDJdb call is two-field and the authority stores four, and against
proofreading/mhc_nonhuman.tsv for the names that database does not cover (murine H2-, macaque
Mamu-, and B2M). So mhc.a.status and mhc.b.status read known, unconfirmed, declared or
unknown.
unconfirmed is a refinement of known, not a rejection (#634): the call resolves in IPD and every
allele under it is Unconfirmed - it rests on a single submission rather than an independent
observation. The existence check cannot make that distinction, because every one of these carries a name
IPD has. Measured 2026-09-29: 14 calls over 105 records, and 80 of those are one call,
HLA-A*02:01:48 - a third-field allele resting on one cell from one submitting group, where
HLA-A*02:01 has 169 Confirmed alleles under it. What to do about that is a curation question about
what the submitters meant, so it is reported and never fatal.
A build carrying an unknown or blank call fails, naming the value, the column, the cell count and
the chunks that report it, so only known, unconfirmed and declared reach a release. The fix is an entry in
patches/mhc.dict when the call is wrong, or a row in proofreading/mhc_nonhuman.tsv when it is a
species IPD-IMGT/HLA does not cover. Whether the allele could present that peptide, rather than only
whether its name exists, is phase 9e: presentation.tsv above for the offline half, and the six
columns below for the model’s ranking.
restriction also carries six promiscuity columns (ROADMAP §10.6). An epitope is often presented by
several alleles and the curated one is not always the best binder; both facts belong in the database
and neither belongs in a key, because a model upgrade would otherwise renumber pmhc_id and break
every external reference while the build passed.
Column |
What it is |
|---|---|
|
distinct |
|
the panel allele that presents this epitope best |
|
where the curated allele sits in that ranking, 1 being the top |
|
the curated allele’s |
|
panel alleles in the strong band for this epitope |
|
the model that produced the five above, per row. Empty on a row scored before the column existed |
All six are a join against the committed proofreading/epitope_promiscuity.tsv (13,510 rows over
1,729 epitopes and 107 alleles). Nothing is predicted during a build: mhcmatch fetches its reference
data from HuggingFace, and a build that downloads a model is neither offline nor deterministic (hard
rule 9), so the table is a reviewed input refreshed by vdjdb promiscuity through its own pull
request.
A curated allele the prediction outranks is not a defect. The curated allele is the one a
publication typed a donor for, which a proteome-background ranking has no access to, and most of these
epitopes are promiscuous. 1,793 of 1,989 class I pairs get a rank; the 196 that do not divide into
four causes, each a different statement: 86 pairs (16,316 records) name an allele outside the panel,
86 (10,395) an epitope the table has not been refreshed to cover since it was written, 24 (381) an
epitope outside the 8-11mer class I range, and the rest resolve only at a depth the panel does not
name. A deeper spelling such as HLA-A*02:01:48 is scored at its two-field molecule, because the
panel is named at two fields and there is no deeper groove.
3.4 vdjdb.parquet - the joined view#
records ⋈ chains ⋈ evidence, one row per chain, with each evidence type pivoted to a boolean. It is
the denormalised convenience table, derived on every build and never authored by hand.
Six boolean columns, all six always present: evidence.motif.tcrnet, evidence.motif.tcremp,
evidence.structure.native, evidence.structure.model, evidence.validation.independent,
evidence.validation.same.study. An evidence type with no producer yet is false, meaning no
evidence of that kind, rather than a missing column, so the shape of the view does not change as
phases 8–11 land.
The names match the five evidence.* columns production vdjdb-web already serves; nothing in this
repo produced them before this view.
3.5 Generated metadata#
vdjdb.schema.json is the field registry in machine-readable form: per column, its vdjdb.meta.txt
attributes, its position in every table that includes it, and its physical dtype. Every other schema
artifact (vdjdb.meta.txt, the legacy column orders, the AIRR mapping in phase 7, the docs tables) is
a projection of the same registry, so none of them can drift.
The dtype is read off the written frame rather than declared, so the schema cannot claim a type the shipped files do not have.
3.6 clusters.parquet and motifs.parquet - the motif tables#
The new-format counterpart of cluster_members.txt and motif_pwms.txt (§2). They resolve the
cluster ids that evidence.parquet stores in evidence_value on its motif_tcrnet and
motif_tcremp rows; in the legacy bundle those ids resolve only into a positionally-parsed text file.
Both methods go into one pair of tables, separated by the method column rather than by a second pair
of files.
clusters has one row per (method, cid, clonotype_id), i.e. a cluster’s membership:
Column |
Note |
|---|---|
|
|
|
|
|
joins to |
|
the scope the clustering ran in |
|
cluster size, in clonotypes |
|
modal allele over the cluster |
|
graph layout coordinates, for |
Membership is keyed on clonotype_id, not on the CDR3/V/J triple. The legacy file repeats cdr3aa v.segm j.segm plus seven annotation columns on every member row, which is why a 55,636-row file has
19 columns of mostly-duplicated epitope metadata. Here the annotation appears once, in records and
epitopes, and the membership row stores a key, so cluster_members.txt is one join away and the
normalisation rule of §3 holds.
motifs has one row per (method, cid, pos, aa), the position weight matrix:
method, cid, pos, aa, len, count, freq, count.bg, total.bg, count.bg.i,
total.bg.i, level.bg, freq.bg, I, I.norm, height.I, height.I.norm.
A residue a cluster never shows has no row, rather than a zero-count one: the logo has no letter there, and the legacy file omits it too.
level.bg names which background stratum supplied count.bg: the (v.gene, j.gene, len) cell when
it has support, otherwise the coarser len-only cell. The legacy schema records the imputation as a
bare need.impute boolean, which says that a fallback happened but not what it fell back to. The
legacy projection derives that boolean from level.bg.
Backgrounds never ship (CLAUDE.md hard rule 5). count.bg and total.bg are derived statistics
computed against a background streamed at build time; no background row reaches any output.
The per-epitope diagnostic for both tables is reports/motifs_per_epitope.tsv (§5), a report rather
than a shipped table, derived entirely from clusters and the corpus.
4. AIRR bundle - vdjdb-airr-<version>.zip#
File |
Level |
Rows (current corpus) |
|---|---|---|
|
one row per chain; AIRR Rearrangement |
286,047 |
|
one row per paired record; AIRR Receptor |
81,003 |
|
one row per record; AIRR Reactivity |
192,753 |
|
the AIRR schema version the files conform to (2.0) |
- |
The files are linked by cell_id, which is the record_id: a VDJdb record is one publication’s
report on one T-cell clone, and a clone is what AIRR’s Cell names. Both are standard AIRR fields.
Reactivity has ligand_type (MHC:peptide), antigen_type (peptide), antigen,
antigen_source_species, peptide_sequence_aa, mhc_class, mhc_allele_1, mhc_allele_2,
reactivity_method, reactivity_readout, reactivity_value, reactivity_unit and
reactivity_refs.
reactivity_readout is confidence and reactivity_value is vdjdb.score, which is what the AIRR
spec asks of a non-physical assay: “for inferred and annotated methods this should indicate a
confidence/quality level”. reactivity_method is MHC_peptide_multimer, native_protein or
annotated; ROADMAP §18 gives the classification and its counts.
Nucleotide fields are present and empty until phase 8 (#461): sequence, junction, the two
alignments and the three cigars. The schema requires the column, not a value, and
airr.validate_rearrangement passes on the full table today.
Receptor requires receptor_variable_domain_{1,2}_aa, the complete mature variable domain and
non-nullable, so both are rebuilt by stitching germline V and J around the inferred nucleotide
junction and translating (ROADMAP §22). Domain 1 is the beta chain, domain 2 the alpha, as the
schema’s controlled vocabularies require. receptor_hash is a sha256 over the two concatenated
domains and is not VDJdb’s TCR_hash, which hashes CDR3s, segments, MHC and epitope and is the key
the structure store uses.
A receptor is a two-domain object, so only paired records appear in that file: 81,003 of the 93,294 paired records have both domains rebuilt. An unpaired record is not dropped; its chain is in the Rearrangement file and it has a Reactivity row of its own, which is where AIRR puts a single rearranged sequence. Overall the AIRR export keeps 1,141 records and 1,501 chains that the legacy build discards (ROADMAP §18).
CI validates the Rearrangement file with the airr package’s own schema validator. airr 2.0.0 has
no Receptor or Reactivity validator, so the Reactivity file is gated against the field list read
from the package’s own airr-schema.yaml instead.
Converting an older release#
vdjdb convert airr --legacy vdjdb.txt produces the same two files from a legacy release zip, for
users who hold one. It runs the same emitter, since legacy vdjdb.txt already uses VDJdb’s column
names, so there is no second mapping to drift. The output has no d_call, because the legacy file has
no D column, and the 1,501 chains the legacy build dropped are absent.
5. Produced but not shipped#
Written to out/reports/ on every build, uploaded as CI artifacts, excluded from every zip. Whether
they ship in future is open (ROADMAP §9).
File |
What it is |
|---|---|
|
records passing the gene, allele and canonical-CDR3 masks |
|
records whose V/J gene name fails the IMGT check |
|
records whose allele number exceeds the gene’s allele count |
|
records whose CDR3 is not biologically valid |
|
the three tables with a |
|
every chunk QC finding: file, row, column, rule, value |
|
one row per rule that fired: level, rule, findings, chunks, whether it is advisory. Written by |
|
text-level findings: encoding, BOM, CRLF, header shape |
|
added / amended / retired records against the previous release |
|
the comparison against a reference release (ROADMAP §5) |
|
one row per (species, gene, method, epitope): clonotypes, clustered, retention, clusters, largest cluster, mean cluster size, singleton clusters, percolation, replicated, tp, precision, lift. 678 rows. Written by |
|
one row per motif stage: the same columns as |
|
one row per spelling of a value that some other value differs from only in case or in a |
|
one row per |
|
one row per chain whose |
|
one row per value the build rewrote, across the four harmonisation passes |
|
one row per chain whose J call some other J gene explains better (#681). A junction’s 3’ end is templated by the J germline, so a record’s J call can be checked against the sequence the same record reports: match the end - the anchor excluded, so a corrupt anchor cannot vote on its own diagnosis - against every J germline of that species and take the longest run. This is not the question |
|
one row per segment call IMGT has at neither allele nor gene level, after harmonisation (#389). This is the report the retired build wrote as |
|
one row per |
|
one row per chain-segment IMGT does not call functional, and |
|
one row per build stage: the parent stage it sits inside, wall seconds exclusive of any stage timed inside it, share of the recorded total, peak RSS in MiB, the record count and the core count. Seconds sum to the wall clock and shares sum to 1, which they did not while a nested stage was counted both in its own row and in its parent’s - the motif report summed 268.7 s over a 198 s step and understated every share by 36 %. Written by |
|
one row per (species, gene, source, axis): eighteen axes for four sources – the last legacy release, the latest release, this build’s two methods, and the partition that clusters nothing. 180 rows. Thirteen axes score a clustering on this build’s cohort (two of them, |
|
the same table as markdown, with each axis against its baseline and against |
|
per-clonotype enrichment statistics, embeddings, cluster labels, eps sweeps, the pooled cross-epitope confusion matrix |
|
the eight dashboard panels tiled, for visual review |
The three *_scored.txt files are currently incorrect: MotifsScoresAssembler.py sets
cluster.member for every record of an (epitope, species, gene) rather than matching on CDR3, which
the rewrite fixes.
6. Record registry and the identity lifecycle#
identity-lifecycle.tsv is one row per id ever published, at any level: id, level, state,
first_release, last_release, replaced_by. No key columns, because a key is recomputable from
chunks/ and an id’s history is not. 15.9 MB for 274,683 derived ids, so it ships as a release asset
beside the zips and is listed in SHA256SUMS. It is written by vdjdb release and never by a build:
a curation branch that adds a clonotype and removes it again has retired nothing. A build with no
previous copy produces exactly the same ids and reports only that the history is unknown.
registry/records.tsv maps record_id to its state, hashes, provenance and amendment history, so an
id survives a curator fixing a typo. It is committed and is an input to the build, not an output of
it: 73.8 MB for 192,883 rows, and kilobytes per amendment in the pack, because a reconciliation rewrites
only the lines it touched. It was going to ship as a release asset the build fetched, and #672 changed
that - without a committed copy the registry went stale between releases and landing one 40-record
chunk moved record_id on 168,723 of 192,753 records (#638). It is written only by
vdjdb identity update, which a chunk branch runs and commits alongside the chunk; vdjdb build reads
it.
Fifteen columns. Thirteen are the state, the two hashes, the chunk provenance, the release and commit at first and last sighting, and the amendment count with the key hash it came from. The other two are the lifecycle a consumer follows:
Column |
What it answers |
|---|---|
|
backwards, and only for an amendment: which key this record used to have |
|
forwards, and only for a retirement the amendment pass refused: which id took over. Two key fields moving is a new record by the rule the registry is built on, so the retirement is right and the pointer is what makes it diagnosable rather than a disappearance (#693, ROADMAP §10.4). Written when one id retires from a line of a chunk, exactly one is allocated against that same line, and the two natural keys name the same receptor; empty otherwise, including on every retirement that is a genuine deletion |
6a. Committed, not produced#
File |
Role |
|---|---|
|
publication years, so the dashboard render is offline and deterministic. A committed, reviewed input refreshed by its own pull request, never written by a build |
|
dashboard event callouts, with no hardcoded coordinates |
|
the differences |
|
where each self epitope sits in its own species’ reference proteome, and how exactly (#632). Written by |
The pinned motif parameters are not a data file. They are the TUNED constants in
vdjdb.motifs.tcrnet and vdjdb.motifs.tcremp, each stated with the scorecard it was chosen from
and the one cost it pays, in the module that uses it.
7. Never produced, never shipped#
TCRvdb / MATCHMAKERS (Messemaker et al., doi:10.1101/2025.04.28.651095) is proprietary: academic, non-commercial, no redistribution, in whole or in part.
It is used only as a held-out validation set, read from a path given by VDJDB_TCRVDB and never from
inside this repository. No file in any bundle, artifact or report may contain its rows, its per-record
labels, or any value derived from them at record granularity. Aggregate validation metrics (counts,
AUROC, recall at a threshold) may be reported; per-record verdicts may not.
Motif clustering is tuned on the independent-study support count: 4,129 human clonotype-epitope pairs
with ≥2 distinct reference.id, across 18 epitopes with ≥20 pairs. TCRvdb is read once, at the end,
to validate. Tuning on it would invalidate it as a validation set, and its 614 labels span only two
epitopes (YLQPRTFLL, GLCTLVAML), too narrow a basis for a global hyperparameter.
CI enforces this: a guard fails the build if any TCRvdb-derived file appears in the repository or in a release bundle.