Identifiers#
A VDJdb record is a receptor against a presented peptide. Both halves recur across records, and consumers reference both halves, so each level carries an identifier of its own.
The antigen is the peptide-MHC complex (terminology), which is why the levels
below stop at the peptide and its presentation: pmhc_id is the finest statement of what a receptor
was shown, and there is deliberately no identifier for a gene, a protein or a species.
antigen.gene and antigen.species are the peptide’s provenance - a real query axis, and not an
identity - so they are absent from every id here and from the row key vdjdb diff uses. A correction
to either is a changed cell on the same record, never a record retired and another allocated.
Level |
Identifies |
Id |
Where it appears |
|---|---|---|---|
clonotype |
one receptor chain |
|
|
clone |
one alpha/beta pair |
|
|
pMHC |
one peptide as presented |
|
|
epitope |
the peptide alone |
|
|
record |
one curated line |
|
|
Counts on the 2026-09 build: 187,935 clonotypes over 286,047 chain rows, 82,266 clones over the 93,294 records reporting two chains, 2,364 pMHCs, 2,118 epitopes and 192,753 records.
What each one keys on#
Level |
Key |
|---|---|
clonotype |
|
clone |
the record’s two |
pMHC |
|
epitope |
|
record |
the record natural key: the complex-information columns, the id fields, and |
mhc.class is not in the pMHC key, because the alleles determine it: keying on it as well gives the
same 2,364 distinct values.
A record reporting one chain has no clone, and clone_id is the empty string rather than a null.
Empty is the only missing marker in VDJdb, so a null in a string column would be the ambiguity that
several shipped bugs came from. 99,459 of 192,753 records are in this state.
Derived and allocated#
Four of the five ids are derived: the id is a hash of its own key and of nothing else. One,
record_id, is allocated from a counter against a registry.
The difference answers the question a curator asks first, which is what happens when a chunk is added or removed. For a derived id, nothing:
the order
chunks/is read in cannot reach it;adding a chunk names that chunk’s new clonotypes and changes no existing id;
removing a chunk retires only the ids nothing else supported;
two hosts, two core counts and two release dates produce the same ids.
A counter has none of those properties. record_id is allocated anyway, because its purpose is to
survive a content change: a curator fixing a CDR3 typo has to keep the record’s id, and no hash of
the content can do that. Every other level keys on content with no separate existence, so a change
there is not an amendment but a different clonotype.
One consequence: exactly one registry is consulted during a build, the record registry, and the four derived levels need no history in order to be correct.
The hash#
sha256 over the key fields joined by \x1f, truncated to 16 hex digits, with the level’s prefix in
front. The separator is the unit separator, so a field containing a tab cannot forge a key boundary.
sha256 and not a library hash function. polars’ Expr.hash is xxhash and its output is not specified
across versions, so a dependency upgrade could renumber every clonotype with nothing failing
anywhere. The algorithm has to be one this repository owns, and
tests/unit/test_identity_levels.py freezes one id against a literal so it cannot move.
Truncating to 16 digits is a size decision rather than a security one. At 187,935 clonotypes the chance of one collision is around 1 in 10⁹, and the invariants below fail a build rather than letting one corrupt a join. Measured cost on the current build: 83 ms to hash the 187,935 distinct clonotype keys and 7 ms to join them back onto 286,047 chain rows.
Lifecycle#
An id that vanishes is the failure a consumer cannot diagnose: a reference that used to resolve returns nothing, and nothing distinguishes a correction from a deletion. So every level keeps a lifecycle row, and the row holds only what a build cannot recompute.
Field |
Meaning |
|---|---|
|
the identifier |
|
|
|
|
|
the release tag the id first appeared in |
|
the newest release carrying it, frozen at retirement |
|
the id that took over, where the retirement was an amendment |
uv run vdjdb identity resolve CT3efa364e7f84aac9 --lifecycle identity-lifecycle.tsv
id CT3efa364e7f84aac9
level clonotype
state active
first_release v2026.09.1
last_release v2026.09.1
replaced_by
Retirement is per release, not per build. A curation branch may add a clonotype and remove it again before anything ships, and neither event is a lifecycle event: only the release job writes these rows. That is also why a curation pull request never touches them.
An id that comes back is the same id. A derived id is the hash of its key, so the same key
returning has to produce the same id. It goes back to active keeping its original first_release,
and vdjdb identity diff reports it under returned rather than under added, because the two mean
different things to a reader: one has been published before and one has not.
The lifecycle table is a release asset rather than a committed file, at 15.9 MB for 274,683 derived ids. The record registry is separate and larger, at 72.7 MB, because it carries each record’s natural key and content hash. The retired rows alone are committed, since they are the part a consumer needs and the part that is otherwise lost.
A build with no lifecycle file still runs and produces exactly the same derived ids. Only the history is unknown, which is the correct state for a fork and for a first build.
Invariants#
vdjdb identity check asserts six of the seven, exits 1 on any failure, and runs in build.yml
after the assembly step.
# |
Invariant |
Needs |
|---|---|---|
1 |
no id is shared by two distinct keys |
the build |
2 |
every stored id equals the hash of the key on its own row, recomputed |
the build |
3 |
a permuted chunk order, and a chunk added then removed, change no id |
two builds |
4 |
every |
the previous registry |
5 |
no id carries a prefix belonging to no level, and no published level empties silently |
the previous lifecycle |
6 |
every id referenced resolves: a clone covers exactly two chains of different |
the build |
7 |
|
the previous tables |
Invariant 3 is a property of two builds rather than of one directory, so it is a unit test over a
fixture in tests/unit/test_identity_levels.py and costs milliseconds instead of two full builds.
Invariants 2 and 5 together are what rules out reuse. Reuse would be an id resolving to a different key than the one it was retired under, and invariant 2 recomputes every id from the key beside it, so an id pointing anywhere else fails. An id merely reappearing is not reuse.
Checks needing history are skipped rather than failed when the history is absent.
uv run vdjdb identity check --tables out/tables
uv run vdjdb identity check --tables out/tables --previous identity-lifecycle.tsv \
--previous-tables previous/tables
uv run vdjdb identity diff identity-lifecycle.tsv --tables out/tables --release v2026.10.1
uv run vdjdb identity lifecycle --release v2026.10.1 --tables out/tables \
--previous identity-lifecycle.tsv --out out/release/identity-lifecycle.tsv
MHC promiscuity is an annotation, never part of a key#
An epitope is often presented by several alleles, and the curated allele is not always the one that binds it best. Both facts belong in the database; neither belongs in a key.
pmhc_id keys on the allele as curated. A prediction is a moving target: mhcmatch ships new
weights, and if a predicted allele were in the key then a model upgrade would renumber pMHC ids and
break every external reference while the build passed. Promiscuity is therefore measured into columns
on restriction, each row carrying the mhcmatch version that produced them, so re-running under a
new model rewrites those columns and no id.
Legacy structure ids#
TCR_hash is the identifier linking a record to a generated structure. It is produced by the legacy
recipe and nothing else: sha256 over a fixed field list, empty unless every required field is present
(vdjdb.assemble.master.add_tcr_hash). That recipe reproduces the values in the released vdjdb.txt
exactly, which is why it is preserved rather than redefined, and it is never recomputed under a new
scheme, renamed, or derived from the levels above. Structure evidence keys on it.
The identifiers on this page are additive, so a consumer holding a TCR_hash is never asked to
migrate, and invariant 7 fails the build if one moves on a clonotype whose key did not.