Building and releasing#

The build is a uv-managed Python package. It reads chunks/, patches/ and proofreading/, and writes everything else. Nothing computed is stored between builds: every output is recomputed from chunks/ on every run.

Local build#

uv sync
uv run vdjdb qc                               # chunk validation, fail-fast
uv run vdjdb build --out out/                 # the definitive tables, then every projection
uv run vdjdb make legacy --tables out/tables  # the legacy files, from the tables that shipped
uv run vdjdb convert airr --tables out/tables # AIRR Rearrangement + Reactivity
uv run vdjdb motifs --tables out/tables       # TCRNET + TCREMP
uv run vdjdb summary --legacy out/legacy      # the dashboard, rendered offline and checked
uv run vdjdb identity check --tables out/tables # the identifier invariants
uv run vdjdb corpus build --tables out/tables # the reference corpus: tf-idf over 12 token families
uv run vdjdb diff <reference.zip> out/legacy  # compare against a released zip
uv run vdjdb release --tag vYYYY.MM.P --dry-run  # the three bundles, without touching the tree
uv run pytest -q

Output goes to out/, not build/: build/ is gitignored as a Python packaging convention.

vdjdb release writes latest-version.txt, which is tracked, so pass --dry-run when the question is the bundle shape rather than a release you are about to publish (#707). The flag leaves the repository copy alone and still puts the tag’s URL in the bundles, so the members are what a real run would ship. Without it, a made-up tag leaves the tree naming a download that 404s - the defect ROADMAP.md §3.2 records as having shipped once already, from the other direction.

Optional dependency groups: motifs for the motif stage, docs for this site, test for the suite, tuning for the clustering bake-off under docs/tuning/.

The junction-nucleotide stage#

annotate.junction.add_junction_nt was 87.2 % of the assembly stage - 407.64 s of 467.72 s on a 4-vCPU runner over 192,793 records - because vdjtools.model.infer_nt wraps a native DP in per-row Python and the wrapper, not the DP, was the cost. That profile is what antigenomics/vdjtools#181 was opened on, and vdjtools 4.5 answers it with infer_nt_batch.

So the stage is one batched call per (species, locus) over the distinct (species, gene, cdr3, v, j) keys - 187,055 of them rather than every row, which is rule 4’s deduplication - and nothing wraps it. infer_nt_batch releases the GIL and partitions the batch across its own kernel threads; a pool of our own would oversubscribe the machine and read as “batching did not help”, which is hard rule 3 and section 0e of CLAUDE.md both.

Measured on 3,000 distinct human TRB keys from the corpus, 16 cores: 1.115 ms/key serial against 0.106 ms batched, 10.5x, and all 3,000 nucleotide sequences identical. End to end on the 4-vCPU runner the stage goes 407.64 s to 108.54 s and the assembly step 467.72 s to 166.54 s, so it is 65.2 % of that step rather than 87.2 %; on a 16-core laptop it is 12.44 s and vdjdb build is 45.9 s wall. The runner wins less because infer_nt_batch defaults to hardware_concurrency - 2 threads, which is two there. tests/unit/test_junction.py asserts both halves of that, because each catches a different failure - identity catches a batch call that is not the same computation, and the ratio catches a regression to the loop or a batch call that loops internally, neither of which changes an answer.

Every inferred sequence encodes the junction it came from, and that is by construction rather than by luck. The DP enumerates (V, delV) x (J, delJ) x (D, delD, position) and picks the best codon assignment within each scenario, so a scenario that cannot spell the given residues has probability zero and is never a candidate; anything the model cannot encode comes back null rather than wrong. Probed on human TRB: a stop codon, an X, a Z, a one- or two-residue junction, an empty string and a true CDR3 with its anchors stripped are all declined. The one input that survives with a difference is a lower-case junction, where the nucleotides are right and the comparison is case-sensitive - and vdjdb qc rejects a residue outside the 20 upper-case letters, with zero such chains in the corpus. Measured on the built corpus: 263,437 of 285,989 chains carry an inferred cdr3nt, 0 mismatches, 0 whose length is not exactly three nucleotides per residue, gated by tests/release/test_tables_contract.py. That check translates the whole column in one threaded native call, vdjtools._core.translate_junctions - 0.023 s against 0.284 s for vdjtools.model.translate in a Python loop, 12.3x, identical on every row.

It previously ran as four worker processes over contiguous parquet slices of the key set, each an ordinary invocation of a subcommand that existed only to be that worker. The subcommand, src/vdjdb/__main__.py, the slice arithmetic and the worker-count argument are all gone: one batched call has no worker count, so rule 7’s “never let worker count change the answer” holds by construction rather than by a tiling test.

Comparing against a release#

vdjdb diff compares a candidate build against a released bundle in three passes: the file set, two digests per file (raw and canonical, rows sorted), and a row-level classification that attributes every changed cell to a declared rule in rules/expected_diffs.toml.

Any cell difference not matched by a rule fails, and so does a rule that fires a different number of times than declared. Every rule therefore declares a measured row count rather than a description.

A correction to an identity column has no cell to attribute: the row leaves the reference bucket and a different row arrives in the candidate one. Those two counts are declared per file in a [[row_delta]] block, with a note giving every reason they are what they are, and the run fails if either moves. A new chunk moves all of them, because its records are rows the reference cannot contain, so landing one means re-measuring the three declarations and extending each note with the chunk and its record count - the entry for PMID_18025130 is the worked example. vdjdb diff --report prints declared against measured per file and flags which one moved.

What the comparison is, and is not, the instrument for#

It runs on the assembled legacy zip, so the thing compared is the file a consumer downloads. Three of that bundle’s twelve members cannot be judged by keying their rows against the reference, and [measured_elsewhere] in rules/expected_diffs.toml names each one with the instrument that gates it instead. They are still read, digested, row-counted and printed; what they are exempt from is the requirement that every changed cell match a rule.

Member

Why keying it says nothing

What gates it

cluster_members.txt

cid is <species>.<chain>.<epitope>.<n> and n is a position in a sorted list, so one renumbered cluster relabels every cluster after it

vdjdb motif-metrics, 18 axes per chain, including partition_neighbours_preserved - of every clonotype the release clustered, the fraction of its cluster-mates this build still gives it. Column count and order by tests/release/test_reference_contract.py

motif_pwms.txt

a row is one PWM cell of one cluster, so it has no identity that survives a re-clustering

the same two

vdjdb_summary_embed.html

a fresh render every build

summary/check_summary.py: the ordered headings and tables, PNG dimensions decoded from the IHDR, ColorBrewer anchors, and SSIM against the last release

LICENSE

not a table

shipped verbatim from the repository root

Measured before those declarations existed: comparing every member of the legacy zip reported 101,877 unattributed cells, and every one was in a motif file or in latest-version.txt while the five legacy tables were fully attributed. The workflow’s answer had been --only naming those five, which is the same exemption with no reason recorded and no digest taken of the other seven.

A member the reference does not contain is declared in [members]. Today that is cluster_members_tcremp.txt and motif_pwms_tcremp.txt; an undeclared one still fails.

Canonical equality is the gate; raw equality is informational. The legacy pipeline iterated os.listdir("../chunks"), which is readdir order and therefore filesystem- and host-dependent, so the 2026-06-03 release cannot be reproduced byte-for-byte on another machine, including by the pipeline that produced it. The current reader sorts, and a release records its chunk order, so raw equality is achievable for releases built by this code.

Testing against older releases#

The comparison currently runs one build against one release, 2026-06-03. A single release cannot distinguish a rule that is correct from a rule that happens to fit that release.

There are 43 published releases, back to 2017-06-13, and each one is a matched pair: the inputs (chunks/, patches/, proofreading/ at that tag, and res/ for a tag older than the #658 retirement) and the outputs (the zip that shipped). Each pair is a regression test: take a git worktree at the tag and build it with the current code:

git worktree add /tmp/vdjdb-2023 2023-06-01      # the corpus as it was
cd /tmp/vdjdb-2023 && uv run --project <repo> vdjdb build --out out/
uv run vdjdb diff <2023-06-01.zip> out/legacy    # does today's code reproduce that release?

What a replay across releases adds:

  • Code changes separated from corpus changes. Every rule in expected_diffs.toml is currently declared against one release. A rule that fires the same way across 2021, 2023 and 2026 is a property of the code; one that fires only on 2026 is a property of today’s data.

  • Format drift that the current corpus no longer contains. The 2017–2019 chunks use header shapes, MHC spellings and nomenclature that later curation cleaned up, and those are the inputs a reader of an old release would hand back to the tool.

  • Dates for known defects. web.cdr3fix.unmp is wrong on 7,973 rows today; replaying it across releases says when it started.

Older releases were produced by older code with different column sets, so the comparison needs era-scoped rules rather than one global set: vdjdb.txt gained columns, the slim table changed shape, and the motif files did not exist at all before 2018. The first replay therefore produces a large diff that is mostly format difference, and the work is in classifying it.

Start with one release, 2023-06-01: recent enough to share most of the schema, old enough that a rule which only fits 2026 will not fit it.

dev and master cannot be deleted#

The repository has delete_branch_on_merge enabled, which is right for feature branches and wrong for the long-lived ones: a dev to master pull request has dev as its head, so merging it used to delete dev. That happened on PR #573 and again on PR #584. Nothing was lost except the ref - every commit stays reachable from master - but a contributor pulling in the window between the merge and the restore gets a confusing error, and the symptom does not look like a branch-protection question.

It is fixed by a ruleset rather than by turning the setting off, so feature branches still tidy themselves up:

Ruleset

Refs

Rules

Bypass actors

no-deletion-dev-master

refs/heads/dev, the default branch

deletion

none

protect-dev

refs/heads/dev

deletion, non_fast_forward, required_status_checks

org admin, admin, triage

vdjdb-master

the default branch

deletion, non_fast_forward, pull_request, required_status_checks

org admin

⚠ The bypass list is why protect-dev alone did not work. It already carried a deletion rule, but a merge runs as the person merging, and an admin’s always bypass covers the deletion too. A GitHub ruleset’s bypass list applies to the whole ruleset and cannot be set per rule, so the deletion rule needs its own ruleset with an empty bypass list. That is what no-deletion-dev-master is, and it holds: a git push origin --delete dev from an admin is now rejected with GH013.

protect-dev keeps its bypass on purpose, so a maintainer can still push a fix to dev directly when a status check is stuck. Deletion is the only thing nobody may do.

If dev is ever absent again, the restore is one command, and the commit is the merge’s second parent:

git push origin $(git rev-parse master^2):refs/heads/dev

Reproducibility#

Same inputs, same bytes: in another process, on another host, at another core count.

  • One seed, vdjdb.config.SEED, passed explicitly to every sampler, shuffler and clustering init. No module-scope random.seed(), no unseeded default.

  • Sort after anything unordered: a group_by without maintain_order=True, a set, a dict keyed on strings under PYTHONHASHSEED, a thread pool, os.listdir. An unstable sort over equal-labelled groups once made the release comparison report three different changed-row counts for one comparison.

  • Worker count never changes the answer. Split into as many big contiguous slices as there are workers, reassemble in slice order, never a pool of small tasks.

Committed inputs#

Three derived tables make the build offline and deterministic. All are committed, reviewed inputs, refreshed by their own pull request and never written by a build:

  • summary/reference_years.tsv - the publication year of every reference, so the dashboard does not call NCBI while it renders. Rebuilt by uv run vdjdb refs.

  • docs/tuning/scorecard.tsv - the clustering bake-off, 252 configurations scored through one harness. Rebuilt by docs/tuning/sweeps.py.

  • proofreading/epitope_promiscuity.tsv - which class I alleles can present each epitope, 13,510 (epitope, allele) pairs over 1,729 epitopes, of which 1,739 are pairings VDJdb records and 11,771 are alleles only the predictor reaches. Rebuilt by uv run vdjdb promiscuity, which loads mhcmatch’s reference panel from HuggingFace. It answers a query-side question - filter VDJdb to a donor’s HLA type and see every record their T cells could have raised - and makes no claim about the mhc.a a publication reported.

Neither is a cache: a cache is a stored answer to the question the build is currently asking, and these are inputs that arrive by review.