Migrating the build onto arda 2.36.0 + vdjtools 4.8.0#

2026-09-30. Everything the build does to annotate a junction is now one call in a library. This note says which call, what it returns, and which of this repository’s modules it replaces. The libraries’ own design note is docs/junction_pipeline.md in antigenomics/vdjtools.

Why#

Five of this build’s modules re-implement, coordinate, or second-guess work the libraries do: annotate/cdr3fix.py (199 lines), annotate/junction.py (203), annotate/segments.py (185), annotate/dgene.py (62) and curate/anchors.py (388) — 1,037 lines that call arda and vdjtools stage by stage, hold the coordinate conversions between them, and in curate/anchors.py compute a junction repair that is then reported and never applied (#711). One library call replaces the annotation part of all five. What stays here is curation: which records to flag, and what a curator does about a contradicted call.

annotate/segments.py exists because “arda.cdr3fix repairs a junction against a named germline and never proposes one”. arda 2.36.0 proposes one, which is why the floor below is 2.36.0 and not 2.33.0.

The one call#

from vdjtools.model import annotate_junctions

out = annotate_junctions(
    keys["cdr3"].to_list(),          # junction space: Cys104 .. Phe/Trp118 inclusive
    keys["v"].to_list(),
    keys["j"].to_list(),
    species=keys["species"].to_list(),   # per row; VDJdb names are accepted as-is
)

One row out per row in, in input order, so it joins positionally. A record the model cannot explain is present with nulls — never dropped, never an exception. Deduplicate to distinct (species, cdr3, v, j) keys first, as the build already does.

What comes back#

column

replaces

notes

cdr3_repaired

annotate/cdr3fix.py

the repaired junction; cdr3_aa is the submission

v_call, j_call

annotate/cdr3fix.py

confirmed or re-called; feeds cdr3fix.vId/jId

v_alts, j_alts

—

every allele the junction cannot separate, chosen one first

proposed

annotate/segments.py

which side the submission left blank and the junction named

v_end, j_start

annotate/cdr3fix.py

residues; VDJdb’s vEnd / jStart

v_end_nt, j_start_nt

annotate/junction.py

nucleotides — no ceil(nt/3) conversion here any more

v_flags, j_flags, good

curate/anchors.py

mismatch is the curator’s list; impossible is a malformed junction

cdr3_nt, pgen

annotate/junction.py

the inferred nucleotide junction and its Pgen

d_call, d_posterior

annotate/dgene.py

use as d.inferred — the D GENE, named by the model, with its posterior

d_start_nt, d_end_nt

annotate/junction.py

where that gene was placed, 1-based closed in junction space

d_start_aa, d_end_aa

—

the residues whose codons the D touches, recomputed from the nt bounds

np1, np2

annotate/junction.py

the N regions either side of the D, sliced from the same bounds

Five things that change in the output, and why#

1. The junction repair is applied, not reported (#711). curate/anchors.py computes a repair and writes it to out/reports/anchors.tsv; 261 of 1,037 flagged chains carry a germline-supported repair and every one ships unrepaired. cdr3_repaired is the repair, already gated: measured over all 187,488 keys of the 2026-06-03 release, arda agrees with that release’s shipped junction on 99.8352 %, reproduces 4,331 of its 4,499 repairs, and ships 677 non-canonical junctions against the release’s 715. Applying cdr3_repaired is not the ungated “apply the proposal” that #711 warns against — the 2,842 rewrites the release never ships are down to 141, of which 35 touch an already-canonical junction.

2. good means “the junction is well formed”, and a germline disagreement no longer contradicts it. A residue differing from germline inside the templated run is flagged mismatch and left exactly as submitted, because a curation error and an allele IMGT does not record are indistinguishable from one junction. Treat mismatch as a proofreading queue (it is the #681 class), and impossible as the record whose anchors cannot be satisfied.

3. d.inferred comes from d_call, and every row that can be drawn has coordinates. Naming the D and placing it are separate questions and one estimator answers each: the recombination model names the gene, the aligner places it greedily and ungated. Measured on 4,000 real human TRB rearrangements whose D and DStart/DEnd come from the nucleotide sequence:

gene right, of all rows

has coordinates

E-value-gated alignment chooses and places

47.93 %

55.75 %

gated alignment, model posterior where it declines

71.40 %

55.75 %

today’s d.inferred (the length-and-prior posterior)

69.67 %

—

model names, greedy alignment places

74.35 %

99.80 %

⛔ arda.dpost does not come back, in either library. The posterior this build reads today is dominated by a group-by over a call the pipeline already makes — 69.67 % against 74.33 %, 56.3 µs per key against 30.1 — and it needs a fitted prior table and a per-locus tempering constant that the replacement does not. annotate/dgene.py is replaced by two columns, not re-pointed at a new module. So fields.py’s comment about d.posterior needs rewriting rather than renaming: the number now comes from the model’s own scenario weights, normalised over D genes.

⚠ The one thing that gets worse is per-row positional precision: d_start_nt is exact on 64.71 % of correctly-called rows against 66.67 % under the gate. It is exact on 1,922 rows rather than 1,278, because it answers 3,992 rather than 2,230. For a view drawing V/N/D/N/J that is the trade to take.

4. The model set is chosen per row, not configured. model_source="auto" runs OLGA’s bundled fit first and arda’s on whatever it left unexplained. Neither alone is best: OLGA’s is better calibrated and cheaper but declines rows outright, arda’s answers everything and is thinner. On 4,000 real human rearrangements the chain wins every column — TRB nucleotide-exact 14.40 / 17.32 / 17.32 % and D gene 72.58 / 74.08 / 74.35 % for arda / OLGA / the chain, and TRA keeps 4,000 of 4,000 nucleotide junctions where OLGA alone loses 143. Leave the default alone unless you are reproducing a published number against one named fit.

⚠ And the chain’s last rung is not a bundled fit at all — it is a germline scaffold, which is what makes the non-human records answer. A fitted model exists for human and mouse and for nothing else, so every rhesus record in VDJdb used to come back empty: 1,457 keys answered zero times, invisible inside a single corpus-wide total. Coverage over the 192,726 curation keys, per species, measured on arda 2.36.0 — the floor this page declares:

species

locus

keys

nucleotide junction

D gene

D coordinates

HomoSapiens

TRB

115,829

115,819

115,819

113,232

HomoSapiens

TRA

58,281

58,203

—

—

MusMusculus

TRB

9,015

8,805

8,805

8,694

MusMusculus

TRA

8,140

8,140

—

—

MacacaMulatta

TRB

1,383

1,379

1,379

787

MacacaMulatta

TRA

74

73

—

—

HomoSapiens

TRD

4

4

4

4

192,423 of 192,726 (99.84 %), and no species or locus at zero. TRA and TRD are VJ loci — there is no D to find. ⛔ Break any coverage check in the build down BY SPECIES: one percentage over this corpus hid an entire organism answering nothing. ⚠ The keys column is as of arda 2.36.0 — the 461 keys naming neither V nor J had no locus before and now count under the one they resolve to, so a per-locus denominator is not comparable to an older run’s.

⚠ Every accuracy percentage on this page is human TRB. The truth set behind them (isalgo/airr_control) carries TRA and TRB and no immunoglobulin, so none of them may be quoted for IGH — and the per-species table above is coverage, not accuracy.

5. The nucleotides are the authority and the amino-acid bounds are recomputed from them. v_end_nt, j_start_nt, d_start_nt and d_end_nt are all read off the inferred nucleotide junction; v_end, j_start, d_start_aa and d_end_aa follow from those. A view showing both alphabets therefore cannot draw them disagreeing, and there is no ceil(nt/3) conversion left in this repository.

Cost#

Measured on this repository’s own chunk files: 192,726 distinct keys in 22.3 s, one process, 116 µs per key. 190,199 get a nucleotide junction, 124,022 of the 124,489 TRB keys get a D gene and 121,336 of those get coordinates; the 65,107 TRA keys have no D to find. The nucleotide inference dominates; naming the D costs ~30 µs and placing it ~1.4 µs.

Compare what it replaces: posterior_d alone was 15.96 s over 119,034 keys via a Python row loop (35.8 % of the build), the nucleotide stage ran as four processes of a since-retired infer-nt subcommand over contiguous slices, and fix_cdr3 was 7.71 s. One call, one process, is now faster than any of them.

⛔ Do not wrap the call in a pool. Every stage is already batched: one markup_batch, then one native threaded infer_nt_batch and one best_aa_scenarios_batch per (organism, locus) on a model loaded once, then a thin Python loop over arda’s C++ aligner. A pool around it would re-import the libraries and re-load the models per worker.

Version floors#

"arda-mapper>=2.36.0",   # cdr3fix repair policy, v_alts/j_alts, blank-call AND blank-locus proposal, map_d_junction(v_end=, j_start=)
"vdjtools>=4.8.0",       # annotate_junctions

arda.dpost is gone in 2.33.0 and does not reappear in vdjtools, so annotate/dgene.py’s from arda.dpost import posterior_d breaks on upgrade — that is the intended failure, not a surprise, and the fix is to read d_call / d_posterior rather than to re-point the import. arda markup --d-posterior and --d-prior are gone with it.

Suggested order#

  1. Bump both floors and delete annotate/dgene.py, reading d_call / d_posterior from the one call instead. The build runs again at this point.

  2. Replace annotate/junction.py’s infer with the cdr3_nt / pgen / d_* columns. Its SPECIES map and its Pgen provenance comment are worth keeping as documentation of what the numbers mean.

  3. Replace annotate/cdr3fix.py with the stage-1 columns. Keep whatever maps them onto VDJdb’s cdr3fix JSON key names — Cdr3Markup.to_cdr3fix() in arda still emits that object key-for-key.

  4. Delete annotate/segments.py and read proposed instead. Of the 3,130 blank-call keys, 2,669 name one side and 461 name neither. arda 2.34.0 resolved the one-sided ones (2,504 good, with both boundaries placed) and refused the rest for want of a locus; 2.36.0 proposes the locus too, so all 461 get one and 459 come back good. The proposed locus agrees with the cdr3.alpha / cdr3.beta column the record was filed under on 457 of 461 (99.13 %); all four disagreements are CACD…DKLIF, TRDV2’s own anchor and TRDJ1’s own ending, in a schema with no δ column. An unresolvable call — TRBVnope*01 — is still refused rather than proposed for, deliberately: a submission that names something wrong is a defect for a curator, not a gap for the junction to fill.

  5. Then #711: curate/anchors.py keeps its classification (which is curation) and drops its repair computation (which is annotation), and the build ships cdr3_repaired.

  6. Re-run vdjdb diff against reference.zip and read the junction-column changes against the table in §”Five things that change” above.