Migrating the build onto arda 2.36.0 + vdjtools 4.8.0#
2026-09-30. Everything the build does to annotate a junction is now one call in a library. This
note says which call, what it returns, and which of this repository’s modules it replaces. The
libraries’ own design note is docs/junction_pipeline.md in antigenomics/vdjtools.
Why#
Five of this build’s modules re-implement, coordinate, or second-guess work the libraries do:
annotate/cdr3fix.py (199 lines), annotate/junction.py (203), annotate/segments.py (185),
annotate/dgene.py (62) and curate/anchors.py (388) — 1,037 lines that call arda and vdjtools
stage by stage, hold the coordinate conversions between them, and in curate/anchors.py compute a
junction repair that is then reported and never applied (#711). One library call replaces the
annotation part of all five. What stays here is curation: which records to flag, and what a curator
does about a contradicted call.
annotate/segments.py exists because “arda.cdr3fix repairs a junction against a named germline and
never proposes one”. arda 2.36.0 proposes one, which is why the floor below is 2.36.0 and not
2.33.0.
The one call#
from vdjtools.model import annotate_junctions
out = annotate_junctions(
keys["cdr3"].to_list(), # junction space: Cys104 .. Phe/Trp118 inclusive
keys["v"].to_list(),
keys["j"].to_list(),
species=keys["species"].to_list(), # per row; VDJdb names are accepted as-is
)
One row out per row in, in input order, so it joins positionally. A record the model cannot
explain is present with nulls — never dropped, never an exception. Deduplicate to distinct
(species, cdr3, v, j) keys first, as the build already does.
What comes back#
column |
replaces |
notes |
|---|---|---|
|
|
the repaired junction; |
|
|
confirmed or re-called; feeds |
|
— |
every allele the junction cannot separate, chosen one first |
|
|
which side the submission left blank and the junction named |
|
|
residues; VDJdb’s |
|
|
nucleotides — no |
|
|
|
|
|
the inferred nucleotide junction and its Pgen |
|
|
use as |
|
|
where that gene was placed, 1-based closed in junction space |
|
— |
the residues whose codons the D touches, recomputed from the nt bounds |
|
|
the N regions either side of the D, sliced from the same bounds |
Five things that change in the output, and why#
1. The junction repair is applied, not reported (#711). curate/anchors.py computes a repair and
writes it to out/reports/anchors.tsv; 261 of 1,037 flagged chains carry a germline-supported repair
and every one ships unrepaired. cdr3_repaired is the repair, already gated: measured over all
187,488 keys of the 2026-06-03 release, arda agrees with that release’s shipped junction on
99.8352 %, reproduces 4,331 of its 4,499 repairs, and ships 677 non-canonical junctions
against the release’s 715. Applying cdr3_repaired is not the ungated “apply the proposal” that
#711 warns against — the 2,842 rewrites the release never ships are down to 141, of which 35 touch an
already-canonical junction.
2. good means “the junction is well formed”, and a germline disagreement no longer contradicts
it. A residue differing from germline inside the templated run is flagged mismatch and left
exactly as submitted, because a curation error and an allele IMGT does not record are
indistinguishable from one junction. Treat mismatch as a proofreading queue (it is the #681 class),
and impossible as the record whose anchors cannot be satisfied.
3. d.inferred comes from d_call, and every row that can be drawn has coordinates. Naming the
D and placing it are separate questions and one estimator answers each: the recombination model names
the gene, the aligner places it greedily and ungated. Measured on 4,000 real human TRB rearrangements
whose D and DStart/DEnd come from the nucleotide sequence:
gene right, of all rows |
has coordinates |
|
|---|---|---|
E-value-gated alignment chooses and places |
47.93 % |
55.75 % |
gated alignment, model posterior where it declines |
71.40 % |
55.75 % |
today’s |
69.67 % |
— |
model names, greedy alignment places |
74.35 % |
99.80 % |
⛔ arda.dpost does not come back, in either library. The posterior this build reads today is
dominated by a group-by over a call the pipeline already makes — 69.67 % against 74.33 %, 56.3 µs per
key against 30.1 — and it needs a fitted prior table and a per-locus tempering constant that the
replacement does not. annotate/dgene.py is replaced by two columns, not re-pointed at a new module.
So fields.py’s comment about d.posterior needs rewriting rather than renaming: the number now
comes from the model’s own scenario weights, normalised over D genes.
⚠ The one thing that gets worse is per-row positional precision: d_start_nt is exact on 64.71 %
of correctly-called rows against 66.67 % under the gate. It is exact on 1,922 rows rather than
1,278, because it answers 3,992 rather than 2,230. For a view drawing V/N/D/N/J that is the trade
to take.
4. The model set is chosen per row, not configured. model_source="auto" runs OLGA’s bundled
fit first and arda’s on whatever it left unexplained. Neither alone is best: OLGA’s is better
calibrated and cheaper but declines rows outright, arda’s answers everything and is thinner. On
4,000 real human rearrangements the chain wins every column — TRB nucleotide-exact
14.40 / 17.32 / 17.32 % and D gene 72.58 / 74.08 / 74.35 % for arda / OLGA / the chain, and
TRA keeps 4,000 of 4,000 nucleotide junctions where OLGA alone loses 143. Leave the default
alone unless you are reproducing a published number against one named fit.
⚠ And the chain’s last rung is not a bundled fit at all — it is a germline scaffold, which is what makes the non-human records answer. A fitted model exists for human and mouse and for nothing else, so every rhesus record in VDJdb used to come back empty: 1,457 keys answered zero times, invisible inside a single corpus-wide total. Coverage over the 192,726 curation keys, per species, measured on arda 2.36.0 — the floor this page declares:
species |
locus |
keys |
nucleotide junction |
D gene |
D coordinates |
|---|---|---|---|---|---|
HomoSapiens |
TRB |
115,829 |
115,819 |
115,819 |
113,232 |
HomoSapiens |
TRA |
58,281 |
58,203 |
— |
— |
MusMusculus |
TRB |
9,015 |
8,805 |
8,805 |
8,694 |
MusMusculus |
TRA |
8,140 |
8,140 |
— |
— |
MacacaMulatta |
TRB |
1,383 |
1,379 |
1,379 |
787 |
MacacaMulatta |
TRA |
74 |
73 |
— |
— |
HomoSapiens |
TRD |
4 |
4 |
4 |
4 |
192,423 of 192,726 (99.84 %), and no species or locus at zero. TRA and TRD are VJ loci — there
is no D to find. ⛔ Break any coverage check in the build down BY SPECIES: one percentage over
this corpus hid an entire organism answering nothing. ⚠ The keys column is as of arda 2.36.0 —
the 461 keys naming neither V nor J had no locus before and now count under the one they resolve
to, so a per-locus denominator is not comparable to an older run’s.
⚠ Every accuracy percentage on this page is human TRB. The truth set behind them
(isalgo/airr_control) carries TRA and TRB and no immunoglobulin, so none of them may be quoted for
IGH — and the per-species table above is coverage, not accuracy.
5. The nucleotides are the authority and the amino-acid bounds are recomputed from them.
v_end_nt, j_start_nt, d_start_nt and d_end_nt are all read off the inferred nucleotide
junction; v_end, j_start, d_start_aa and d_end_aa follow from those. A view showing both
alphabets therefore cannot draw them disagreeing, and there is no ceil(nt/3) conversion left in
this repository.
Cost#
Measured on this repository’s own chunk files: 192,726 distinct keys in 22.3 s, one process, 116 µs per key. 190,199 get a nucleotide junction, 124,022 of the 124,489 TRB keys get a D gene and 121,336 of those get coordinates; the 65,107 TRA keys have no D to find. The nucleotide inference dominates; naming the D costs ~30 µs and placing it ~1.4 µs.
Compare what it replaces: posterior_d alone was 15.96 s over 119,034 keys via a Python row loop
(35.8 % of the build), the nucleotide stage ran as four processes of a since-retired
infer-nt subcommand over contiguous slices, and fix_cdr3 was 7.71 s. One call, one process, is now faster than any of them.
⛔ Do not wrap the call in a pool. Every stage is already batched: one markup_batch, then one
native threaded infer_nt_batch and one best_aa_scenarios_batch per (organism, locus) on a model
loaded once, then a thin Python loop over arda’s C++ aligner. A pool around it would re-import the
libraries and re-load the models per worker.
Version floors#
"arda-mapper>=2.36.0", # cdr3fix repair policy, v_alts/j_alts, blank-call AND blank-locus proposal, map_d_junction(v_end=, j_start=)
"vdjtools>=4.8.0", # annotate_junctions
arda.dpost is gone in 2.33.0 and does not reappear in vdjtools, so annotate/dgene.py’s
from arda.dpost import posterior_d breaks on upgrade — that is the intended failure, not a
surprise, and the fix is to read d_call / d_posterior rather than to re-point the import.
arda markup --d-posterior and --d-prior are gone with it.
Suggested order#
Bump both floors and delete
annotate/dgene.py, readingd_call/d_posteriorfrom the one call instead. The build runs again at this point.Replace
annotate/junction.py’sinferwith thecdr3_nt/pgen/d_*columns. ItsSPECIESmap and its Pgen provenance comment are worth keeping as documentation of what the numbers mean.Replace
annotate/cdr3fix.pywith the stage-1 columns. Keep whatever maps them onto VDJdb’scdr3fixJSON key names —Cdr3Markup.to_cdr3fix()in arda still emits that object key-for-key.Delete
annotate/segments.pyand readproposedinstead. Of the 3,130 blank-call keys, 2,669 name one side and 461 name neither. arda 2.34.0 resolved the one-sided ones (2,504good, with both boundaries placed) and refused the rest for want of a locus; 2.36.0 proposes the locus too, so all 461 get one and 459 come backgood. The proposed locus agrees with thecdr3.alpha/cdr3.betacolumn the record was filed under on 457 of 461 (99.13 %); all four disagreements areCACD…DKLIF, TRDV2’s own anchor and TRDJ1’s own ending, in a schema with no δ column. An unresolvable call —TRBVnope*01— is still refused rather than proposed for, deliberately: a submission that names something wrong is a defect for a curator, not a gap for the junction to fill.Then #711:
curate/anchors.pykeeps its classification (which is curation) and drops its repair computation (which is annotation), and the build shipscdr3_repaired.Re-run
vdjdb diffagainstreference.zipand read the junction-column changes against the table in §”Five things that change” above.