Motif clustering algorithms, parameters and scorecards#
Which algorithm turns a set of candidate clonotypes into motif clusters, what each one’s parameters
do, and what each one measures at. docs/denoising.md defines what to optimise and this document
catalogues what is available to optimise over, so read docs/denoising.md first: a scorecard here is
read against its §7.1 two-stage rule.
Rendered on the documentation site from this file by myst-parser. It stays Markdown rather than becoming .rst, because it is cited by section number from ROADMAP.md, from the package docstrings and from docs/tuning/, and converting it would break those references.
Status key: shipped - the default in TUNED · wired - implemented, tested, measured, not
default · rejected - measured and ruled out, kept so the ruling stays checkable ·
measured - scored on the same instruments from a script under docs/tuning/, not in the package.
Every number below comes from docs/tuning/scorecard.tsv, 252 configurations scored through one
harness on one cohort, and the markdown is generated from it by docs/tuning/report.py, so no figure
in this document is transcribed by hand. §11 has the commands.
The cohort is chunks/ as of 2026-09-26, and the legacy bar below is that cohort’s. The released
cluster_members.txt never changes, but the corpus it is scored against does, so the bar moves when a
chunk merges: three merges on 2026-09-28 took TRA’s Q from 0.1691 to 0.1668, its purity from 0.8658
to 0.8641 and its precision from 0.8567 to 0.8548, and TRB’s Q from 0.4433 to 0.4391. The numbers
here are not restated for that, because every configuration in the scorecard was scored against the
same bar in the same run and re-stating one column would break the comparison the document exists to
make. The current values of every axis for the released files, the latest release and the build are
measured on every build and recorded in rules/motif_metrics.tsv, which is what a regression is gated
against (vdjdb motif-metrics).
Percolation is recorded there and not gated, and it is not a score. Per epitope it is the share
of the epitope’s clustered clonotypes that sit in its largest cluster, so it describes the shape of
the repertoire before it describes the clustering. An epitope with one prominent motif (A*02 GIL)
should percolate: one cluster is the right answer, and a high value is correct. An epitope with two
or three prominent motifs sits lower, and a featureless or bulged one with many disjoint motifs
(A*02 NLV) sits low because that is what its repertoire looks like. So neither a higher nor a lower
median is better or worse by itself. What percolation can indicate is a spurious edge fusing two
different motifs, and that is a question about the purity of the largest cluster, which purity,
precision and Q measure and vdjdb motif-metrics gates. The median over epitopes also moves by
a whole rank step whenever the set of covered epitopes changes, which it did on 2026-09-29 and
2026-10-01 with every quality axis flat or up. Wherever this document says an epitope “percolates” it
means the largest cluster holds most of it, which is a measurement of the epitope and, only when its
purity falls, a defect.
1. Connected components - shipped (TCRNET)#
vdjdb.motifs.cluster._components. The transitive closure of the similarity graph: two clonotypes
are in one cluster if a path of within-scope edges joins them. Parameters: none, so there is
nothing to tune and nothing to overfit; it is what the legacy Rmd did.
Knob |
Where |
Default |
Effect |
|---|---|---|---|
|
|
|
edge definition: |
|
|
|
enrichment threshold that decides which clonotypes enter the graph |
|
|
|
smallest surviving cluster, post length-split |
The failure mode to watch is percolation: one spurious edge merges two motifs, and at scope 2 a chain
of such edges can span an epitope. Whether a large largest cluster is that failure or the epitope’s
own single motif is read per epitope, from the cluster’s purity (see the note on percolation at the top of this document). At scope 1,0,0,1 the largest component holds ≥ 90 % of a TRA epitope’s
clustered clonotypes in 22 of 118 epitopes; median percolation 0.3958.
2. CPM Leiden - rejected#
vdjdb.motifs.cluster._leiden, reached by setting tcrnet.TUNED[gene]["resolution"] to a positive
number. It runs inside each connected component, so it subdivides what §1 found: it never merges
across components and never recovers a clonotype components missed.
2.1 Parameters#
igraph.Graph(n=n, edges=edges).community_leiden(
objective_function="CPM", resolution=resolution, n_iterations=-1)
Parameter |
Value |
Notes |
|---|---|---|
|
|
Not modularity: modularity’s resolution is scaled by the graph’s total edge count, so one number is a different filter on a 30-clonotype epitope than on a 30,000 one. CPM’s resolution is an absolute internal-density threshold and transfers across epitopes. |
|
|
A community is kept while its internal edge density exceeds \(\gamma\). \(\gamma = 0\) reproduces connected components exactly, asserted in |
|
|
Iterate to convergence rather than a fixed count, which would make the answer depend on how long it ran. |
seed |
|
Leiden randomises node visit order and the refinement step, and python-igraph draws from Python’s own global RNG rather than accepting a seed argument, so the seed must be set on the |
CPM maximises
over partitions, where \(e_c\) is the edge count inside community \(c\), \(n_c\) its size, and \(\gamma\) the resolution. A community survives while its internal density \(e_c / \binom{n_c}{2}\) exceeds \(\gamma\), so \(\gamma\) reads directly as the minimum density of a motif.
2.2 The resolution ladder#
Human, scope 1,0,0,1, p 0.01, min_cluster 5. res = None is connected components, the shipped
default and the first row of each table.
TRA (legacy bar: Q 0.1691, purity 0.8658, precision 0.8567):
resolution |
lift |
Q |
h |
parsimony |
purity |
precision |
retention |
cids |
clonotypes clustered |
|---|---|---|---|---|---|---|---|---|---|
None |
2.051 |
0.1798 |
0.9600 |
0.0992 |
0.8754 |
0.8679 |
0.2323 |
859 |
15,608 |
0.05 |
2.136 |
0.1218 |
0.9653 |
0.0650 |
0.8943 |
0.8841 |
0.2211 |
1,376 |
14,554 |
0.1 |
2.289 |
0.0929 |
0.9696 |
0.0488 |
0.8961 |
0.8853 |
0.2030 |
1,573 |
12,798 |
0.2 |
2.949 |
0.0538 |
0.9759 |
0.0276 |
0.8871 |
0.8747 |
0.1581 |
1,380 |
8,591 |
0.3 |
4.869 |
0.0259 |
0.9838 |
0.0131 |
0.8824 |
0.8560 |
0.1001 |
526 |
3,803 |
0.5 |
5.349 |
0.0233 |
0.9849 |
0.0118 |
0.8852 |
0.8576 |
0.0951 |
452 |
3,352 |
0.8 |
5.349 |
0.0233 |
0.9849 |
0.0118 |
0.8852 |
0.8576 |
0.0951 |
452 |
3,352 |
TRB (legacy bar: Q 0.4433, purity 0.9790, precision 0.9756):
resolution |
lift |
Q |
h |
parsimony |
purity |
precision |
retention |
cids |
clonotypes clustered |
|---|---|---|---|---|---|---|---|---|---|
None |
1.501 |
0.4428 |
0.9920 |
0.2850 |
0.9790 |
0.9761 |
0.3337 |
657 |
37,898 |
0.05 |
1.518 |
0.2243 |
0.9937 |
0.1264 |
0.9785 |
0.9758 |
0.3181 |
1,949 |
35,822 |
0.1 |
1.488 |
0.1429 |
0.9944 |
0.0770 |
0.9782 |
0.9751 |
0.3011 |
3,374 |
33,570 |
0.2 |
1.465 |
0.1067 |
0.9958 |
0.0564 |
0.9799 |
0.9774 |
0.2471 |
3,263 |
26,572 |
0.3 |
1.293 |
0.0879 |
0.9976 |
0.0460 |
0.9856 |
0.9769 |
0.1868 |
2,103 |
20,011 |
0.5 |
1.198 |
0.0854 |
0.9977 |
0.0446 |
0.9856 |
0.9767 |
0.1782 |
1,967 |
19,168 |
0.8 |
1.198 |
0.0854 |
0.9977 |
0.0446 |
0.9856 |
0.9767 |
0.1782 |
1,967 |
19,168 |
Lift is on the full cohort here, so these rows are comparable to each other but not to
denoising.md’s non-display figures.
2.3 Ladder results#
Homogeneity rises while parsimony falls: on TRA h goes 0.9600 → 0.9849 as parsimony goes
0.0992 → 0.0118, an 8.4-fold fall, and on TRB h goes 0.9920 → 0.9977 against parsimony
0.2850 → 0.0446, a 6.4-fold fall. That is shattering: every cluster gets purer because every
cluster gets smaller, and past some point they are small enough to be uninformative. Q is the only
one of the seven columns that records it: purity and precision barely move (TRA precision falls,
0.8679 → 0.8576) and lift rises, so a criterion built on purity, precision or lift alone ranks the
degenerate end top.
On TRA lift rises at the cost of coverage. Lift 2.051 → 5.349 is a 2.6× improvement in how often a clustered clonotype is independently replicated, but the clustered set shrinks from 15,608 clonotypes to 3,352, because Leiden keeps the densest cores and discards the rest. Dropping 78 % of what components clustered is not a trade §7.1 can accept.
On TRB there is no lift gain: the maximum is 1.518 at resolution 0.05, a 1.1 % gain over components’
1.501, and it falls to 1.198 by 0.5 while Q drops to a fifth of its value. No resolution improves
TRB over connected components.
The ladder saturates between 0.5 and 0.8, where the two rows are bit-identical on both chains. Once
every surviving community is dense enough to clear \(\gamma\), raising \(\gamma\) further changes
nothing, so the ladder has a top, resolution = 0.8 is not more aggressive than 0.5, and
extending the sweep upward measures the same partition twice.
No Leiden cell is admissible on either chain. resolution stays in TUNED, set to None, and the
implementation and its tests stay, so the measurement can be re-run when the corpus changes.
3. DBSCAN - shipped (TCREMP)#
vdjdb.motifs.tcremp.cluster_labels. Density clustering in the prototype-distance embedding, run
per epitope at one radius estimated on the chain’s pooled geometry.
Parameter |
Where |
Default |
Effect |
|---|---|---|---|
|
|
TRA 1.8, TRB 1.15 × |
the radius, and the primary knob |
|
|
|
the TCREMP paper’s Table-1 value, and the smallest number that can be a cluster |
|
|
|
PCA dimensions before clustering |
|
|
|
post length-split filter |
eps is pooled and DBSCAN is per-epitope. Re-estimating the radius on a per-epitope \(n\) of 30–300
is the degenerate regime a knee fails in (§5), so the geometry is estimated once per chain and the
clustering runs per epitope (ROADMAP §8.4).
The frontier is monotone, so tuning is a bisection rather than a search: lift falls as eps widens
and Q rises, so the best admissible cell is the smallest eps clearing the Q bar, and both
chains therefore sit on the boundary by construction, TRB clearing by 0.0004 and TRA’s next step
down missing by 0.0010. min_cluster 5 dominates 3 on lift, purity and precision at every coef
measured, so there is no trade-off between them: clusters of 3–4 contribute low-confidence members,
not signal.
4. HDBSCAN - rejected#
vdjdb.motifs.tcremp.hdbscan_labels, sklearn.cluster.HDBSCAN, with no dependency beyond the
sklearn already used for the PCA and DBSCAN.
4.1 Motivation#
DBSCAN commits to one radius for every epitope, and VDJdb’s epitopes do not share a density: a
display-selected library yields thousands of receptors one substitution apart by construction (29,688
of 192,753 records) while a 30-record epitope is sparse. HDBSCAN condenses a cluster hierarchy over
mutual-reachability distance and selects by stability, so there is no global radius to be wrong, and
it folds min_samples and the min_cluster post-filter into one min_cluster_size.
4.2 Parameters#
Parameter |
Swept |
Effect |
|---|---|---|
|
3, 5, 10, 20 |
smallest cluster the hierarchy will emit, and HDBSCAN’s primary knob |
|
|
excess-of-mass takes the most stable cut (parsimonious); |
|
2, 5 |
how conservative the core-distance estimate is; higher declares more points noise |
|
not swept |
the anti-shatter knob, and not usable; see below |
|
|
left alone; |
cluster_selection_epsilon is broken in sklearn 1.9.1. On a 12-cell probe at 0.5 and 1.0 × the
chain’s mean 1-NN distance, 5 cells raised
TypeError: only 0-dimensional arrays can be converted to Python scalars from
sklearn/cluster/_hdbscan/_tree.pyx:588 (traverse_upwards), and the cells that did run returned
bit-identical results to epsilon = 0.0, so the one parameter that would directly counter
shattering is unavailable and the leaf rows below cannot be repaired.
4.3 Scorecard - 56 cells per chain, human, min_cluster fixed at 5#
min_cluster_size × {3, 5, 8, 10, 15, 20, 30, 50}, cluster_selection_method × {eom, leaf},
min_samples × {2, 3, 5, 10}, cells with min_samples > min_cluster_size skipped: 56 per chain.
The full grid with every instrument is docs/tuning/scorecard.tsv; the extremes are below.
TRA - bar: Q 0.1691, purity 0.8658, precision 0.8567, coverage 103 epitopes.
Each row is the grid’s extreme on one column; p is Q’s parsimony term.
config |
lift |
f1 |
q |
p |
purity |
retention |
epitopes |
perc_med |
|---|---|---|---|---|---|---|---|---|
leaf mcs5 ms5 |
1.397 |
0.0799 |
0.1876 |
0.104 |
0.9002 |
0.3344 |
114 |
0.377 |
eom mcs3 ms2 |
1.389 |
0.0789 |
0.1700 |
0.093 |
0.9061 |
0.3244 |
114 |
0.355 |
eom mcs5 ms2 |
1.257 |
0.0734 |
0.2616 |
0.152 |
0.9017 |
0.4366 |
116 |
0.306 |
leaf mcs50 ms5 |
1.143 |
0.0669 |
0.3467 |
0.213 |
0.8869 |
0.4304 |
108 |
0.241 |
eom mcs30 ms5 |
0.921 |
0.0548 |
0.6489 |
0.507 |
0.8833 |
0.5623 |
111 |
0.286 |
TRB - bar: Q 0.4433, purity 0.9790, precision 0.9756, coverage 103 epitopes.
config |
lift |
f1 |
q |
p |
purity |
retention |
epitopes |
perc_med |
|---|---|---|---|---|---|---|---|---|
leaf mcs3 ms3 |
1.538 |
0.0884 |
0.1346 |
0.072 |
0.9337 |
0.3478 |
147 |
0.455 |
leaf mcs15 ms3 |
1.222 |
0.0743 |
0.3331 |
0.201 |
0.9326 |
0.6270 |
130 |
0.215 |
eom mcs5 ms2 |
1.218 |
0.0744 |
0.4427 |
0.287 |
0.9389 |
0.7240 |
165 |
0.357 |
eom mcs20 ms3 |
1.119 |
0.0689 |
0.6254 |
0.464 |
0.9433 |
0.8326 |
125 |
0.244 |
eom mcs30 ms2 |
1.055 |
0.0649 |
0.6397 |
0.479 |
0.9424 |
0.8336 |
115 |
0.247 |
4.4 Scorecard results#
Under the legacy-relative bar, 51 of 56 TRA cells are admissible and 0 of 56 TRB cells are: on TRB purity tops out at 0.9433 against a bar of 0.9790, so every cell regresses it by 3.6 points. The 51 admissible TRA cells are not on their own a result, since §8 shows the do-nothing partition is admissible on TRA too.
Under the absolute floor of 0.94 the two chains swap places, 0 of 56 on TRA and 11 of 56 on TRB.
TRA’s highest HDBSCAN purity is 0.9061, so the floor excludes the grid there, and §8.1 shows it
excludes every other algorithm’s TRA cells with it. Both verdicts are columns in scorecard.tsv
(admissible, admissible_legacy_bar); neither replaces the other.
The eom / leaf axis behaves as theory predicts: leaf raises lift and lowers parsimony, TRB
leaf parsimony spanning 0.068–0.286 against eom’s 0.148–0.500, while eom raises parsimony,
retention and coverage and gives up lift, and neither end is admissible on TRB.
HDBSCAN’s strength is coverage and structure rather than selectivity. TRB eom mcs5 ms2 reaches 165
of 178 epitopes at retention 0.7240 and median percolation 0.357, against legacy’s 103 epitopes,
0.3218 and 0.769, and it leads on every axis except the one being optimised, its lift being 1.218
against shipped TCREMP’s 3.586.
4.5 Under the absolute purity floor of 0.94#
The floor §7 lists as an open question was measured; §8.1 gives the measurement behind choosing 0.94 on both chains rather than 0.93.
gene |
purity floor |
HDBSCAN cells admissible of 56 |
best admissible cell |
lift |
F1 |
retention |
epitopes |
|---|---|---|---|---|---|---|---|
TRA |
legacy, 0.8658 |
51 |
|
1.397 |
0.0799 |
0.3344 |
114 |
TRA |
0.9400 |
0 |
– |
– |
– |
– |
– |
TRB |
legacy, 0.9790 |
0 |
– |
– |
– |
– |
– |
TRB |
0.9300 |
25 |
|
1.194 |
0.0731 |
0.7367 |
160 |
TRB |
0.9400 |
11 |
|
1.155 |
0.0709 |
0.7698 |
155 |
At a floor of 0.94 HDBSCAN becomes admissible on TRB and loses stage 2 by 3.1×, lift 1.155 against the shipped TCREMP’s 3.586 and F1 0.0709 against 0.1888, so the floor can be set without changing what ships. Rewriting stage 2 as well would give a different database: 155 of 178 epitopes carrying motifs instead of 105, at 77.0 % retention instead of 37.5 %, with a third of legacy’s percolation, and a clustered clonotype 16 % more likely than chance to have been seen by a second laboratory rather than 259 % more likely.
On TRA the same floor admits nothing, HDBSCAN included, since its best purity is 0.9061 against the floor’s 0.94. TRA’s admissible reading stays the legacy-relative one, and §8.1 measures why no absolute floor can replace it there. HDBSCAN is rejected on the objective rather than on the bar; it stays wired and tested, and §8 records what the bar can and cannot do.
5. Choosing eps by a knee - wired, not in the operating path#
vdjdb.motifs.tcremp.knee. The published TCREMP protocol picks eps from the knee of the sorted
k-distance curve; this repo does not.
The reference implementation is degenerate at production scale: Kneedle with a degree-10 polynomial
fit on pooled human TRB, 112,983 clonotypes, returns knee index 1, fraction 0.000. A degree-10 fit
over 113,000 points oscillates, and the reported knee is the first oscillation rather than a feature
of the data, which puts eps below the data entirely and retention falls to a few per cent.
knee states Kneedle’s difference curve directly, with no polynomial:
for a sorted curve \(c\) (signs reversed for a convex curve). Three guards, each against an observed failure:
Guard |
Default |
What it catches |
|---|---|---|
resample onto a fixed grid |
1,000 points |
the degree-10 oscillation; makes the knee a property of the curve’s shape, measured invariant at 0.799 / 0.800 / 0.800 for one elbow at \(n\) = 300 / 30,000 / 300,000 |
|
0.05 |
a curve with no knee. \(\max(y-x)\) is exactly 0 for a straight line, so the statistic that locates a knee is the same one that reports there is none, and no separate test is needed |
|
0.05 / 0.95 |
a knee pinned to an end. At the bottom |
strength is scale-invariant by construction, so it is comparable across chains with different
embedding scales (TRA mean 1-NN 9.66, TRB 9.19). What runs instead is
eps = coef × mean(1st-NN distance), with coef fitted against independent replication, which is
not a knee method and must not be described as one (ROADMAP §8.3).
6. The per-epitope view#
vdjdb.validate.motif_bench.per_epitope, counted on clonotypes so a clonotype reported by forty
papers does not weigh forty times in its own epitope’s retention, and reported over epitopes, the
unit curation acts on. A pooled retention of 0.33 is a different database depending on whether every
epitope is a third clustered or a third of them are fully clustered and the rest not at all; on TRB
it is closer to the second, and no pooled column reports that.
6.1 The shipped clusterings, per epitope#
TRA - 118 epitopes with ≥ 30 records:
method |
epitopes covered |
none at all |
retention median |
retention q75 |
≥90 % percolated |
percolation median |
clonotypes clustered |
cids |
epitopes with lift > 1 |
|---|---|---|---|---|---|---|---|---|---|
legacy |
103 |
15 |
0.0647 |
0.1603 |
24 |
0.4717 |
15,063 |
2,214 |
29 |
TCRNET |
103 |
15 |
0.0706 |
0.1624 |
22 |
0.3958 |
15,480 |
2,173 |
33 |
TCREMP |
104 |
14 |
0.0850 |
0.2115 |
23 |
0.4834 |
17,380 |
2,926 |
34 |
TRB - 178 epitopes:
method |
epitopes covered |
none at all |
retention median |
retention q75 |
≥90 % percolated |
percolation median |
clonotypes clustered |
cids |
epitopes with lift > 1 |
|---|---|---|---|---|---|---|---|---|---|
legacy |
103 |
75 |
0.0089 |
0.0941 |
42 |
0.7692 |
37,210 |
1,074 |
42 |
TCRNET |
109 |
69 |
0.0142 |
0.0941 |
41 |
0.7059 |
37,109 |
1,051 |
43 |
TCREMP |
105 |
73 |
0.0128 |
0.1139 |
40 |
0.6667 |
42,802 |
1,156 |
42 |
6.2 Per-epitope results#
TRB is a two-population database and TRA is not. 75 of 178 TRB epitopes get no motif from the shipped annotation, and the median epitope has 0.9 % of its clonotypes clustered against a pooled retention of 32 %, the pooled figure coming from a minority of large, deeply-sampled epitopes such as GILGFVFTL, NLVPMVATV and YLQPRTFLL. On TRA the median epitope has 6.5 % clustered against a pooled 21 %, still skewed but far less bimodal. A single coverage number for TRB describes about sixty epitopes and implies 178.
On TRB the structural failure these scorecards measure is percolation, not shattering. The median TRB epitope has 77 % of its clustered clonotypes in a single cluster (many TRB epitopes do have one prominent motif, so the figure is read with purity, not on its own), 42 of 178 are ≥ 90 % collapsed, and both rewrites reduce it (TCRNET to 0.7059, TCREMP to 0.6667) without removing it. It is the target for a future motif change, and the reason §2 measured Leiden: subdividing percolated components addresses it, but the resolutions that subdivide enough also shatter the rest.
Coverage and pooled lift move in opposite directions, and that is the finding that changed the
shipped parameters. See denoising.md §7.1, where TCREMP TRB at coef 1.15 maximised pooled lift at
4.490 and covered 88 of 178 epitopes, fifteen fewer than the annotation it replaces; at 1.55 it
covers 105 and leads legacy by 25.6 % on lift.
lift > 1 per epitope is the coverage-weighted read. Of the epitopes where a local base rate exists
at all, legacy beats chance in 42 of them on TRB and 29 on TRA, while the rewrites reach 43 (TCRNET)
and 42 (TCREMP) on TRB and 33 and 34 on TRA.
6.3 The report#
out/reports/motifs_per_epitope.tsv, 888 rows, one per (gene, method, epitope), written by the
build and uploaded as a CI artifact, never shipped in a bundle (outputs.md §5). Columns:
clonotypes, clustered, retention, clusters, largest_cluster, mean_cluster_size,
singleton_clusters, percolation, replicated, tp, precision, lift.
lift is null, not zero, where an epitope has no replicated clonotype: there is no base rate to
lift against, and averaging a zero there would put the per-epitope mean below the pooled figure. 47
of 118 TRA epitopes and 60 of 178 TRB epitopes have a defined lift.
7. Summary#
Algorithm |
Status |
TRA lift |
TRA F1 |
TRB lift |
TRB F1 |
TRB epitopes |
Distinguishing property |
|---|---|---|---|---|---|---|---|
the shipped 2026-06-03 annotation |
baseline |
1.703 |
0.0932 |
2.855 |
0.1489 |
103 / 178 |
the source of every bar |
connected components |
shipped, TCRNET |
1.977 |
0.1086 |
3.194 |
0.1669 |
109 |
no parameters to overfit |
CPM Leiden |
rejected |
– |
– |
– |
– |
– |
lift 2.6× on TRA at parsimony ÷ 8.4 |
DBSCAN |
shipped, TCREMP |
1.855 |
0.1032 |
3.586 |
0.1888 |
105 |
monotone frontier; tuning is a bisection |
HDBSCAN |
rejected |
1.397 |
0.0799 |
1.155 |
0.0709 |
155 at floor 0.94 |
coverage 160/178 at retention 0.74 |
Lumbermark |
measured |
– |
– |
1.148 |
0.0706 |
175 |
percolation 0.224 vs legacy’s 0.769 |
TCRNET gate + Lumbermark |
measured |
3.618 |
0.1702 |
4.089 |
0.1943 |
82, below the floor |
best objective measured; fails coverage |
A dash is where no admissible configuration exists, so the rejected rows show the best admissible
number or nothing. The two measured rows show their distinguishing configuration instead; for
Lumbermark that is not one of the six TRB cells that clear the 0.94 floor, which sit at lift
1.009–1.069 (§9.3). Across 252 configurations and six algorithms, under both the legacy-relative bar
and the absolute 0.94 floor, nothing admissible beats TCREMP at coef 1.55 on TRB or TCRNET on TRA,
and §§8–9 record the supporting measurements.
What would change a verdict#
~~A bar that is absolute rather than relative.~~ Settled at 0.94 on both chains, and it changes nothing that ships; measured in §4.5 and §8.1. On TRB the floor takes stage 1 from 1 admissible configuration of 122 to 23, namely 11
hdbscan-eom, 6 Lumbermark, 3 wider DBSCAN radii and 2hybrid-len-all, and stage 2 rejects all 22 newcomers, because the shippedcoef1.55 still has the highest admissible lift at 3.585 and the selected radius does not move on either chain. On TRA the floor admits nothing at all, including both shipped methods; §8.1 measures the window as empty by 0.0030 of purity and states what guards TRA instead.~~Per-epitope coverage entering the criterion.~~ Settled as a fourth admissibility axis (
denoising.md§7.1). Measured, TCREMP TRB atcoef1.15 covered 88 of 178 against legacy’s 103; the shipped TRB radius moved to 1.55 as a result.A gate with the recruited set’s coverage and the enriched set’s selectivity. §9 shows the gate and the partition are separable, and that the gate is what moves the objective, TRA F1 0.1086 → 0.1702 on the enriched gate alone; nothing measured keeps both, and that is the open question this document leaves.
cluster_selection_epsilonbecoming usable in a later sklearn, the one knob that would give HDBSCAN’sleaflift without its shattering.
A rejected algorithm stays implemented and tested, so the measurement that made it non-default can be re-run.
8. Instrument audit - the do-nothing partition#
Every sweep above includes one extra row, trivial: one cluster per epitope, every clonotype in it,
nothing excluded. It is not an algorithm but the reading an instrument gives when no clustering has
happened at all, and §7.1’s admissibility axes do not survive it.
gene |
axis |
do-nothing partition |
legacy bar |
does the bar exclude it? |
|---|---|---|---|---|
TRA |
|
0.7940 |
0.1691 |
no |
TRA |
purity |
0.9204 |
0.8658 |
no |
TRA |
precision |
0.9237 |
0.8567 |
no |
TRA |
epitope coverage |
118 |
103 |
no |
TRA |
lift |
1.0000 |
1.7027 |
yes |
TRA |
F1 |
0.0600 |
0.0932 |
yes |
TRB |
|
0.8632 |
0.4433 |
no |
TRB |
purity |
0.9343 |
0.9790 |
yes |
TRB |
precision |
0.9455 |
0.9756 |
yes |
TRB |
epitope coverage |
178 |
103 |
no |
TRB |
lift |
1.0000 |
2.8552 |
yes |
TRB |
F1 |
0.0622 |
0.1489 |
yes |
vdjdb.validate.motif_bench.trivial_members builds it, so the audit re-runs from the package and
not from a sweep script. Two forms of it exist, and every shipped clustering is split by CDR3 length
before it reaches a release (§0), so the split form is the one a bar has to clear; the unsplit form,
one cluster per epitope, scores higher on Q (0.9208 on TRA, 0.9654 on TRB) and lower on purity
(0.9144, 0.9253). A floor of 0.93 would exclude the unsplit partition and admit the split one; the
table above is the split form.
Three consequences:
Q is not an admissibility axis, since doing nothing scores Q 0.7940 on TRA and 0.8632 on TRB,
higher than any clustering measured here on either chain across 252 configurations. Its parsimony
term charges for shattering, and the partition that shatters least is the one with one cluster per
class, so Q discriminates between clusterings of comparable retention and says nothing across
retentions. It stays in the scorecard as a shatter guard and comes out of the admissibility rule.
Epitope coverage behaves the same way. Coverage is maximised by claiming every clonotype, so a coverage floor is a floor against narrowness rather than a quality bar; it still serves the purpose §7.1 added it for, stopping stage 2 from raising lift by abandoning epitopes, but it is not evidence that a clustering is good.
On TRA the four-axis rule already admits the do-nothing partition. Legacy TRA purity is 0.8658,
below the 0.9204 that doing nothing scores, so trivial clears all four axes today, and what
excludes it is stage 2: its lift is exactly 1.000 by construction and its F1 sits at the floor.
8.0 It also beats the shipped clustering at reproducing the shipped clustering#
The tables a release ships were never compared against the previous release’s, because vdjdb diff
keys cluster_members.txt on cid and a cid is <species>.<chain>.<epitope>.<n> with n a
position in a sorted list, so one renumbered cluster reads as the entire file replaced. Since
2026-09-29 vdjdb motif-metrics measures the partition instead
(vdjdb.compare.clustering.agreement): of every clonotype the 2026-06-03 release clustered, what
fraction of its cluster-mates does a source still give it?
gene |
source |
clonotypes the release clustered |
also placed here |
cluster-mates kept |
co-clustered pairs kept |
ARI |
|---|---|---|---|---|---|---|
TRA |
TCRNET (shipped) |
15,016 |
12,011 |
0.7024 |
0.7194 |
0.9945 |
TRA |
TCREMP |
15,016 |
8,530 |
0.3430 |
0.4039 |
0.8161 |
TRA |
the do-nothing partition |
15,016 |
13,108 |
0.7490 |
0.7345 |
0.1872 |
TRB |
TCRNET (shipped) |
36,906 |
35,095 |
0.9224 |
0.9952 |
0.9999 |
TRB |
TCREMP |
36,906 |
32,047 |
0.7921 |
0.9966 |
0.9977 |
TRB |
the do-nothing partition |
36,906 |
35,649 |
0.9394 |
0.9991 |
0.9939 |
So doing nothing reproduces the released clustering’s neighbourhoods better than either shipped
method does, on both chains. That is a statement about the released clustering, not about ours: it
puts 19,971 of its 36,906 TRB clonotypes in one cluster, H.B.SLLMWITQV.1, which is 54 % of the
chain’s clonotypes and 199,410,435 of the file’s 210,634,576 co-clustered pairs, 94.7 %. A
partition that groups every clonotype of an epitope reproduces that blob exactly.
Two consequences, the same shape as the three above.
Agreement with the last release is not an admissibility axis. It is gated against its own
committed baseline, not against latest and not against trivial, for the reason every other row of
§8 gives: the do-nothing partition clears it.
A pair-weighted agreement is a measurement of one cluster. partition_pairs_preserved and the
adjusted Rand index read 0.9991 and 0.9939 for doing nothing on TRB, while the per-clonotype figure
separates it from TCRNET by 1.7 points. Both are recorded, because the blob is worth watching, and
only the per-clonotype one is gated.
8.1 The purity floor - 0.94 on both chains#
On TRB the legacy purity bar of 0.9790 is the one axis the do-nothing partition fails, and relaxing it to 0.93 removes that, because doing nothing scores 0.9343, 0.0043 above the proposed floor. An absolute floor has to clear the measured do-nothing purity per chain rather than be a round number chosen in advance.
The floor is 0.94 on both chains (docs/tuning/sweeps.py, BAR), the lowest round number above the
do-nothing purity of either chain, 0.9204 on TRA and 0.9343 on TRB. Every cell in scorecard.tsv
has two verdicts, admissible under that floor and admissible_legacy_bar under the
legacy-relative one, so neither reading has to be recomputed to be checked.
TRB purity floor |
HDBSCAN cells admissible of 56 |
do-nothing partition admitted? |
|---|---|---|
legacy, 0.9790 |
0 |
no |
0.9300 |
25 |
yes |
0.9343 |
23 |
yes (the floor equals its purity) |
0.9400 |
11 |
no |
Floor sweep and admissible window, 122 configurations per chain#
Both tables are generated by docs/tuning/report.py from scorecard.tsv; a cell is admissible at an
absolute floor when purity and precision clear it and Q and epitope coverage clear legacy’s.
gene |
purity floor |
cells admissible |
of |
do-nothing admitted? |
best by lift |
lift |
F1 |
retention |
epitopes |
|---|---|---|---|---|---|---|---|---|---|
TRA |
0.8658 (legacy) |
70 |
122 |
yes |
|
1.854 |
0.1032 |
0.2509 |
104 |
TRA |
0.9204 |
0 |
122 |
yes |
– |
– |
– |
– |
– |
TRA |
0.9300 |
0 |
122 |
no |
– |
– |
– |
– |
– |
TRA |
0.9400 |
0 |
122 |
no |
– |
– |
– |
– |
– |
TRB |
0.9300 |
37 |
122 |
yes |
|
3.585 |
0.1887 |
0.3747 |
105 |
TRB |
0.9343 |
35 |
122 |
yes |
|
3.585 |
0.1887 |
0.3747 |
105 |
TRB |
0.9400 |
23 |
122 |
no |
|
3.585 |
0.1887 |
0.3747 |
105 |
TRB |
0.9790 (legacy) |
1 |
122 |
no |
|
3.585 |
0.1887 |
0.3747 |
105 |
The floor takes TRB from one admissible configuration to twenty-three, and stage 2 selects the same
one: under the legacy bar exactly one candidate clears stage 1, dbscan coef 1.55, which is what
ships, and at 0.94 it is joined by 11 hdbscan-eom cells, 6 Lumbermark, 3 wider DBSCAN radii and 2
hybrid-len-all, each of which loses on lift. The relaxation adds 22 candidates and changes nothing.
A floor is usable only if it excludes the do-nothing partition and admits something, so its ceiling
is the highest purity reached by any configuration that already clears the other two axes, Q and
epitope coverage, both legacy-relative.
gene |
do-nothing purity |
cells clearing |
highest purity among them |
window |
width |
|---|---|---|---|---|---|
TRA |
0.9204 |
74 of 122 |
0.9174 |
empty |
−0.0030 |
TRB |
0.9343 |
37 of 122 |
0.9829 |
[0.9343, 0.9829] |
0.0486 |
On TRB the window is 0.0486 of purity wide and 0.94 sits inside it with margin at both ends, 5.4× the [0.9343, 0.9433] window HDBSCAN alone offers, because the incumbent DBSCAN radius sits at 0.9829 and the ceiling is set by the best method rather than by the method being tested against it.
On TRA the window is empty by 0.0030: across six algorithms and 122 configurations, the highest
purity any cell reaches while clearing Q and coverage is 0.9174 (hybrid-len-all mcs5 M1), below
the 0.9204 that doing nothing scores. Every absolute purity floor that excludes the do-nothing
partition on TRA therefore also excludes every configuration that clears the other two axes,
including both shipped methods, whose TRA purity is 0.8754 (TCRNET) and 0.9104 (TCREMP), and at 0.94
specifically 0 of 122 TRA cells are admissible. Purity is therefore not the axis that separates
signal from nothing on TRA, and §8’s third consequence identifies stage 2 as the axis that is; TRA
keeps admissible_legacy_bar as its usable reading, and the 0.94 column records what an absolute
floor would cost there.
9. Lumbermark and the TCRNET-gated hybrid - measured, not shipped#
9.1 Motivation#
§6.2 measured the defect: TRB’s dominant failure is percolation, not shattering, with the median
epitope putting 0.706 of its clustered clonotypes in one cluster under the shipped TCRNET and 41 of
178 epitopes ≥90 % collapsed. DBSCAN percolates because one radius chains through density bridges;
Leiden (§2) and HDBSCAN leaf (§4) break the bridges by shattering, which is the other failure.
Lumbermark (Gagolewski 2026, arXiv:2604.07143, lumbermark on
PyPI, AGPL-3) addresses that gap. It cuts the k−1 longest edges of the M-mutual-reachability
minimum spanning tree, HDBSCAN’s own internal structure, subject to every resulting component being
at least min_cluster_size, after pruning the tree’s leaves. Mutual reachability pulls low-density
points apart so a bridge becomes a long edge, the leaf pruning puts the cut on a protruding edge
rather than an outlier’s stalk, and the size floor stops the cut sequence shattering. It beats Genie
and HDBSCAN on 61 reference datasets by adjusted Rand index. Requesting k = n−1 makes
min_cluster_size the only granularity knob: the algorithm saturates and warns rather than failing,
so there is no cluster count to choose.
It has no noise label: it partitions every point, and the paper lists outlier detection as future work. VDJdb already supplies that half, since TCRNET’s enrichment test against a matched background repertoire is a statistical noise model rather than a density heuristic. The variants measured each differ from their comparator in exactly one factor:
variant |
vertex set |
partition |
isolates |
|---|---|---|---|
|
everything |
Lumbermark |
the partition’s ceiling with no gate |
|
TCRNET enriched |
Lumbermark |
the gate, against |
|
TCRNET enriched + neighbours |
Lumbermark |
the partition, against shipped TCRNET |
|
as above, within CDR3-length strata |
Lumbermark |
the emit path’s length split |
hybrid-recruited and the shipped TCRNET partition the same vertices, so the only difference between
them is the partition step.
9.2 Measurements#
Lumbermark gives the lowest percolation measured in this document. TRB mcs10 M1: median
per-epitope percolation 0.224 against legacy’s 0.769 and shipped TCREMP’s 0.667, at retention 0.8345
over 175 of 178 epitopes, purity 0.9445, which is the highest purity of any configuration that
retains four-fifths of the cohort (23 of the 115 TRB cells clear retention 0.80, and it leads all of
them), above HDBSCAN’s best of 0.9433 and above the do-nothing partition’s 0.9343. That applies to
the inclusive regime only, since DBSCAN reaches purity 0.9920 on TRB at retention 0.1189, the trade
this document is about, and Lumbermark sits at the other end with lift 1.148.
The gated hybrid gives the highest selectivity measured in this document. On TRA, hybrid mcs3 M5
reaches F1 0.1702 and lift 3.618, against the shipped TCRNET’s F1 0.1086 and lift 1.977 and legacy’s
F1 0.0932 and lift 1.703, which is +57 % F1 over the best shipped TRA configuration on the objective
§11.1 tunes against, and the largest single improvement measured on either chain.
gene |
configuration |
lift |
F1 |
|
purity |
retention |
epitopes |
percolation |
|---|---|---|---|---|---|---|---|---|
TRA |
legacy release |
1.703 |
0.0932 |
0.1691 |
0.8658 |
0.2105 |
103 |
0.472 |
TRA |
shipped TCRNET |
1.977 |
0.1086 |
0.1798 |
0.8754 |
0.2323 |
103 |
0.396 |
TRA |
|
2.368 |
0.1244 |
0.0759 |
0.9025 |
0.1816 |
97 |
0.375 |
TRA |
|
3.618 |
0.1702 |
0.0452 |
0.9040 |
0.1255 |
86 |
0.500 |
TRB |
legacy release |
2.855 |
0.1489 |
0.4433 |
0.9790 |
0.3218 |
103 |
0.769 |
TRB |
shipped TCRNET |
3.194 |
0.1669 |
0.4428 |
0.9790 |
0.3337 |
109 |
0.706 |
TRB |
shipped TCREMP |
3.586 |
0.1888 |
0.4947 |
0.9829 |
0.3745 |
105 |
0.667 |
TRB |
|
3.234 |
0.1649 |
0.1110 |
0.9786 |
0.2864 |
102 |
0.500 |
TRB |
|
4.089 |
0.1943 |
0.1388 |
0.9816 |
0.2847 |
82 |
0.527 |
Three readings:
The gate moves the objective and the partition moves percolation. hybrid-recruited against the
shipped TCRNET - same vertices, Lumbermark instead of connected components - moves TRB percolation
0.706 → 0.500 and leaves F1 flat (0.1669 → 0.1649) while Q falls 0.4428 → 0.1110 on 4,702 cids
against 657, so swapping the partition alone de-percolates and costs parsimony, and tightening the
gate from recruited to enriched is what moves lift and F1.
Both hybrids fail the coverage axis, by 17 epitopes on TRA and 23 on TRB, which is the narrow corner §7.1’s coverage floor exists to exclude.
Clustering inside CDR3-length strata does not recover the coverage. hybrid-len was run because the
emit path splits every cluster by length before applying min_cluster (§0), so a cluster of 8 across
three lengths is removed entirely at the floor of 5 and clustering within a stratum should have
recovered it. It recovers some coverage and gives back the same amount of objective, TRA 86 → 93
epitopes at F1 0.1702 → 0.1436 and TRB 82 → 91 epitopes at F1 0.1943 → 0.1832.
9.3 Status#
Neither is shipped, and the reason is coverage rather than quality. The gated hybrid is inadmissible
under both rules on both chains, 0 of 15 cells each for hybrid and hybrid-recruited under the
legacy bar and under the 0.94 floor alike, because the enriched gate abandons epitopes: the two
variants that clear the objective cover 82–97, and hybrid-recruited tops out at 102 on both chains,
one epitope short of the 103 floor.
Lumbermark is the exception, on TRB and under the new floor only. Six of its 15 TRB cells, mcs20
and mcs50 at every M, clear all four axes at 0.94, with purity 0.9405–0.9439, Q 0.4800–0.5681
against legacy’s 0.4433, and 124–156 epitopes against a floor of 103. Stage 2 then rejects all six at
lift 1.009–1.069 against shipped TCREMP’s 3.586, the same shape as HDBSCAN, where the floor admits
the inclusive regime and the objective rejects it.
The gate and the partition are separable and were measured separately. The shipped methods each couple one representation to one partition, TCRNET to connected components on the Hamming-1 graph and TCREMP to DBSCAN in the embedding, while the hybrid shows the two choices are independent, that VDJdb’s best noise model is its enrichment test rather than any density heuristic, and that the embedding’s MST is where percolation can be fixed. A method that kept the recruited gate’s coverage and the enriched gate’s selectivity would beat everything in this document, and nothing measured here does both.
One thing deliberately not tried: admitting clusters individually by the within-stratum
sum-of-hypergeometrics null in vdjdb.validate.noise. Admitting a cluster because its members are
enriched for independent replication, and then scoring the result by independent-study lift, fits on
the test statistic, so the apparent improvement would not be informative.
11. Figures#
Rendered from the committed tables by gnuplot; regenerate with the commands in each script header.
figure |
script |
shows |
|---|---|---|
|
|
all six instruments against retention, every configuration, both chains |
|
|
the per-epitope retention and percolation ECDFs behind §6 |
The first figure shows §8’s measurements: lift and F1 fall monotonically as retention
rises, Q and coverage rise toward the grey do-nothing marker at retention 1.0, and purity is the
only panel where that marker sits below the measured clusterings.
Reproducing the tables#
uv run python docs/tuning/sweeps.py all # -> out/reports/tuning/*.csv (~50 min)
uv run python docs/tuning/report.py # -> docs/tuning/*.tsv, tables.md
gnuplot -e "chain='TRB'" docs/tuning/tuning.gp
docs/tuning/scorecard.tsv is the full 252-row measurement, one row per configuration and every
instrument, a committed and reviewed input to this document, refreshed by its own pull request and
never written by a build (hard rule 9).