The shipped models#
mhcmatch rank scores a candidate with a fitted aggregate artifact, one per
(cls, species, mode), vendored under src/mhcmatch/data/. There is no fallback: asking
for a combination that was never fitted raises rather than scoring it with a neighbour’s
coefficients, which is the mistake the lookup exists to prevent.
This page is what you have. Every table on it is generated from the artifacts themselves on each docs build — the coefficients used to be typed into six pages and all six went stale together the first time the model was refitted, so nothing here is written by hand.
Three identifiers, and only one of them moves with the library#
model_id—mhc1.human.neoantigen. Which cell of the lookup this is.version— an integer, the model version. It moves when the specification changes: a term added, a column respecified, a population redefined.release— the dotted package version the fit was accepted in, stored rather than derived. A manuscript pins a fit while the library keeps moving underneath it, somhc1.human.neoantigen v12 (release 1.20.0)is a citation andmhcmatch 1.20.1is not.
mode is neoantigen on four and pathogen on four. It is a key rather than a covariate
because a tumour neoantigen and a pathogen epitope are two mechanisms, not two values of one
variable. All eight cells are fitted; the refusal branch stays, because a cell can
leave the registry again — a refit withdrawn, an artifact not vendored — and what has to survive
that is the honest refusal, not a table that happens to be complete today. mhcmatch models --all
prints all eight and marks an unfitted one --.
The four pathogen fits are one family and two term sets. All four are fitted on the IEDB
T-cell corpora, where a negative is a peptide that was assayed and did not respond, and all four
carry one global intercept rather than a grouping unit. What they do not share is the term list:
class I is fitted on binder and C_corpus_self, class II on binder and the
physicochemical pair, so the four are not comparable coefficient by coefficient. Read
fit.deposit and features on the artifact; infer neither from the mode.
At a glance#
model |
model version |
release |
terms |
rows |
immunogenic |
intercepts |
AUROC |
protocol |
|---|---|---|---|---|---|---|---|---|
|
12 |
1.20.0 |
9 |
339,595 |
594 |
7 per screen |
0.7094 |
leave-one-screen-out, mean |
|
6 |
1.20.0 |
9 |
921 |
379 |
61 per reference |
0.6335 |
in-sample, within reference |
|
2 |
1.20.0 |
6 |
1,112 |
656 |
157 per reference |
0.6020 |
in-sample, within reference |
|
4 |
1.20.0 |
6 |
468 |
177 |
30 per reference |
0.5741 |
in-sample, within reference |
|
2 |
1.20.0 |
2 |
16,790 |
7,002 |
1 per corpus, global |
0.5988 |
in-sample, pooled off the logit |
|
2 |
1.20.0 |
2 |
10,404 |
2,196 |
1 per corpus, global |
0.5561 |
in-sample, pooled off the logit |
|
2 |
1.20.0 |
3 |
7,946 |
5,148 |
1 per corpus, global |
0.5824 |
in-sample, pooled off the logit |
|
2 |
1.20.0 |
3 |
11,725 |
3,324 |
1 per corpus, global |
0.6446 |
in-sample, pooled off the logit |
The AUROC column is three different protocols and must not be read down. The human class-I
neoantigen fit spans seven independent screens, so it can hold one out whole and be scored on it;
that is the 0.7102. The three single-deposit neoantigen fits have no second screen to hold out,
so what they record is an in-sample within-reference figure. The four pathogen fits are
whole-corpus GLMs with one global intercept and no grouping unit at all — neither a screen to hold
out nor a per-reference intercept to exclude — so each reports in-sample, pooled off the logit
against its own prevalence, and those run from 0.2111 to 0.6479. Averaging the column, or ranking
the eight fits by it, compares three different questions.
What “in-sample, within reference” means#
For the three single-deposit fits it is a precise thing, and it is not the naive apparent number:
scored on the slope term \(X\beta\) alone — the 61 / 157 / 30 fitted per-reference intercepts are excluded from the score;
macro-averaged within reference, over the references carrying at least three of each class;
on the fitting rows, no fold and no holdout.
Both exclusions are load-bearing. Reading the same 921-row mouse class-I fit three ways:
score |
AUROC |
AUPRC |
|---|---|---|
slopes and the 61 fitted intercepts, pooled |
0.9267 |
0.8910 |
slopes only, pooled across references |
0.4771 |
0.3989 |
slopes only, within reference — what is recorded |
0.6335 |
0.5931 |
The 0.9267 is 61 free intercepts on 921 rows reproducing base rates that run from 0 % to 90 % between publications. The 0.4771 is the same slopes judged on an axis they carry no information about — which publication a row came from. Only the third is a statement about the model.
The same reasoning is why the per-reference intercept exists at all. Fitted against a single pooled intercept instead, every mouse class-I coefficient came out at or below zero: the slopes were spending themselves on the base rates.
mhc1.human.neoantigen#
The fit the manuscript pins. Seven human neoantigen screens, 339,595 rows, 594 immunogenic, 523 (patient, screen) bootstrap clusters. Nine terms in four blocks.
rows |
339,595 |
immunogenic |
594 (0.17 %) |
screens |
7 (GBM, HiTIDE, IEDB_neoag, ITSNdb, NCI, TESLA, VACCIMEL) |
bootstrap clusters |
523 |
ridge \(\tau\) |
0.25 |
bootstrap |
1000 resamples of (patient, screen) cluster |
deviance |
– |
BIC |
3,107.6 |
held out |
leave-one-screen-out |
corpus block |
k = 3, kernel |
block |
term |
coef |
boot sd |
z |
p |
95 % CI |
sign stability |
|---|---|---|---|---|---|---|---|
|
|
+0.7596 |
0.1004 |
+6.65 |
2.9e-11 |
\([+0.584, +0.966]\) |
1.00 |
|
|
+0.1694 |
0.1044 |
+1.24 |
0.214 |
\([-0.070, +0.335]\) |
0.92 |
|
|
+0.5000 |
0.0953 |
+3.12 |
0.00179 |
\([+0.338, +0.712]\) |
1.00 |
|
|
+0.2222 |
0.1150 |
+1.77 |
0.0774 |
\([-0.021, +0.427]\) |
0.96 |
|
|
+0.2187 |
0.0712 |
+3.10 |
0.00195 |
\([+0.096, +0.375]\) |
1.00 |
|
|
-0.1446 |
0.0736 |
-1.92 |
0.0555 |
\([-0.285, +0.001]\) |
0.97 |
|
|
+0.1754 |
0.0820 |
+2.14 |
0.0327 |
\([+0.008, +0.340]\) |
0.98 |
|
|
-0.4525 |
0.1159 |
-3.74 |
1.9e-04 |
\([-0.706, -0.253]\) |
1.00 |
|
|
+0.2033 |
0.1082 |
+1.84 |
0.0651 |
\([+0.005, +0.433]\) |
0.98 |
screen |
rows |
immunogenic |
AUROC |
|---|---|---|---|
GBM |
150 |
33 |
0.6333 |
HiTIDE |
1,505 |
37 |
0.7364 |
IEDB_neoag |
420 |
231 |
0.6962 |
ITSNdb |
197 |
128 |
0.5718 |
NCI |
336,300 |
100 |
0.9703 |
TESLA |
930 |
39 |
0.8564 |
VACCIMEL |
93 |
26 |
0.5011 |
mean over the 7 |
0.7094 |
grouping |
folds |
groups |
mean over deciding screens |
median |
pooled |
|---|---|---|---|---|---|
|
5 |
337,692 |
0.7103 |
0.7148 |
0.9631 |
|
5 |
5 |
0.6960 |
0.6962 |
0.9817 |
What it delivers#
Each screen held out whole and scored by a model that never saw it: mean 0.7094, median 0.6962, over seven screens that each carry at least 20 held-out positives.
Two grouped cross-validations agree with that, and they are the ones that could have disagreed. Peptide-grouped 5-fold over 337,696 groups gives mean 0.7158 over the deciding screens; twin-grouped, which deletes the shared candidate-generation lineage (TESLA, NCI and HiTIDE draw on one) as a single group, gives 0.6957. A fold that shares a laboratory with its training set and one that does not read the same.
binderis the largest coefficient at +0.7596 (\(p\) = 1.3 × 10⁻¹¹, sign stable in 400 of 400 resamples), thenexpr_lvl+0.5000 andC_corpus_self−0.4525. Six of the nine terms are sign-stable in at least 97.5 % of resamples; the three that are not arelog10a(0.921),expr_norm(0.962) andC_phys_charge(0.973).
Caveats#
Two screens read near chance, and on both of them the design is the reason. A screen that pre-selected its candidates on one of EPIC’s blocks cannot test that block, and the composite carries the block anyway.
ITSNdb, 0.5714 on 197 rows. It admits a peptide, positive or negative, only on experimentally validated MHC-I binding, so binding and presentation are equalised by construction — presentation alone reads 0.5165 there. It applies no expression filter, and abundance being free is why the set has an answer at all: the shipped nine-term score reaches AUPRC 0.7256 against the set’s own prevalence of 0.6497, and puts 10 of 10 at the head (precision@10 = 1.00). AUROC is the wrong statistic on a set built to hold binding constant.
VACCIMEL, 0.5011 on 93 rows. An allogeneic whole-cell vaccine cohort: nothing about the construct was chosen by a presentation-and-abundance pipeline, and the 93 candidates were enumerated afterwards, so those two axes do not separate responders from non-responders here. Read on the recognition blocks alone — presentation and expression dropped, same leave-one-screen-out protocol, same model — it reaches 0.6447. That is the reading the manuscript prints for this cohort, marked as recognition-only.
Pooled AUROC over the whole corpus is not the number to quote. cv_peptide pools to 0.9636
and cv_twin to 0.9816; both are mostly NCI, which is 336,300 of the 339,599 rows at a
prevalence of 0.03 %. The per-screen column is what carries a claim.
The artifact’s own verdict block reads "ship": false. Against v10 it records four
improvements, one tie and two regressions (IEDB_neoag −0.0254 on 424 rows, VACCIMEL −0.0448 on 93);
it shipped over that bar on the author’s decision. Reading it beside the two caveats above is the
point of printing it here rather than leaving it to be found in the JSON.
BIC compares within a fit population; leave-one-screen-out compares across. The population
moved at both of the last two version bumps. v10 to v11: 342,432 rows / 741 positives / 8 screens
to 339,599 / 597 / 7, because parent genes were resolved for the 51.2 % of rows that deposited none
and Gfeller_GBM left the corpus as 96.5 % Gfeller. v11 to v12: 339,599 / 597 to 339,595 /
594, entirely in one screen — IEDB_neoag 424 to 420 rows and 234 to 231 positives, every other
screen identical. So BIC is not the comparison for either pair. The leave-one-screen-out mean is:
0.6998 → 0.7102 → 0.7094.
An ablation is the other case, and there BIC is the right instrument, because
bench/epic/fit.py --drop holds the population fixed and changes only the term list. Dropping
all three corpus channels gives BIC 3089.4 and held-out mean 0.7001 against 3107.6 and 0.7094 for
the nine-term fit — the corpus block costs 18.2 BIC and buys +0.0093 of held-out mean.
v12 is v11’s specification on a frame five scorer epochs newer, and it was shipped for
reproducibility rather than for accuracy. Held out it is a wash: 0 improvements, 7 ties, 0
regressions, each screen judged at its own resolution 1/n_pos, and no coefficient moving by
more than 0.018. What v11 cannot offer is a fit population that still exists — 424 IEDB_neoag rows
is not something the current chain produces at any scorer epoch, so v11 is citable but no longer
reproducible. v12 additionally carries the CR1-corrected cluster sandwich, the Laplace posterior
and a fit.gof block, which the other seven cells already had.
This fit is pinned and does not get regenerated. Its coefficients, bootstrap, loo,
cv_peptide and cv_twin blocks are what the manuscript cites; mhcmatch build --check
compares version stamps and cannot see a hand-copied replacement, so the artifact is guarded by a
test that digests (coef, mu, sigma) instead.
mhc1.mouse.neoantigen#
Nine terms on 921 rows from the IEDB mouse neoantigen deposit, over 61 publications and 6 H-2 allotypes. One screen, so the publication is where prevalence lives and the intercept goes there.
rows |
921 |
immunogenic |
379 (41.2 %) |
screens |
1 (IEDB_neoag) |
references |
61 |
allotypes |
6 |
bootstrap clusters |
61 |
ridge \(\tau\) |
0.25 |
bootstrap |
400 resamples of reference_id cluster |
deviance |
599.5 |
BIC |
1,077.3 |
in-sample within-reference AUROC |
0.6335 |
corpus block |
k = 3, kernel |
block |
term |
coef |
boot sd |
z |
p |
95 % CI |
sign stability |
|---|---|---|---|---|---|---|---|
|
|
+0.5347 |
0.1689 |
+4.02 |
5.9e-05 |
\([+0.222, +0.971]\) |
1.00 |
|
|
+0.0686 |
0.2467 |
+0.40 |
0.692 |
\([-0.241, +0.633]\) |
0.65 |
|
|
+0.2330 |
0.2331 |
+1.52 |
0.128 |
\([-0.320, +0.674]\) |
0.91 |
|
|
-0.3139 |
0.2012 |
-2.38 |
0.0173 |
\([-0.702, +0.079]\) |
0.96 |
|
|
+0.1481 |
0.1636 |
+1.17 |
0.241 |
\([-0.241, +0.469]\) |
0.84 |
|
|
+0.1188 |
0.1582 |
+1.11 |
0.267 |
\([-0.321, +0.411]\) |
0.81 |
|
|
+0.2919 |
0.2154 |
+1.95 |
0.0507 |
\([-0.121, +0.810]\) |
0.94 |
|
|
+0.0009 |
0.2284 |
+0.01 |
0.995 |
\([-0.474, +0.527]\) |
0.55 |
|
|
-0.3120 |
0.2816 |
-1.68 |
0.093 |
\([-0.976, +0.292]\) |
0.90 |
What it delivers#
In-sample within-reference AUROC 0.6335, AUPRC 0.5931, over the 8 of 61 references carrying at least three of each class (448 rows).
binder+0.5347 is the term whose interval excludes zero (\(p\) = 1.6 × 10⁻³, sign stable in 399 of 400 resamples).The best single term as a univariate ranker, on the same within-reference axis, is
log10aat 0.5865; the joint nine-term fit gains +0.047 AUROC over it.
Caveats#
All three corpus channels read the human tables.
mhcmatch.mimicry.reference_species() routes thymus, self and viral alike to
human, so a mouse query is matched against the identical mhc1|…|human|3 tables the human
artifact scores against. Nothing is trained on human data — a corpus channel is a k-mer
density lookup, and all nine coefficients are fitted on mouse neoantigens. The reason is deposit
composition, not sample size: every one of the 2,663 allele-annotated peptides in the mouse thymic
deposit is H-2Db or H-2Kb, so the channel built from it measured one groove rather than
thymic selection, and the mouse viral deposit samples 9 allotypes against human’s 129. self
agrees across species at r = 0.9990 regardless. Corpus complementarity: what the repertoire was shaped by carries the matched-mass control
that rules out thinness as the explanation.
Expression is not covered by that and must not be. Human and mouse organs and tumours are
different tissues, so mhcmatch.expression stays species-keyed at every rung. The mouse
floors come from FANTOM5 CAGE tag density over the tissue the syngeneic model arose in and run
0.60–2.00 against human’s 0.10–0.40, and they are normal-tissue floors — there is no mouse
TCGA, so the human convention of taking the floor from the tumour’s own transcriptome has no
counterpart. Compare floors within a species, never across one.
Only ``binder`` is resolved, so a sign disagreement on any other term is not a finding. Four of
the nine take the opposite sign from the human fit (expr_norm, C_phys_charge,
C_corpus_self, C_corpus_viral), and all four of those have intervals spanning zero on the
mouse side.
Abundance is deposited on 315 of 921 rows (34 %). Elsewhere expr_lvl falls back to the
gene’s tissue median, which is what expr_norm already is, so on those rows the two terms are
the same column. Four alternative expression arms were measured — an availability indicator, a
pan-tissue contrast, both together, and a genuine mouse tumour transcriptome (GEO GSE245293, 6
syngeneic models) — and the shipped artifact carries vanilla, the human block unchanged.
No held-out split is fitted and none is reported. The uncertainty on every slope is the
400-resample cluster bootstrap over the 61 reference_id clusters.
The deposit is exhausted at this size. It holds 968 class-I rows, 966 in the length range, 921 fitted; what limits the model is not rows but deciding references, and 8 of the 61 carry enough of both classes to score.
mhc2.human.neoantigen#
Six terms on 1,112 rows from CEDAR, over 157 publications and 72 allotypes.
rows |
1,112 |
immunogenic |
656 (59.0 %) |
screens |
1 (CEDAR) |
references |
157 |
allotypes |
72 |
bootstrap clusters |
157 |
ridge \(\tau\) |
0.25 |
bootstrap |
400 resamples of reference_id cluster |
deviance |
452.5 |
BIC |
1,595.8 |
in-sample within-reference AUROC |
0.6020 |
corpus block |
not in this design – the deposited references are class I |
block |
term |
coef |
boot sd |
z |
p |
95 % CI |
sign stability |
|---|---|---|---|---|---|---|---|
|
|
+0.3773 |
0.3416 |
+2.33 |
0.0197 |
\([+0.032, +1.334]\) |
0.98 |
|
|
-0.1193 |
0.4138 |
-0.36 |
0.719 |
\([-0.813, +0.787]\) |
0.63 |
|
|
-0.2240 |
0.7327 |
-0.41 |
0.68 |
\([-1.806, +1.154]\) |
0.65 |
|
|
-0.2120 |
0.6114 |
-0.48 |
0.634 |
\([-1.260, +1.256]\) |
0.61 |
|
|
+0.1710 |
0.1237 |
+2.07 |
0.0385 |
\([+0.000, +0.493]\) |
0.97 |
|
|
-0.1530 |
0.1588 |
-1.28 |
0.2 |
\([-0.424, +0.199]\) |
0.81 |
What it delivers#
In-sample within-reference AUROC 0.6020, AUPRC 0.6811.
binder+0.3773 (95 % CI [+0.032, +1.334], sign stable in 98 % of resamples) andC_phys_buried+0.1710 ([+0.000, +0.493], 97 %) are the two terms whose intervals exclude zero.The best univariate is
binderat 0.5645 AUROC / 0.6560 AUPRC; the joint six-term fit gains +0.038 AUROC over it.
Caveats#
This is a CD4 response model over human self proteins, not a tumour-neoantigen model. Every
antigen in the deposit is a human self protein and every row is a CD4 response to one. 143 of the
1,112 rows are a cancer and 260 are healthy donors; the single largest disease is type 1
diabetes at 364 rows and 91 positives. The whole composition ships inside the artifact as
fit.population, so a consumer can read it without this page. Ranking class-II tumour
neoantigens with it is an extrapolation from that population, and worth stating as one.
Both expression terms come out negative and neither is resolved (sign stability 0.65 and 0.61,
intervals spanning zero either side). The disease mix is the visible reason: expr_norm is the
gene’s median in the disease’s target tissue, and 426 of the 1,112 rows map to no tissue and take
the pooled floor.
The default carries no corpus block, and from 1.20.0 that is a measurement rather than a
specification. A C_corpus_* channel is a density over a reference set of peptides —
displayed self, encoded self, viral — and until 1.20.0 all three published sets were class I, so
the block left the design rather than being fitted against a 9-mer density the 15-mer register does
not match. The HLA Ligand Atlas now supplies a class-II reference: 132,818 peptides over 29
benign tissues, against the 27,987 of the class-II thymic fraction alone.
Fitting the block on the same 1,112 rows answers in the negative. Every one of the three
coefficients spans zero — C_corpus_thymus +0.0170 (95 % CI −0.4944 to +0.3995, sign held in
53 % of 400 cluster resamples), C_corpus_self +0.5227 (−0.3044 to +0.9815), C_corpus_viral
−0.3612 (−0.8419 to +0.2928) — and the two that carry any weight take the opposite sign to the
human class-I fit’s −0.4525 and +0.2033. At class II the corpus block adds nothing, so the
six-term fit stays the default: blocks lists three entries and the corpus-geometry keys are
absent rather than declared and unused.
The nine-term fit ships anyway, reachable by name so the measurement can be reproduced:
from mhcmatch import rank
rank.aggregate("mhc2", "human", variant="corpus") # nine terms, four blocks
It is registered in mhcmatch.rank.AGGREGATE_VARIANTS and never in
mhcmatch.rank.AGGREGATE_ARTIFACTS, so no default moves and mhcmatch rank is unaffected.
The same corpus rebuild is what made the mouse class-I cell fittable, where the block does carry.
No held-out split is fitted. This is a GLM whose deliverable is a coefficient and the interval around it, and the interval already resamples whole publications. Cutting the corpus into folds would answer a different question at the precision of the 11 references (of 157) that carry at least three of each class.
mhc2.mouse.neoantigen#
Six terms on 468 rows from the IEDB mouse neoantigen deposit, over 30 publications and 7 H-2 allotypes.
rows |
468 |
immunogenic |
177 (37.8 %) |
screens |
1 (IEDB_neoag) |
references |
30 |
allotypes |
7 |
bootstrap clusters |
30 |
ridge \(\tau\) |
0.25 |
bootstrap |
400 resamples of reference_id cluster |
deviance |
334.9 |
BIC |
556.2 |
in-sample within-reference AUROC |
0.5741 |
corpus block |
not in this design – the deposited references are class I |
block |
term |
coef |
boot sd |
z |
p |
95 % CI |
sign stability |
|---|---|---|---|---|---|---|---|
|
|
+0.2340 |
0.6946 |
+0.56 |
0.575 |
\([-1.720, +0.748]\) |
0.70 |
|
|
-0.0456 |
0.7233 |
-0.11 |
0.909 |
\([-0.682, +1.978]\) |
0.55 |
|
|
+0.2009 |
0.5572 |
+0.84 |
0.399 |
\([-0.495, +1.723]\) |
0.80 |
|
|
-0.3904 |
0.6063 |
-1.51 |
0.132 |
\([-2.178, +0.171]\) |
0.92 |
|
|
+0.0604 |
0.5040 |
+0.18 |
0.856 |
\([-1.221, +0.409]\) |
0.53 |
|
|
-0.0748 |
0.2762 |
-0.38 |
0.704 |
\([-0.743, +0.393]\) |
0.61 |
What it delivers#
In-sample within-reference AUROC 0.5741, AUPRC 0.5917, over the 7 of 30 references carrying at least three of each class.
It completes the lookup: all four
(cls, species)neoantigen cells are fitted, so a mouse class-II run scores against a mouse class-II fit instead of refusing.
Caveats#
This is the thinnest of the four, and its own intervals say so. All six 95 % CIs span zero and
every \(|z|\) is below 0.7. expr_norm (−0.3904, sign stable in 92 % of resamples) is the
only term whose sign holds in more than nine tenths of the bootstrap. Use it as a ranking prior
over mouse class-II candidates; it is not evidence about any individual term, and this page will
not present it as any.
Abundance is deposited on 59 of 468 rows (13 %) — the same fallback collinearity described for mouse class I, at a quarter of the coverage.
No corpus block. The channel is computable here — corpus_tables.npz ships
mhc2|thymus|mouse|3 and mhc2|self|mouse|3 — and it is left out because it was measured
and adds nothing, not because a reference is missing.
No held-out split is fitted, for the same reason as the other two single-deposit fits.
What “mouse” means, component by component#
A mouse fit is not a mouse model end to end, and the difference is worth stating once rather than inferring it from six places. Three questions, never conflated: is the coefficient fitted on mouse observations, is the model or table it indexes built from mouse data, and is the reference it reads mouse?
component |
coefficient |
model / table |
reference read |
|---|---|---|---|
|
mouse |
mouse anchor models (5 shipped |
mouse panel, H-2 pseudosequences |
|
mouse |
species-agnostic |
mouse anchor model as the class-II register oracle |
|
mouse |
mouse |
mouse — FANTOM5 mouse, GEO GSE245293 for tumour |
|
mouse |
species-free by construction — Rose scale, Atchley AF5 |
— |
|
mouse |
human |
human — all three, both classes |
recognition / complement heads |
mouse |
mouse ( |
— (not in the EPIC aggregate) |
So the short answer is: presentation and expression are mouse; physicochemistry is species-free; the corpus block is human. Every coefficient is fitted on mouse observations either way — what crosses the species line is the table being indexed, never the fit.
The routing has no class key. mhcmatch.mimicry.CORPUS_REFERENCE is keyed
(species, component), so mouse class II is routed to the human tables exactly as class I
is. It happens not to matter today only because both class-II artifacts carry no corpus term at
all — but a class-II fit that grew one would inherit the substitution silently, so the rule is
stated here rather than scoped to class I.
Why the substitution, per component — Pearson r between the same peptide’s density under the
two species’ tables, measured on the 921-row mouse class-I fit population:
component |
|
what the mouse deposit is |
verdict |
|---|---|---|---|
|
0.9990 |
112,565,681 mouse against 121,968,158 human proteome windows |
the same table twice over; substitution is free |
|
0.8382 |
9 mouse allotypes ( |
a 9-of-129 allotype sample; transfers, with a caveat |
|
0.3245 |
2 mouse allotypes ( |
the H-2b motif and nothing else; does not transfer |
It is not a sample-size effect, and that was measured rather than assumed. Thinning the human
deposit to the mouse table’s window count still reproduces the full human column at r = 0.8933
and still disagrees with the mouse table at 0.2903. A human table cut to mouse’s size does not
become the mouse table — what differs is which grooves each deposit sampled, so depositing more
mouse thymic peptides from the same two allotypes would not close it.
Expression is the one rung that must not transfer, and for the opposite reason: human and mouse
organs and tumours are different tissues, so a human expression level is not a stand-in for a mouse
one at any sample size. mhcmatch.expression stays species-keyed at every rung. A corpus
channel transfers because a k-mer table over a TCR face is shared geometry. A tissue is not.
The second mode: --epitope pathogen#
Every artifact above is a neoantigen fit. A pathogen epitope is answered by a different
mechanism — autoimmunity is not inflammation — so it is a second model rather than the same
model with an extra covariate, and mhcmatch rank --epitope pathogen selects it. All four
(cls, species) cells are fitted in this mode:
mhcmatch rank pairs viral.tsv --epitope pathogen --score features
Six or seven of the nine terms leave the design, depending on the class, and only one block leaves for a reason that belongs to the mode.
Expression is undefined, not missing. A peptide from an organism the host does not transcribe
has no source-gene abundance and no matched normal to compare it against. expr_lvl would be a
number for a quantity that does not exist, so mhcmatch.rank._expression_for() returns NaN
with imputed=False — imputed=True would claim a substitution rung was walked — and
--tissue / --tumor / --expr-floor are not read.
Which of the remaining channels a fit carries is the artifact’s answer, not the mode’s. Nothing
in the library selects channels by mode. rank.stand_in(mode, cls) supplies the column list for
--score features, where there is no artifact to read one from, and what it supplies is the
admissible set rather than the shipped one — four columns, C_phys_buried,
C_phys_charge, C_corpus_thymus and C_corpus_self, keyed by class in
mhcmatch.rank.FEATURES_ONLY_PATHOGEN, so that the next arm is fittable on any subset of
them. Everything that scores reads the artifact’s own fitted features list, and
mhcmatch.rank.TERMS_PATHOGEN_EXPECTED is the specification a shipped fit is checked
against: binder and C_corpus_self at class I, binder, C_phys_buried and
C_phys_charge at class II. Adding or removing a term stays additive.
C_corpus_viral is in neither list, and that is a statement about deposits rather than about the
mode. A viral reference is “foreign in origin, observed presented on MHC”, which is the
compartment a pathogen corpus is itself drawn from; where the two overlap, the channel reads
membership rather than similarity and its coefficient transfers to nothing, and a fit on a deposit
that does not overlap the reference would keep it. The host channel is the one that asks what the
mode is for: whether a foreign epitope resembles the proteome the repertoire was tolerised against.
The routing is a default, not a constant. mhcmatch rank --native-corpus scores the two
host components — self and thymus — against the query species’ own tables instead.
All twelve tables ({mhc1,mhc2}|{self,thymus,viral}|{human,mouse}|3) ship, so it is a routing
switch and nothing is fetched. It is off by default and warns on every run, for two reasons that
are both measurements rather than conventions: the mouse thymic table is the H-2b motif and
correlates with the human one at r = 0.3245 with thinness ruled out, and every shipped mouse
artifact was fitted against the human tables — so under the flag its coefficients meet a column
they never saw. viral is not a host compartment and stays human either way. Measured on
SIINFEKL / H-2Kb, switching moves C_corpus_thymus 0.000137 -> 0.001004 (7.3x, an H-2Kb
epitope meeting an H-2b table), moves C_corpus_self by 0.3 %, and leaves C_corpus_viral
bit-identical.
The four pathogen fits#
All four rest on one corpus construction: immunogenicity/pathogen_tcell_{cls}_{species}.tsv.gz
on isalgo/pmhc_data, one row per peptide, built from the full IEDB T-cell export, where the
epitope’s source organism is a pathogen by its NCBI lineage and the label is the assay itself —
Positive* against Negative. So a negative here is a peptide that was tested and did not
respond, and three of the four cells were unfittable for want of one. Each fit is a whole-corpus
GLM at ridge \(\tau\) = 0.25 with one global intercept and no grouping unit, 400 row resamples
for the interval and a 5-fold row cross-validation beside it.
Warning
Three of these cells looked unfittable, on the explanation that IEDB’s T-cell
export is positives-only. It is not. The full export (dump/tcell_full_v3_tsv.zip,
577,219 assay rows) carries 363,181 assays labelled Negative. What is positives-only is
a query-filtered download of it, and reading that download’s absence as the database’s own is
what left three of the four cells unfitted. Every fit below is built on those negatives.
The allotype composition is corrected for, not conditioned on. Positives and negatives do not
carry the same allotypes, and left alone that difference is read as chemistry. Each fit is a
weighted GLM whose prior weights take the corpus to its pooled marginal allotype frequency —
no peptide is subsampled away, and each class is normalised against the pooled target, so
prevalence is untouched. The balancing allotype is NetMHCpan-4.2 / NetMHCIIpan-4.3’s
%Rank_EL argmin over each record’s own recorded restrictions and never mhcmatch’s own, because
a weight that is a function of the column being fitted conditions the design on that column. The
whole block ships as fit.allele_balance: weights run 0.46 to 7.01 over the 107 balancing
allotypes of the human class-I corpus and 0.42 to 2.21 over the 7 of the mouse class-II one.
mhc1.human.pathogen
rows |
16,790 |
immunogenic |
7,002 (41.7 %) |
allotypes |
104 |
ridge \(\tau\) |
0.25 |
bootstrap |
400 resamples of row |
deviance |
– |
BIC |
– |
corpus block |
k = 3, kernel |
block |
term |
coef |
boot sd |
z |
p |
95 % CI |
sign stability |
|---|---|---|---|---|---|---|---|
|
|
+0.3418 |
0.0190 |
+17.93 |
7.3e-72 |
\([+0.311, +0.384]\) |
1.00 |
|
|
-0.0221 |
0.0173 |
-1.25 |
0.211 |
\([-0.056, +0.013]\) |
0.91 |
mhc1.mouse.pathogen
rows |
10,404 |
immunogenic |
2,196 (21.1 %) |
allotypes |
8 |
ridge \(\tau\) |
0.25 |
bootstrap |
400 resamples of row |
deviance |
– |
BIC |
– |
corpus block |
k = 3, kernel |
block |
term |
coef |
boot sd |
z |
p |
95 % CI |
sign stability |
|---|---|---|---|---|---|---|---|
|
|
+0.1819 |
0.0262 |
+6.88 |
6.2e-12 |
\([+0.135, +0.236]\) |
1.00 |
|
|
-0.1481 |
0.0267 |
-5.39 |
7.2e-08 |
\([-0.200, -0.101]\) |
1.00 |
mhc2.human.pathogen
rows |
7,946 |
immunogenic |
5,148 (64.8 %) |
allotypes |
80 |
ridge \(\tau\) |
0.25 |
bootstrap |
400 resamples of row |
deviance |
– |
BIC |
– |
corpus block |
not in this design – the deposited references are class I |
block |
term |
coef |
boot sd |
z |
p |
95 % CI |
sign stability |
|---|---|---|---|---|---|---|---|
|
|
+0.2947 |
0.0344 |
+8.52 |
1.6e-17 |
\([+0.223, +0.361]\) |
1.00 |
|
|
-0.1143 |
0.0298 |
-3.89 |
9.9e-05 |
\([-0.178, -0.055]\) |
1.00 |
|
|
-0.0094 |
0.0280 |
-0.33 |
0.742 |
\([-0.065, +0.043]\) |
0.64 |
mhc2.mouse.pathogen
rows |
11,725 |
immunogenic |
3,324 (28.3 %) |
allotypes |
7 |
ridge \(\tau\) |
0.25 |
bootstrap |
400 resamples of row |
deviance |
– |
BIC |
– |
corpus block |
not in this design – the deposited references are class I |
block |
term |
coef |
boot sd |
z |
p |
95 % CI |
sign stability |
|---|---|---|---|---|---|---|---|
|
|
+0.5240 |
0.0237 |
+21.09 |
1.0e-98 |
\([+0.480, +0.571]\) |
1.00 |
|
|
-0.2224 |
0.0214 |
-10.29 |
8.0e-25 |
\([-0.266, -0.184]\) |
1.00 |
|
|
+0.0125 |
0.0230 |
+0.57 |
0.566 |
\([-0.033, +0.056]\) |
0.71 |
What they deliver#
The lookup closes at eight of eight.
mhc2.human.pathogenandmhc2.mouse.pathogenare the first class-II pathogen fits andmhc1.mouse.pathogenthe first mouse one; earlier in development those three cells raised.binderis the largest coefficient in all four, from +0.1819 onmhc1.mouse.pathogen(10,404 peptides) to +0.5240 onmhc2.mouse.pathogen(11,725), and its sign holds in 400 of 400 resamples on every one of them.One non-presentation term resolves per class, and it is a different term in each. At class I it is
C_corpus_self, negative on both cells — resolved on the mouse one (−0.1481, \(p\) = 7.2 × 10⁻⁸ over 10,404 peptides) and not on the human one (−0.0221, \(p\) = 0.21 over 16,790). At class II it isC_phys_buried, negative and resolved on both (−0.1143, \(p\) = 9.9 × 10⁻⁵ over 7,946 peptides; −0.2224, \(p\) = 8.0 × 10⁻²⁵ over 11,725).Precision is the readout at these base rates, and AUROC is not. Pooled AUPRC against the cell’s own prevalence is a lift of 1.234 on
mhc1.human.pathogen(AUPRC 0.5147 against a prevalence of 0.4170, 7,002 immunogenic peptides of 16,790), 1.347 onmhc1.mouse.pathogen(2,196 of 10,404), 1.099 onmhc2.human.pathogen(5,148 of 7,946) and 1.571 onmhc2.mouse.pathogen(3,324 of 11,725).Each fit is stable to which rows it saw. Row-resampled 5-fold, mean and SD over the folds, against the in-sample AUROC it reproduces: 0.5986 ± 0.0108 against 0.5988, 0.5553 ± 0.0221 against 0.5561, 0.5818 ± 0.0100 against 0.5824, and 0.6445 ± 0.0059 against 0.6446, in the order the tables are printed above.
Caveats#
The two term sets are disjoint beyond presentation, so the four coefficient tables are not read
across. Class I drops the physicochemical pair because it did not earn a parameter: on the
balanced corpora, adding Rose burial and Atchley AF5 on top of binder and C_corpus_self
moved ROC-AUC by +0.0008 (human) and +0.0003 (mouse). Class II drops the corpus block, and drops it
completely — neither class-II artifact declares corpus_k, corpus_mask, corpus_kernel
or corpus_shapes, because mhcmatch.mimicry.corpus_geometry() reads a bare aggregate()
when it is passed no artifact, and a geometry no term of the fit uses is a claim it would act on.
Of the two host channels, one ships. C_corpus_thymus and C_corpus_self correlate at
+0.683 to +0.792 on these corpora and took near-equal-and-opposite coefficients wherever both were
fitted, so a design carrying the pair was fitting their difference rather than two mechanisms.
Fitted alone, self keeps the tolerance direction — a foreign epitope resembling the host
proteome is the one a tolerised repertoire is least likely to have responded to — and there is no
circularity in it: the rows are foreign peptides and the table is self.
Conditioning on presentation is what was given up for the sample size. The negatives are
assay-negatives of unknown presentation, so binder separates the two classes partly for a
reason that is not recognition: a peptide that was never presented could not have been responded to
either. Read the coefficient as a ranking term over candidates, not as a measurement of
recognition.
C_phys_charge resolves on neither class-II cell — \(p\) = 0.742 with sign stability
0.64 on the human one, \(p\) = 0.566 with 0.71 on the mouse one, and the two point opposite
ways. It is kept because dropping it would leave the class-II design a single non-presentation
term, and that is a refit’s call rather than this page’s.
log10a is absent because it duplicates binder, not because a pathogen has no wild
type. It is log10([P]/Kd) of the candidate itself (mhcmatch.rank._logit10()) and needs no
germline counterpart at all — the wild-type-dependent quantities are agretopicity,
d_occupancy and wt_absent, and those are degenerate in this mode rather than absent. The
duplication is not a property of the mode either: on the two class-II neoantigen populations the
same pair correlates at +0.7006 (mhc2.human, 1,081 of 1,112 rows carrying both columns
finite) and +0.7505 (mhc2.mouse, 468 of 468), and neither class-II log10a coefficient
resolves in its own fit — -0.1193 at \(p\) = 0.773 and -0.0456 at \(p\) = 0.950. Those
two artifacts still carry the term; whether they should is a refit decision, not a reading of these
numbers.
The groove-specific rows of mhc2.mouse.pathogen are 87.8 % H-2-IAb. The 7,784
haplotype-restricted peptides are what carry the d, k and s haplotypes. Check any pooled statistic
over this fit against that skew — it is the same shape as the class-II ligand pool that once
produced a below-chance AUROC.
These are in-sample, pooled-off-the-logit figures, unlike the four neoantigen fits above. Read each against its own prevalence: 0.4170, 0.2111, 0.6479 and 0.2835.
Note
These are the only non-neoantigen artifacts, and the only ones fitted with a global
intercept. The intercept is recorded and shipped as null like every other fit: what ships
is a ranking, and calibration to a population is mhcmatch.rank.probability(), which the
caller owns because only the caller knows their prevalence.
They require a candidate set at scoring time. Each was fitted on the best-presenting allotype
among those its record supports, so each declares
allele_policy = {"kind": "panel", "select": "presentation", "source": "record"} and
mhcmatch.rank._check_allele_policy() refuses a bare rank pairs run:
mhcmatch rank fasta prot.fa --alleles H-2Kb,H-2Db --epitope pathogen # works
mhcmatch rank pairs p.tsv --allele-panel H-2Kb,H-2Db --epitope pathogen # works
mhcmatch rank pairs p.tsv --epitope pathogen # refused, by name
That refusal is a measurement, not a convention: the best-of-record binder sits a mean
-0.2146 below the same peptide’s single-recorded-allele binder on human class I, and
-0.1403 / -0.1248 / -0.0533 on the other three. Between 52.7 % and 73.7 % of peptides carry exactly
one candidate and are unaffected; for the rest the coefficients would meet a shifted column with no
NaN, no error and an unchanged header. All four pathogen fits declare such a policy.
Reading any of this yourself#
The artifact is the record, and both interfaces print it:
mhcmatch rank --coefficients # every term, its block, its coefficient
mhcmatch rank --holdout # held-out AUROC, the grouped CVs
mhcmatch rank --coefficients --species mouse # the mouse class-I fit
mhcmatch rank --coefficients --cls mhc2 --species mouse
from mhcmatch import rank
rank.models() # every shipped fit: model_id, version, release, rows
a = rank.aggregate("mhc1", "human") # the artifact itself
a["coef"], a["ci95"], a["fit"], a["loo"]
Standardisation (mu, sigma) travels inside the artifact, so a caller reproduces the
score exactly, and a feature you cannot supply contributes its training mean — which is what “no
information” should do. A candidate with no expression value is scored on the terms it has rather
than dropped.
The blocks, the terms and what each one is computed from are Ranking neoantigens; the recognition axis is The recognition axis, reduced and Corpus complementarity: what the repertoire was shaped by.