Corpus complementarity: what the repertoire was shaped by#
What it is. Three of the nine terms fitted by mhc1.*.neoantigen — C_corpus_thymus,
C_corpus_self and C_corpus_viral — each measuring how densely a candidate sits among a
reference set a real repertoire was actually shaped by. Thymic immunopeptidome reads as danger, the host proteome as
tolerance, the foreign ligandome as a reference never seen during selection.
Why it is cheap. Each channel is a 64 KB k-mer table contraction, not a neighbour search, so all three together cost three table lookups and the ranking path builds no proteome index at all. The tables ship in the wheel.
Why it is label-free. No immunogenicity label is anywhere in the fit — the channels are densities against deposits, so nothing here can memorise a screen’s outcome.
Complementarity is two factors (Complementarity: the recognition axis); the chemistry one, C_phys, is a single
imported scale and is The recognition axis, reduced. This page owns the other one — and the reason the term that
used to carry it was fitted on the wrong question.
That term was C_aa, the residue-identity half of mhcmatch.complement: forty log-odds
cells estimated on the Chowell corpus. Chowell separates peptides that are foreign from peptides
that are self and presented — a statement about passing thymic selection. A neoantigen is a
self peptide carrying a somatic mutation, and whether a T cell responds to it is a different
question. So C_aa imports a selection discriminator into a neoantigen model.
mhcmatch.mimicry.corpus_R() is the label-free replacement, and its thymus channel is what
the shipped aggregate scores as C_corpus_thymus. It reads how close a candidate
sits to reference peptide sets a real repertoire was actually shaped by.
Three references, separated by when a T cell meets them#
The three channels are not three flavours of one measurement, and the difference predicts their fitted signs before anything is estimated.
channel |
what a T cell does with it |
reads as |
sign in the shipped fit |
|---|---|---|---|
|
The thymic immunopeptidome — self displayed on MHC in the thymus. The only one of the three that enters selection. |
danger |
|
|
The host proteome. Encoded, with no guarantee of presentation; the self a mature T cell meets in the periphery, where tolerance is maintained rather than established. |
the block’s background, not tolerance — see below |
|
|
A foreign presented ligandome. A thymocyte never sees this during selection. A hit is about peripheral priming — a different mechanism. |
reference only |
|
Why the thymic channel is positive#
Read as tolerance, a positive coefficient is backwards: clonal deletion should make thymic similarity reduce immunogenicity. The sign is right and the reading was wrong.
The thymus is not a random sample of self, because it cannot afford to be. A medullary epithelium a few million cells across cannot display the whole proteome to every passing thymocyte — there is not enough presentation capacity, and each cell shows only a small slice of what the tissue as a whole can show. Something has to choose what makes the cut. Dedicated machinery does: medullary thymic epithelial cells promiscuously express tissue-restricted antigens under the control of Aire and, independently, Fezf2, and losing either produces organ-specific autoimmune disease rather than a general failure of tolerance.
So the working hypothesis this term rests on is a selection argument. If display is scarce and its purpose is to prevent autoimmunity, the peptides that get displayed are the ones whose escape would be most damaging — the self antigens that would drive a destructive response if a reactive clone survived. The thymic immunopeptidome is then a curated list of what self looks like when it is dangerous, not a uniform sample of self.
That inverts what a thymic hit means for a neoantigen. A candidate resembling a thymic ligand is not being flagged as tolerated; it is being flagged as built like the self peptides the immune system was specifically defended against — which is exactly the shape a T cell responds strongly to. Escaping deletion is a property of the individual’s repertoire; looking like something worth deleting against is a property of the peptide, and it is the peptide the model scores.
The prediction that follows is the sign dissociation, and it is what is measured: thymus and
self are both similarity to self peptide sets, and they take opposite signs. A tolerance
account predicts both negative. A “typicality” account predicts both the same sign. Only a
curated-display account predicts one of each, and no single-mechanism account does.
Measured on the burial axis of The recognition axis, reduced (mean Rose propensity over the TCR face, human 9-mers):
set |
n |
mean face ρ |
|---|---|---|
thymus MHC-I ligands |
17,546 |
0.7303 |
presented self, non-thymic |
60,000 |
0.7222 |
human proteome, random windows |
60,000 |
0.7204 |
Thymus against random proteome, Cohen’s d = +0.1842; thymus against non-thymic presented self, d = +0.1650 (p = 1.0×10⁻⁸⁰). The second comparison holds the MHC-I presentation filter constant — both sides are eluted human self 9-mers — so the enrichment is thymus-specific.
The formula#
The term is the exact Łuksza sum, evaluated as a table lookup. The paragraphs below give it in four steps and then say why the lookup is exact rather than an approximation.
1. The face, and the window that slides along it. Mask the anchor positions
ANCHORS = (0, 1, 2, -2, -1) to leave the TCR face. For class I the anchor set is the first three
and last two positions, so the face is contiguous — literally p[3:L-2] — and W = L − 5
residues wide, at every length from 8 to 15. A width-k window slides along it, giving
\(m_k(p) = W - k + 1\) windows per peptide. Class II gathers its face around the floating core
and the window slides over that projection instead.
Sliding rather than taking the whole face is what lets a query of one length be compared against a reference of another: the table is keyed on the k-mer, not on the length.
2. The reference table. \(T_k[x]\) counts, over the whole reference corpus D and over
every length it contains, how many sliding windows equal the k-mer x — with multiplicity,
one increment per reference peptide per window, which is the published Łuksza form.
\(N_k = \sum_x T_k[x]\) is the total reference window mass. corpus_counts().
3. The sum, over every reference. For each query window u, weight every k-mer in the table by \(\beta^{d_H(u,x)}\) with \(\beta = e^{-\kappa}\), and add:
There is no radius and no k-nearest cutoff. \(\beta^d\) is the threshold, and it is applied to every reference in the corpus.
4. Normalisation, twice, and each divisor has a job.
\(N_k\) makes the value a density, so thymus (140,482 reference windows) and self
(121,968,158) land on one scale and “does thymus make the others redundant” is a comparison rather
than a deposit-size effect. \(m_k(q)\) makes it per query window, which is what removes the
length artefact: the fixed-face column this replaced varied 17× in mean across lengths 8–11 and
correlated with length at Spearman −0.502, against +0.036 here (3,600 real epitopes, 900 per
length, none of them in the reference).
So \(\rho\) is the expected mismatch weight between a uniformly chosen query window and a uniformly chosen reference window. The Łuksza \(Z/(1+Z)\) saturation is gone as redundant: it existed to bound an unbounded count, \(\rho\) is already bounded, and the old column never left its linear regime anyway. \(a_0\) went with it — it was a scale the standardizer absorbed, and the length compensation \(e^{\kappa(L-a_0)}\) it carried is now the explicit \(m_k\) divisor.
Why the sum over the whole corpus costs a table lookup#
Hamming distance is additive over positions, so the weight factorises:
and the inner sum becomes a multilinear contraction — one 20×20 matrix applied along every axis of
the table (contract()):
Contract once, then every query is a single array index. Measured against a literal all-vs-all
over every reference k-mer, the contraction agrees to 5.5×10⁻¹⁶ — floating-point round-off,
not an approximation. The radius-2 trie search it replaced captured a median 0.4999 of the same
sum (IQR 0.4115–0.5556, min 0.1539, n = 600 real 9-mers), and cost ~46 s against 2.3 ms for
340,876 queries. The ~7.5 GB proteome index the self channel needed became a 64 KB table.
Any position-additive, ungapped score factorises the same way: pass a BLOSUM62 kernel
\(K[a,b] = e^{\kappa\sigma(a,b)}\) and the graded Łuksza form is exact at identical cost
(verified to 4.4×10⁻¹⁶). Gapped alignment does not factorise, which is the one real limit and why
features() and safety() keep their index — they
also have to report which reference was hit, which a weighted sum cannot.
What \(\kappa\) actually controls#
\(K\) has exactly two eigenvalues: \(1 + 19\beta\) on the constant direction of each position’s 20-vector space, and \(1 - \beta\) on its 19-dimensional complement. On the tensor product, a mode that is informative on a set S of positions and flat on the rest is scaled by \((1-\beta)^{|S|}(1+19\beta)^{k-|S|}\). Divided by the total mass — which is the all-constant mode — an order-|S| interaction survives with weight
So \(\kappa\) is a single scalar bandwidth and \(\rho\) is a \(\gamma\)-weighted
ANOVA sum over interaction orders. \(\gamma \to 0\) as \(\kappa \to 0\) (every reference
weighted alike; the table says nothing) and \(\gamma \to 1\) as \(\kappa \to \infty\) (exact
k-mer matching; nothing smoothed away). The law is verified against the real tables to 10⁻¹⁵ in
bench/results/kmer_spectrum.md.
Why k = 3#
k is bounded below by coverage, not chosen by fit. The face is L − 5 wide and the shortest class-I ligand is an 8-mer, so an 8-mer supplies exactly three face residues: at k = 4 it has no window at all, and at k = 5 neither an 8-mer nor a 9-mer does. That reads as a low score and is really a structural zero, which is the failure mode the \(m_k\) divisor removes. Restricted to the rows every k can score, a wider window buys nothing (profile deviance 375.7 at k = 3, 4 and 5).
Fitted shapes#
\(\kappa\) is profiled per component by within-screen deviance over the whole fit corpus and
vendored as SHAPES — those are the Hamming-kernel values, and the table
below is read under that kernel. The shipped BLOSUM62 artifact carries its own
\(\kappa\) = 1.65 (thymus) / 0.65 (self) / 1.35 (viral) in its corpus_shapes
field, which corpus_shapes() reads at score time and which
SHAPES only falls back to. The profiles are shallow, so the rule is the
smallest \(\kappa\) within 0.05 deviance units of the minimum — an argmin on a flat
likelihood is noise, not a fit.
component |
\(\kappa\) |
\(\gamma\) |
profile range (deviance) |
what it says |
|---|---|---|---|---|
|
3.0 |
0.49 |
1.05 |
A genuine interior optimum. Tolerance is graded: a near-miss against the thymic immunopeptidome counts, and forcing exact matching makes the column worse. |
|
5.0 |
0.88 |
0.12 |
Not identified. The profile barely moves over a 24-fold range of \(\kappa\), because the
human proteome occupies essentially every cell of the 3-mer table — smoothing cannot
reorder anything. That is the same fact as |
|
8.0 |
0.99 |
4.94 |
Runs to the exact-match limit. What matters is sharing an actual 3-mer with a viral ligandome, not resembling one. |
Using it#
from mhcmatch import mimicry
spec = mimicry.corpus_spectrum(cls="mhc1") # all three components; memoised
rows = mimicry.corpus_R(["GILGFVFTL"], spec)
rows[0]["thymus"], rows[0]["self"], rows[0]["viral"]
There is no cache to manage and no index to build. The two halves are split on purpose:
corpus_counts() builds the count table — the expensive part, 0.5 s for an
immunopeptidome and ~50 s for the 122-million-window human proteome — and
contract() applies \(\kappa\) in about a millisecond. The counts are
memoised per (class, component, k, species) and not keyed on \(\kappa\), so profiling
the decay costs one build per component rather than one per grid point.
The memo needs no lock. A table is built into a local, frozen read-only, and published with a single dict assignment; two threads racing both build the same array, because the table is a pure function of the reference deposit and the key. Nothing partially built is ever visible and nothing shared is ever mutated. There is deliberately no disk cache: the artifact is 64 KB and the largest build is under a minute, so a cache directory would only add a staleness mode.
Species — and why a mouse run reads the human corpus tables#
self_species picks which species’ reference table a channel is contracted against, and all six
class-I tables ship — thymus, self and viral for each of human and mouse:
mimicry.corpus_spectrum(cls="mhc1", self_species="mouse") # explicit: mouse everywhere
That call does exactly what it says. What the scorer does is not the same thing, and it is
decided by mhcmatch.mimicry.reference_species():
from mhcmatch import mimicry
{c: mimicry.reference_species("mouse", c) for c in ("thymus", "self", "viral")}
# {'thymus': 'human', 'self': 'human', 'viral': 'human'}
A mouse class-I query is matched against the same three human tables a human query is —
mhc1|thymus|human|3, mhc1|self|human|3, mhc1|viral|human|3, the identical tables
behind aggregate_mhc1.json. A human query is unaffected, and the scoring path is bit-identical
to what it was.
Note
Nothing is trained on human data by this. A corpus channel is a k-mer density lookup: the
query peptide’s TCR-facing windows are matched against a counted reference table, and the table
is the only thing that is human. Every coefficient in aggregate_mhc1_mouse.json is a GLM
coefficient fitted on mouse neoantigens, and presentation, expression and physicochemistry all
read mouse sources.
Why: the reference deposit must not be one groove#
The corpus channel asks how much does this candidate look like peptides the immune system has
already seen in context X. It answers that by pooling a deposit into a k-mer table, and a pooled
table is only about context X if the pool is not dominated by one MHC allotype — otherwise it is
that allotype’s binding motif wearing the channel’s name, and it is collinear with binder,
which already scores the groove.
The mouse deposits are too small and too groove-skewed to be a reference, and the thymic one is the
extreme: of its 6,661 class-I peptides, every one of the 2,663 that carries an allele annotation
is H-2Db (1,574) or H-2Kb (1,089) — no other haplotype appears at all — and
that column was being applied to a fit spanning six H-2 allotypes. The human thymic deposit (HLA
Ligand Atlas thymus, 25,891 class-I peptides pooled over donors) has no single groove to encode.
The three components sit at three points on that axis. r is the Pearson correlation between the
same peptide’s density under the two species’ tables, over the 921-row mouse class-I fit
population:
component |
distinct class-I allotypes |
|
what the mouse table is |
|---|---|---|---|
|
none — an unpresented proteome window has no groove |
0.9990 |
the same table twice over. At r = 0.9990 across 113 M against 122 M proteome windows the substitution is free in either direction, and taking human keeps one reference source |
|
9 ( |
0.8382 |
a 9-allotype sample of a 129-allotype space |
|
2 — |
0.3245 |
the H-2b motif, and nothing else |
It is not a sample-size effect#
The mouse thymic table stands on 25,264 reference windows against human’s 140,482, so thinness is the obvious explanation — and it is the wrong one. Thinning the human deposit at the peptide level to the mouse table’s window count, 40 draws, still reproduces the full human column at r = 0.8933 (range 0.8728–0.9109) and still disagrees with the mouse table at 0.2903 (0.2467–0.3310). A human table cut to mouse’s size does not become the mouse table. What differs is which grooves each deposit sampled, and no number of extra mouse thymus peptides from the same H-2b studies would close it — which is what makes substitution the fix rather than a stopgap.
This is the same failure background="ligand" had, and the same one that took the corpus block
out of the mouse class-II model: a pooled statistic over a pool dominated by one allotype measures
that allotype. Mouse class II is the extreme case — its thymic deposit is 1,490 peptides, all
I-Ab — and there the block was dropped rather than substituted, because no class-II arm made
it pay. Both released class-II artifacts, human and mouse, are six-term models with no corpus
block. From 1.20.0 the human one has a nine-term sibling, variant="corpus", fitted against the
HLA Ligand Atlas’s 132,818 class-II peptides now that such a reference exists; every one of its
three coefficients spans zero, which is why it is a variant and not the default. See
The shipped models.
Where the corpus block is resolved#
The three channels are fitted with their own coefficients on both class-I populations, and the two
populations resolve them very differently. On the human artifact’s 339,595 rows / 594 positives /
523 (patient, screen) clusters all three separate from zero:
channel |
coefficient |
p |
sign stability |
|---|---|---|---|
|
−0.4578 |
7.9 × 10−5 |
1.00 |
|
+0.1733 |
0.026 |
0.98 |
|
+0.2091 |
0.049 |
0.98 |
The mouse population is 921 rows / 379 immunogenic over 61 references, of which 8 carry at least
three of each class; it resolves binder (+0.5347, p = 0.0016) and reports wide intervals on
everything else. What the human tables buy there is sign coherence with the fit above:
C_corpus_thymus is +0.2919 at sign stability 0.94, agreeing with the human artifact’s
+0.1733, where the mouse tables give −0.0056 at 0.53 on the same rows and the same terms.
Warning
This does not extend to expression, and must not. Human and mouse organs and tumours are
different tissues, so a human expression level is not a stand-in for a mouse one at any sample
size; mhcmatch.expression stays species-keyed throughout, reading FANTOM5 mouse and
GSE245293 rather than GTEx and the human tumour half. Mapping a gene to its mouse orthologue is
a different operation — it fixes gene identity and still reads a mouse transcriptome.
The corpus channels transfer because a k-mer table over a TCR face is a shared geometry —
self proves it at r = 0.9990 across 122 M against 113 M proteome windows. A tissue is not.
The measurements are bench/epic/corpus_transfer.py and bench/epic/corpus_paired.py in the
benchmark repo, recorded in bench/results/epic_mouse_corpus_transfer_mhc1.md and
bench/results/epic_mouse_corpus_human_vs_mouse_mhc1.md.
All three channels, and why all three are scored#
components= selects the channels, and since EPIC v3 the shipped aggregate has read all
three;
since artifact v4 it reads them under the graded BLOSUM62 kernel.
mimicry.corpus_spectrum(components=("thymus",)) # one channel
mimicry.corpus_spectrum() # all three -- what rank uses
They were not all scored before, and the reason was cost rather than worth: self needed a
~7.5 GB proteome trie, which is why the corpus term was once thymus alone. The contraction
removes that cost — three channels are three 64 KB tables — and the held-out numbers say to
carry them: adding the corpus block moves leave-one-screen-out mean AUROC from 0.6840 to
0.6927, the largest gain of any recognition block.
The three are not independent, and the next section is what that turns out to mean.
self is the block’s background term, not a third measurement#
C_corpus_self fits with a large, highly significant negative coefficient while its own marginal
AUROC is 0.4662, below chance. A large, highly significant coefficient on a column that predicts nothing
by itself has two readings, and they matter for how the score should be explained to a user:
tolerance — resembling the proteome genuinely lowers the odds of a response; or
background —
selfis the reference level the other two channels are read against, the term an R formula removes with~ 0 +.
The three channels correlate +0.70 to +0.79, so both readings fit the full model equally well.
What tells them apart is dropping partners: a tolerance measurement keeps its sign and its size
alone, a background term does not. Every non-empty subset, entered on the same base block, at the
shipped \(\kappa\) = (1.65, 0.65, 1.35), and read off the same 400 cluster resamples —
one bootstrap for all eight designs, so a coefficient that grows when a partner is added grew on the
same resampled patients (bench/results/epic_corpus_decor.md):
channels in the model |
BIC |
LOO |
|
|
|
|---|---|---|---|---|---|
|
4162.8 |
0.6475 |
−0.018 (z −0.40, p 0.69, 63 %) |
||
|
4158.9 |
0.6527 |
+0.085 (z +2.05, p 0.041) |
||
|
4160.3 |
0.6508 |
+0.065 (z +1.75, p 0.081) |
||
|
4171.6 |
0.6521 |
+0.080 (z +1.14, p 0.25) |
+0.006 (z +0.09, p 0.93) |
|
|
4163.4 |
0.6599 |
+0.216 (z +3.77, p 1.7×10⁻⁴) |
−0.188 (z −2.95, p 3.2×10⁻³) |
|
|
4164.6 |
0.6541 |
−0.215 (z −2.45, p 0.014) |
+0.220 (z +2.98, p 2.9×10⁻³) |
|
all three |
4172.4 |
0.6602 |
+0.155 (z +2.29, p 0.022) |
−0.270 (z −3.11, p 1.9×10⁻³) |
+0.146 (z +1.70, p 0.090) |
Read it in three lines.
Alone, ``self`` is nothing — −0.018, p = 0.69, and it holds its sign in only 63 % of resamples, which is a coin flip. There is no tolerance effect to measure on its own.
Give it any partner and it is significant and ten times larger, and so is the partner:
thymusgoes +0.085 → +0.216 beside it (2.5×),viral+0.065 → +0.220 (3.4×). Neither channel is readable until the background is in the model.Take it away and the block dies.
thymus+viralwithoutselfis the decisive cell: both fall to non-significant (p = 0.25 and 0.93) and the held-out mean drops to 0.6521, below either channel on its own. Two channels sharing an uncorrected composition background cancel each other.
Refitting \(\kappa\) per subset does not soften it — thymus + viral without self
then reads p = 0.55 and p = 0.58, and self alone is unchanged at p = 0.69.
So self is the intercept of the corpus block. The human proteome is the null distribution of
peptide-like sequence: a candidate scores high against thymus or viral partly because it looks like
a thymic or viral ligand and partly because it is made of common amino acids in common
arrangements, and self is the best available estimate of that second part. Its negative sign is
the subtraction, not a tolerance measurement — which is also why it is negative on a corpus where
similarity to self should, on a naive tolerance account, be protective.
Three consequences that a user should take away.
Never quote ``C_corpus_self`` on its own. It is not “how self-like this peptide is, and that is bad”. Out of the block it is meaningless, and its marginal direction is the opposite of the story its coefficient tells.
The block is one term with three columns. Drop any of the three and the remaining two are worth less than they look; this is why the shipped model carries all three even though
viral’s own p is 0.090.The sign dissociation still stands and still needs the mechanism above. A background term explains why
selfis negative and large; it does not explain whythymus— similarity to a self peptide set, measured against the same background — comes out positive. Nothing about a shared composition axis produces opposite signs. That is the curated-display argument at the top of this page, and the ladder is what isolates it: onceselfabsorbs what the two channels have in common, what is left inthymuspoints the other way.
Re-run on the shipped v11 base: the same finding, and a sharper reading#
The ladder above was fitted on the v4 base block. Re-entered on the base the shipped v11 model
carries, on 339,599 rows and 597 positives over seven datasets, every qualitative conclusion holds
and two get stronger. Both readings are kept, and both regenerate: they are the decor and
decor-v11 stages of bench/run_epic.sh in the benchmark repository, writing
epic_corpus_decor.md and epic_corpus_decor_v11base.md; the written analysis behind this
section is epic_corpus_thymic_rescue.md beside them. The v11 arm was outside that chain until
2026-09-20, which is how a page here came to cite a file only a branch carried. Running it
confirmed every figure below reproduces from the recorded frame.
Note
The shipped base is v12. This section is left as measured on the v11 base, and
the stage is still named decor-v11, because v12 is v11’s specification on a newer scorer
epoch and the three corpus coefficients moved by at most 0.0058 between them
(C_corpus_thymus +0.1733 → +0.1754, C_corpus_self −0.4578 → −0.4525, C_corpus_viral
+0.2091 → +0.2033). Every conclusion below is about the ordering of those channels, which
does not move. Re-running the arm on the v12 base is a stage rename, not a new finding.
The decisive cell reproduces.
thymus+viralwithoutselfis again the worst subset of the seven — both channels non-significant (+0.0311, p = 0.68 and +0.0366, p = 0.60) and the worst BIC of any subset, 3116.0 against 3098.4 for the best.Each partner still resolves only beside self, and by more than before:
thymusgoes +0.0596 (z +1.20, 85 %) alone to +0.2534 (z +3.91, p = 9.4×10⁻⁵, 100 %) beside it, andviral+0.0584 (z +1.26, 88 %) to +0.2923 (z +3.85, 100 %).In the full block all three are now individually significant with the expected signs —
thymus+0.1556 (z +2.14),self−0.4350 (z −4.41),viral+0.2191 (z +2.53). At v4 that held only under themaxreduction; at v11 it is the shippedmeanconfiguration.
The sharper reading comes from one further arm. Replacing self by a constant reproduces
dropping it exactly — thymus returns to +0.0596, z +1.20, 85 % in both cases. Since
optimize.standardise mean-centres, a constant column carries the level and none of the
variation, so what the block needs from self is its variation across candidates, not its
level. That is narrower than “reference level”: self is a correlated covariate doing
suppression — it absorbs the composition the three channels share, and the partner coefficients
are what is left once it has. The consequences above are unchanged; only the mechanism is stated
more precisely.
The correlation is a property of the density scale, not of \(\kappa\)#
The obvious repair is to sharpen the kernel until the channels separate. It does not work. Sweeping one \(\kappa\) across all three, the pairwise r on the raw \(\rho\) saturates — it stops falling past \(\kappa\) = 3 and sits at +0.760 / +0.699 / +0.696 forever, because a bounded density is dominated by the same handful of high-mass k-mers in every reference. On \(\log\rho\) the same sweep keeps falling, to +0.359 / +0.365 / +0.294 at \(\kappa\) = 8.
The less obvious repair is to change coordinates, and that is worth stating because it is a trap.
Four representations were fitted — raw, \(\log\rho\), enrichment over self
(\(\log(\rho_c/\rho_{\text{self}})\)), Gram–Schmidt, and principal components. The last four
are exact rotations of each other, and they return the identical BIC of 4177.7 and the identical
held-out mean of 0.6522, with Gram–Schmidt and PCA reporting max |r| = 0.000. A rotation
relabels a linear model’s coefficients; it does not change what the model predicts. Orthogonalising
the block makes self collapse to −0.019 (z −0.35, 71 %) and hands its weight to
thymus_perp — which is the same finding as the ladder above, arrived at by rotation, and buys
nothing in fit.
What does change the numbers is changing the measurement. Reducing the query’s face windows by
their maximum rather than their mean — the nearest-window reading, same references and same
\(\kappa\) — gives the best BIC of any arm, 4167.8, and is the only configuration in which
all three channels are individually significant with the expected signs: thymus +0.1501
(z +2.53, p = 0.011, 99 %), self −0.2610 (z −3.05, p = 2.3×10⁻³, 100 %), viral
+0.1918 (z +2.24, p = 0.025, 99 %). Its held-out mean, 0.6557, is below the mean-reduced 0.6602,
so what ships is not settled by this arm alone. Both are recorded.
The check that the channels are behaving: they do not solve Chowell#
Chowell separates foreign immunogenic peptides from self eluted ligands. Similarity to self is close to that label definition run backwards, and the thymic deposit is a presented subset of the same negative set — so a corpus channel that scored well there would be reading how the negative set was built rather than measuring immunogenicity. The expected result is nothing, and that is what is measured:
corpus |
host |
n |
positives |
|
|
|
|---|---|---|---|---|---|---|
|
human |
464,161 |
14,712 |
0.442 |
0.452 |
0.525 |
|
mouse |
47,140 |
5,154 |
0.448 |
0.433 |
0.507 |
|
pooled |
9,888 |
5,035 |
0.467 |
0.472 |
0.554 |
|
human |
58,789 |
17,346 |
0.533 |
0.506 |
0.480 |
|
mouse |
6,948 |
5,267 |
0.536 |
0.507 |
0.500 |
|
mouse |
1,393 |
1,053 |
0.486 |
0.423 |
0.487 |
Every Chowell cell is at or below chance, self most of all — the right direction when the
negatives are self ligands. On Kešmir, whose negatives are foreign ligands that failed to be
immunogenic, thymus moves the other way (0.533 / 0.536). Compare the chemistry term, whose
largest deviation on the same corpora is 0.160 against the corpus block’s 0.077: chemistry
transfers to a selection corpus and the corpus channels do not, which is the separation the two
factors are supposed to have.
On the neoantigen screens the block does carry signal on its own: fitted with screen intercepts and nothing else, leave-one-screen-out mean 0.5781 against 0.6602 for the full v4-base model.
The matching option on the chemistry side is mhcmatch.complement.burial()’s scale=, which
selects the residue basis; The recognition axis, reduced owns it, together with the 576-candidate selection that
settled on "Rose".
What the retired parameterisation got wrong, recorded#
Both of these were measured under the radius-2 search and are kept because the corrections they motivated are now structural rather than optional.
The published threshold \(a_0\) was not identified. Z stayed below 1.320×10⁻³ over
328,276 cached peptides, so R = Z/(1+Z) never left its linear regime and \(a_0\) only
multiplied Z by a constant that any standardizing fit absorbs — the correlation between the
feature at \(a_0\) = 14 and at \(a_0\) = 26 was 1.000000. What \(a_0\) did carry was a
per-row factor \(e^{\kappa(L-a_0)}\) spanning \(e^{2\kappa}\) = 90× between a 9-mer and an
11-mer, and dropping it saturated R. The per-window divisor replaces that compensation with an
explicit one, so the parameter is gone rather than absorbed.
Grading the substitutions “bought nothing”, and that verdict was wrong. The recorded arm reported the fitted weight running to zero and the score collapsing to the raw hit count (r = 0.998108). It scored the graded form with a radius-capped neighbour search against an exact contraction, so it compared a truncation to the real thing: 18.5 % / 16.0 % of queries returned no hit at all, “running to zero” was the first point of its own grid, and it carried two channels of three because a proteome index was unaffordable under a search.
Re-run as a contraction with an identity-normalised kernel,
\(K[u,x] = e^{\kappa(\sigma(u,x) - \sigma(u,u))}\) so \(K[u,u] = 1\) exactly, the graded
kernel wins on leave-one-screen-out mean and median alike, under an identical
\(\kappa\)-refit protocol on identical rows (bench/results/epic_corpus_kernel.md). blosum62_kernel()
builds it and contract() takes it. The corpus channels have been BLOSUM62 since
artifact v4; Hamming is kept so earlier results reproduce.
Two variants do not pay, and both are informative. Wildcarding the anchors in place instead of slicing them out costs at least 0.014 of leave-one-screen-out mean under both kernels, and is not recovered at k = 4 or k = 5 — the anchors carry nothing this term can use. And the unnormalised kernel, taking \(\sigma(X,a) = \sigma(a,a)\) literally, pins \(\kappa\) at the grid floor in all three channels, because BLOSUM62’s diagonal spans 4–11 half-bits and that is the only way to neutralise a wildcard row of \(e^{\kappa\sigma(a,a)}\).
Scope#
Length is no longer in it. The uncorrected column correlated −0.3493 with peptide length; the per-window density correlates +0.0399 on the same rows. The correction was the \(m_k\) divisor, and the same correction had to be made on the chemistry side — see The recognition axis, reduced.
The three channels share an axis. Their joint contribution is what the held-out number rewards; which individual one carries it depends on entry order (above).
The reference is the adult HLA Ligand Atlas immunopeptidome, whose TCR-face cysteine content is 0.024 % against 2.085 % in the proteome — an ~85× mass-spectrometry depletion. A danger term read off a mass-spectrometry ligandome inherits that depletion.
Promiscuous expression is not itself selective: the gene array has been reported as random rather than chosen. The enrichment above is measured at the level of presented peptides, where processing and MHC binding intervene between transcript and ligand.
All three channels route a mouse run to the human reference —
mhcmatch.mimicry.reference_species()maps every component that way, and the section above says so in full.