Complementarity — the I and C of EPIC#
The recognition axis: chemistry, identity and the corpus channels.
mhcmatch.complement module#
The recognition axis as one score: whole-peptide physicochemistry and length, the same components
split MHC-facing vs TCR-facing, MJ1996 / repertoire-marginalised TCRen contact potentials,
hydrophobic-run and dipeptide motifs, and per-role residue log-odds — pooled, per length bin
(8/9/10/11+ at class I, quartiles at 14/16/19 at class II) and per position zone (relative thirds
of the TCR face at class I, the nflank/core/cflank register zones at class II, selected
by cls=), whose pooled aa_anchor/aa_tcr pair
reproduces mhcmatch.posbayes.llr() exactly, so that model is a strict special case. Linear
head, because
a diagonal Gaussian cannot represent a summed log-odds; the EM Gaussian parameters ship alongside
for comparison. Vectorised — pass a list, not a loop. See Complementarity — the I and C of EPIC.
Complementarity: how well a presented peptide complements a T-cell repertoire.
This is the recognition axis. A whole-peptide physicochemical predictor – summed two principal components of the amino-acid property matrix over the whole peptide and added its length: 13 parameters, fitted by EM. That construction is kept and five blocks are added, each answering something the pooled version provably cannot:
block |
what it adds |
|---|---|
|
PC1/PC2 of the property matrix summed over the peptide, plus length. The whole-peptide feature set, kept as the floor everything else is measured against. |
|
The same components computed separately over MHC-facing and TCR-facing residues, plus Kidera KF4 (hydropathy) per role. The two channels carry opposite-sign contributions for several residues; a pooled sum reports their difference, weighted by the corpus’s composition. |
|
Contact potentials, one per side: MJ1996 on the anchors (burial in a groove – MJ is 96.4% one-body with hydropathy as its dominant mode) and TCRen marginalised over a real CDR3 repertoire on the TCR-facing residues. TCRen is only 3.29% one-body, so no per-residue scale can be extracted from it and the unknown receptor side is integrated out instead. |
|
Contiguity of the hydropathy stretch: longest run, number of runs and
above-threshold fraction of hydrophobic TCR-facing residues. A run of 3-4 is
a different object from the same residues scattered, and no sum can express
the difference. |
|
Residue identity as a log-odds per amino acid per role, and the block that
knows the peptide’s geometry: a pooled anchor/TCR pair whose sum is exactly
|
|
The same over adjacent TCR-facing residue pairs – a preference for a specific dipeptide that no marginal composition feature can express. |
The head is linear, and that is a measured choice. The shipped posbayes score is a sum of
the two role log-odds – weights fixed at 1 on two of these columns. A diagonal-covariance Gaussian
classifier cannot represent that: it maps each column through its own quadratic and re-weights by
inverse class variances, so the additive form is outside its hypothesis space and the extra blocks
are paid for out of a worse fit to the term carrying most of the signal. On the training corpus the
EM Gaussian reaches 0.657 grouped-CV AUROC on the aa block where the plain sum reaches 0.711.
A linear head contains the sum as a special case, so whatever it adds is genuinely an addition.
The EM Gaussian parameters are vendored alongside anyway, so the comparison stays re-checkable.
It emits a log-odds, not a probability. Like posbayes, the corpus’s own base
rate is divided out, so a caller supplies whatever prevalence their setting actually has:
logit P(immunogenic) = score(peptide) + log(prior / (1 - prior))
and posterior() does exactly that.
Vectorised. score() takes an iterable and returns an array; the whole feature set is two
(n, 20) count matrices times a handful of property vectors, so scoring a full deposit is seconds,
not minutes. Pass a list, not a loop.
>>> from mhcmatch import complement
>>> complement.score(["GILGFVFTL", "SIINFEKL"])
array([1.79, 0.42])
Parameters are vendored per species in mhcmatch/data/complement_mhc1_{human,mouse}.json and
never refitted at import. The two hosts are never pooled: different MHC, different thymic
repertoires, so one fit across them is fitting a mixture. score(peps, species="mouse") selects.
Provenance, the block-by-block cross-validation, the corpus-transfer matrix and the size-matched
cross-species transfer are in the benchmark repo (bench/results/complementarity.md).
Both classes, and they are two fits rather than one with a parameter. score(peps) is class
I: the role split is P1-P3, PΩ-1, PΩ at fixed peptide positions. score(peps, cls="mhc2") takes
the anchors from the P1/P4/P6/P9 core of the floating 9-mer register
(mhcmatch.store.anchor_indices()) and reads its own vendored tables, fitted on 603,781 human
and 50,258 mouse class-II peptides from the IEDB export. Passing a class-II ligand to the class-I
path labels the wrong residues as anchors and returns a confident, wrong number, so the class is an
argument and never inferred from the length.
Parameters per class and species in complement_{mhc1,mhc2}_{human,mouse}.json. Validation is in
bench/results/complementarity.md and complementarity_mhc2.md.
The shipped model’s recognition term lives here too, under a different name. score() is
the thirty-column fitted head above; burial() is C_phys, one of the two Complementarity
factors the fitted aggregate carries (mhcmatch.rank) – a single imported residue scale
(PHYS_SCALE) summed over the TCR face, with no fitted residue parameters, which is why
it cannot memorise the corpus. What it buys over the fitted head, and how the scale was chosen, is
on burial() itself and in docs/burial.rst. Its partner C_corpus is
mhcmatch.mimicry.corpus_R().
- mhcmatch.complement.ANCHORS = (0, 1, 2, -2, -1)#
MHC-facing positions, signed, matching
mhcmatch.immuno.ANCHOR_SCHEMES"pockets"andmhcmatch.posbayes.ANCHORS.
- mhcmatch.complement.PARATOPE = {'A': (0.0468, 0.3668), 'C': (0.1994, 0.5024), 'D': (0.1813, 0.5419), 'E': (0.1014, 0.313), 'F': (0.0754, 0.3616), 'G': (-0.0, 0.2369), 'H': (0.2189, 0.5621), 'I': (0.0486, 0.416), 'K': (0.1234, 0.4516), 'L': (-0.0251, 0.1864), 'M': (0.0345, 0.3921), 'N': (0.1193, 0.3291), 'P': (0.0469, 0.2501), 'Q': (0.0384, 0.3411), 'R': (0.0515, 0.421), 'S': (0.0124, 0.2607), 'T': (0.1307, 0.3778), 'V': (0.0143, 0.2991), 'W': (0.0847, 0.4427), 'Y': (0.0172, 0.2474)}#
TCRen marginalised over 28,250,990 TRB IMGT CDR3 loops –
paratope(a) = sum_b f(b) TCRen(b,a)– and its spread over the same distribution. A residue can have a mild mean energy and still be highly discriminating across receptors, so both are features. Measured in the benchmark repo (bench/results/paratope_basis.md); the 32M-clonotype repertoire is not needed at runtime.
- mhcmatch.complement.PARATOPE_CONTACT = {'A': (0.1288, 0.4503), 'C': (0.1266, 0.5443), 'D': (0.1674, 0.6565), 'E': (0.0623, 0.3612), 'F': (0.0375, 0.3135), 'G': (-0.0177, 0.2136), 'H': (0.3078, 0.5235), 'I': (0.0233, 0.4399), 'K': (0.1341, 0.4223), 'L': (0.0396, 0.1878), 'M': (0.0131, 0.375), 'N': (0.0678, 0.3208), 'P': (0.0365, 0.2677), 'Q': (0.005, 0.2787), 'R': (0.0805, 0.4688), 'S': (0.031, 0.2542), 'T': (0.0965, 0.3768), 'V': (0.0596, 0.2969), 'W': (0.0887, 0.4643), 'Y': (0.037, 0.2607)}#
The same potential marginalised over the residues that actually contact peptide, rather than over the whole loop –
f_contact(a) = sum_{clonotypes, i} P(contact | locus, i, L) * 1[cdr3_i = a], geometry from the 370 structures and identity from the same 28M clonotypes.TCRen is a directed contact potential, so the flat loop composition behind
PARATOPEis the wrong measure for it, and wrong in a biasing direction: only 35.4% of TRB loop residues ever contact a peptide, and the germline V-encoded head and J-encoded tail carrying most of the flat mass are exactly the positions the structures put atP(contact) = 0.00(TRB, length 12: positions 1-2 and 9-12). Conditioning on contact is the germline-flank trim; it needs no separate rule. Composition is never taken from the crystals – 370 complexes are a poor estimate of which residues sit mid-CDR3 – only the geometry is.It is a materially different vector, not a rescaling: Spearman against
PARATOPEis +0.7549, 19 of 20 residues change rank, and the spread widens from 0.2440 to 0.3255. A structure-free second route restricted to the N-D-N insert agrees at Spearman +0.7173.Generated by
bench/immuno/paratope_contact_basis.py;bench/results/paratope_contact.mdhas the composition andparatope_contact_basis.mdthis vector. Opt in withscore(..., paratope="contact"); the default stays"loop"so no recorded number moves.
- mhcmatch.complement.PARAMS: dict = {'alphabet': 'ACDEFGHIKLMNPQRSTVWY', 'anchors': [0, 1, 2, -2, -1], 'arm': 'chowell_rebuilt/human', 'features': ['pc1', 'pc2', 'length', 'pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr', 'mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr', 'kd_run_max', 'kd_run_n', 'kd_run_frac', 'aa_anchor', 'aa_tcr', 'aa_anchor8', 'aa_tcr8', 'aa_anchor9', 'aa_tcr9', 'aa_anchor10', 'aa_tcr10', 'aa_anchor11', 'aa_tcr11', 'aa_tcr_n', 'aa_tcr_m', 'aa_tcr_c', 'kmer_llr'], 'fits': {'em': {'immunogenic': {'mean': [-0.06616284744384984, 0.033257527759931516, 0.7425823626840402, 0.009944481875500932, -0.09744510738436167, -0.10270366444392376, 0.15899611033358801, -0.021620240473703266, 0.0453647892974406, 0.017723454003337618, 0.5022583696379227, 0.029067682576747078, -0.011969435994234565, 0.15024378096183383, 0.23894006374608343, -0.01841692959247472, 0.03140174804271794, 0.03149765470224604, 0.10795577805722431, 0.09403711588635565, 0.285996840399637, 0.25304989654414084, -0.4283587039003903, -0.5300476240404369, 0.08119225093371836, 0.1213505372650956, 0.044784546090369255, -0.0738356488326727, 0.03588980019953098, 0.026272257168881406], 'var': [1.1516692663151664, 1.1084522262605425, 0.19860447709908008, 1.0869453773015285, 1.0310605727504196, 1.1253315125231618, 1.1444420360343173, 1.0927980035842606, 1.1519899538076086, 1.1071525977980514, 0.6586179938029688, 0.8947099943206341, 0.8880246294471223, 1.1369114207966735, 1.0532738268222466, 0.8999572184564513, 1.1610496236862218, 1.3275002839607049, 0.09581536907339404, 0.10353666897503208, 0.2807555549028624, 0.34450507409073905, 4.564916481275358, 4.441787944835681, 0.10721622172558731, 0.1657772376791387, 1.0902456488673937, 1.559667039606054, 1.156856098160668, 1.4032918413480189]}, 'mixing': 0.20847508289499742, 'non_immunogenic': {'mean': [0.017426242443356672, -0.008759504226350764, -0.19558439193947458, -0.0026192184712624652, 0.025665492520388408, 0.027050512871895235, -0.04187704842323269, 0.005694427714860107, -0.011948364485953862, -0.004668076093047893, -0.13228687181118653, -0.007655965597128482, 0.003152559202088637, -0.03957182397370027, -0.06293301514573582, 0.004850726541647888, -0.00827072134405988, -0.008295981633823916, -0.028433836122047553, -0.0247678817254575, -0.07532702220934928, -0.06664932084712008, 0.11282287439673347, 0.1396061197470374, -0.021384748449546936, -0.03196180280491581, -0.011795537647336141, 0.01944713637972825, -0.009452802950278383, -0.006919694974540168], 'var': [0.9585960947283462, 0.9710673715761963, 1.0275840819740554, 0.9770670750613483, 0.9886594572583193, 0.9634797671746005, 0.9535443010666477, 0.9754029519903206, 0.9592834760347848, 0.9716731960596869, 1.0059726022242426, 1.0274505716934175, 1.029444872547767, 0.9564283707210262, 0.9669707503174307, 1.0262368262064476, 0.9572539814923111, 0.9134115157580363, 1.234269793626622, 1.2331716577819047, 1.1622201173805702, 1.1513392221513465, 1.000785366389977e-08, 1e-08, 1.232951486845345, 1.214820890534065, 0.9755633505553665, 0.850778526544332, 0.9582579947100118, 0.8935496732434686]}}, 'supervised': {'immunogenic': {'mean': [0.48458438715948515, 0.06298948453314321, -0.029654583164330322, 0.39128355525065656, 0.1014682115471637, 0.24996067426166946, -0.016188312542218187, -0.3491597579256158, -0.2717712962243248, 0.3675560892923596, 0.18135148896737566, -0.042275395210148434, -0.006615296723916411, 0.20049069236962921, -0.019114653444797746, 0.19690361209248805, 0.6315466359237181, 0.6564536683522727, 0.14271720635612214, 0.12582837414290413, 0.5589047848821895, 0.5188691769230301, 0.26888146312369027, 0.3491639056385894, 0.11709064950644237, 0.18159236648548993, 0.354008713848338, 0.3918615009934163, 0.4387206547144678, 0.7714373541380991], 'var': [1.3930230496478904, 1.1602006424759057, 0.6030177894148014, 1.2263787688544587, 0.9778518154042607, 1.0985013400029648, 1.1708403457939764, 1.1354769271622067, 1.095528758069639, 1.3172639009536289, 0.9862524597940981, 1.054503535191982, 1.0096275333834448, 1.1582998494436165, 0.9218909598402559, 1.0523029114404265, 1.5847444990351371, 2.0242241415857265, 0.6287863446403338, 0.6798051670099481, 1.7587954787066233, 2.182601004693286, 1.5692679535176246, 2.000733627789036, 0.7036790214229164, 1.0860954313051818, 1.986962383550572, 1.9224728597493752, 1.669564234269702, 2.1575359637691927]}, 'mixing': 0.03169589862138353, 'non_immunogenic': {'mean': [-0.01586210115917164, -0.00206186084840703, 0.0009706957352459654, -0.012808046440973456, -0.003321400933777193, -0.008182066129269654, 0.0005298987295771788, 0.011429190761580657, 0.008896002238425973, -0.01203136548461819, -0.005936253291110097, 0.001383817995791796, 0.00021654124361396656, -0.006562744752223986, 0.0006256878566463744, -0.006445327369643187, -0.020672677228585953, -0.021487969422078916, -0.004671621340607238, -0.004118792210880399, -0.018294861475243836, -0.016984359362011706, -0.008801408136358138, -0.011429326530386547, -0.0038327766566152344, -0.005944138035086989, -0.011587913641138271, -0.012826964578101358, -0.014360824636687601, -0.02525177796385514], 'var': [0.9791968914743827, 0.994621969619256, 1.0129648653826153, 0.9874142364597183, 1.0003769476971038, 0.9946635842540551, 0.9943989636734838, 0.9914441481583227, 0.9943762110655474, 0.9850479244412838, 0.9993382267703844, 0.998155507163318, 0.9996833889205654, 0.9934594770280398, 1.0025444345477688, 0.9969773070126338, 0.9673762074590005, 0.9519060687198755, 1.011462555639168, 1.0099458548589053, 0.9646022565242619, 0.9521883135121277, 0.9789219266960645, 0.9631212447947094, 1.0092361349801686, 0.9960670700841824, 0.9634568602941078, 0.9646134000808131, 0.9715762636565185, 0.9419920898218387]}}}, 'kd_threshold': -0.8500000000000001, 'log_odds': {'aa_anchor': [-0.17195912826031323, 1.8451940731045058, -0.3383326314979689, -0.5309006197745685, -0.10548669598002158, 0.14203178413119355, -0.4438454689097804, 0.08188954910316015, -0.12307214034493219, 0.2531804504988795, 0.5553004601198759, 0.1305669975254693, -0.035223985277909264, -0.2042807334750738, -0.3067910455939242, -0.08833926936175907, 0.060135842421909835, 0.15719944510128148, 0.1361199476438575, 0.06362187397922714], 'aa_anchor10': [-0.2022520811905535, 1.6582365394881133, -0.21191815390528346, -0.4273384044275228, -0.19388742686109328, 0.15314775491206678, -0.5778472050268788, 0.0421627706836043, 0.012321574897647736, 0.19777070487777615, 0.5231004866997857, 0.15504047994184544, -0.03311038630277352, -0.07875792042935048, -0.3743890548054498, 0.10555445424948529, 0.1198990535270954, 0.2992389719106292, -0.15750866458101953, -0.04342257958925888], 'aa_anchor11': [-0.16065601260790485, 1.6174140397801606, -0.27465433156603636, -0.14863832184342884, 0.06167147632298109, -0.07367551716753917, -0.1010235917043052, -0.028861400128323833, -0.230528091540513, 0.03736191683968837, 0.36971437324133927, -0.07392299797558932, 0.21726648668495496, -0.4227304106649723, -0.24945682828088422, 0.09779000750694689, 0.15277128278115182, 0.06643388308785925, -0.058436472170595, 0.2639780408845138], 'aa_anchor8': [-0.13629873025390715, 2.6257263950162546, -0.7974534382634779, -0.4950730959299041, -0.08115130790852332, 0.20165648504369704, -0.3868574556412123, -0.028843247661880955, -0.09865958906304018, 0.04648481655892489, 0.4944428841186088, -0.03053823305143455, 0.22466831742077442, -0.16959297760528402, 0.40538983996537015, 0.039470325723848454, 0.11370982206898317, -0.021221858621291112, 0.4424026125901088, 0.16730566685241888], 'aa_anchor9': [-0.17862695719588118, 1.9054169060709576, -0.27658399902294883, -0.5987303966114133, -0.09066545816289695, 0.17642552217525376, -0.43654794822798815, 0.10186616378673508, -0.1596994472827098, 0.2892885908046945, 0.5735795678820441, 0.12129354135612713, -0.06062207561426547, -0.23424807832612737, -0.31999365078425024, -0.18613321931811377, 0.021565292739611497, 0.12556288364885804, 0.3091340996907128, 0.06728014303186924], 'aa_tcr': [0.1269650827576445, 2.0508651169873318, -0.1278672175910449, -0.3965923638827822, 0.312048362043968, 0.05456623245180037, -0.28960607601331123, -0.08651718328776914, -0.3822461012648932, 0.07643556676200847, 0.5557280465342154, 0.13994712385397712, -0.16182712771524965, -0.39178104663270386, 0.13865507523536413, -0.09380087552214711, 0.1356944452714206, -0.06006938642577664, 0.8066684257701158, 0.3563580715841401], 'aa_tcr10': [0.11041695229878457, 1.9695372867631002, -0.27353380791584314, -0.5717224209998402, 0.47369453970111497, 0.004229081709119953, -0.31082693732559763, 0.13940711867842825, -0.33756958013949534, 0.13552090933857208, 0.5602298437823778, -0.003674095757914664, -0.2440546440131821, -0.467221239478429, 0.12067295968179748, 0.030798063215855098, 0.08170025394600078, -0.006905284398639466, 0.9126222570779312, 0.31409292051353965], 'aa_tcr11': [0.034307092582310794, 2.1276411057690137, -0.32117676334623724, -0.22987957712244578, 0.2692742510626802, -0.09214137696067315, -0.1046682419251912, 0.038192708748248094, -0.4901140111953084, -0.00674018569624657, 0.4210891741674745, -0.17233903117264138, 0.08571050978170502, -0.3319401982231489, 0.3854749883302704, -0.06271559344998678, 0.13828374712029357, 0.002301475026118993, 0.6788054088668662, 0.5037665937803593], 'aa_tcr8': [-0.29697431523976325, 2.442364901686199, 0.21768694002602063, -0.06801892407850607, 0.6132789072842479, 0.2080022184566439, -0.6029684606121299, -0.18840420706895777, 0.17071023514109074, -0.18143216168115472, 0.6777731326144099, 0.040877527639617384, -0.004906500768449895, -0.31094769397583377, 0.07652626579107435, -0.25894400383096405, -0.009474325818346951, -0.22749187598759413, 1.0455387632473006, 0.6328266631195407], 'aa_tcr9': [0.17945770071163913, 2.0619637052103297, -0.029981044514179267, -0.3654405343298275, 0.20675501336828495, 0.12656639409957382, -0.2782317593896546, -0.1917114470680965, -0.4077037365276799, 0.07032137075823908, 0.5592548998295923, 0.23170647822326806, -0.17681564687095896, -0.3842853433911029, 0.12292084043406026, -0.12364457611197066, 0.15333852825101246, -0.07875146626659957, 0.7355366135993782, 0.31904688155671534], 'aa_tcr_c': [0.16333550161931853, 1.8707098333760452, -0.1148505690154531, -0.4670323791682849, 0.3314358608017822, 0.0754790845130513, -0.3641240402505592, -0.12193367623157014, -0.43474646569214226, 0.08468515767806473, 0.5125890659921324, -0.016602970922443117, -0.0327462463587751, -0.47481451107669725, -0.20126018609099328, -0.058127301973135204, 0.1656247749868056, -0.05152218711154033, 0.797024826894055, 0.2857599761897811], 'aa_tcr_m': [0.008107412078389054, 2.177840329799901, -0.04579142184387397, -0.3610209167881262, 0.28258237212748005, -0.017691991710642174, -0.2506611603667275, -0.06943184967596983, -0.37064317866220176, 0.14042794973508954, 0.4553026314937103, 0.07890375752309797, -0.4131976989638213, -0.29321347265154074, 0.42957503508856787, -0.156967566477761, 0.1484842222490097, -0.10365523055072456, 0.797835467602722, 0.34723132268913437], 'aa_tcr_n': [0.1954528096969046, 2.2986892206804814, -0.1450903656778295, -0.3338366929044301, 0.26025426122320683, 0.14542760252377196, -0.17975415714020437, -0.0486790082582389, -0.3081890476685425, -0.10720632699825039, 0.7943039763213995, 0.42803131629758884, -0.10253553972469698, -0.3429217934402682, 0.27063702790624644, -0.07733311107678142, 0.0014274289169473597, -0.06758518949137082, 0.8294865346693872, 0.5232115791339824], 'kmer_llr': [-0.017874478710091957, 2.4432388959174407, 0.06144394763864014, -0.6114595687515676, 0.35940076483606376, 0.15473638788561495, -0.24590417844759394, 0.0811799423628008, -0.6759640849584203, 0.07704089966355099, 0.6876447031574555, 0.3596030991356125, 0.2937716689416785, -0.4560810125888013, -0.05588165663171196, 0.18203482103036883, -0.03050596545520179, -0.00279635952314905, 0.9071851400401734, 0.8170942241911412, 2.039704290339335, 3.9193938320740926, 2.1948431125394867, 1.4795343601766895, 2.914371891493655, 2.0803334022280726, 1.4968059466853498, 1.8137364482633638, 1.5286670487543343, 2.1504166725145826, 2.725471363601658, 2.170793857306818, 2.0533775922395456, 1.4943698918874695, 1.7214750513083974, 2.0980755606044923, 2.3600115901071916, 2.1948431125394867, 2.725471363601658, 2.4070176324831225, -0.06015697214592741, 1.5358872967278216, -0.020396323000508865, -0.3722189044643205, 0.0833748258927951, 0.16644368827812972, -0.6500520212923888, -0.1094628095243877, -0.42434483187710015, -0.09896229280251223, -0.02967895563648515, -0.301205713801755, -0.1342283435993199, -0.4390103892007451, 0.12164621760607641, -0.4682697866954566, -0.2576821277454737, -0.23685898356808455, 0.4260161636165396, 0.025040878014332968, -0.20391580887762562, 1.155254164320839, -0.16269973320831443, -0.5903989533073615, -0.01164859093226589, -0.3190510741217656, -0.6669814276717467, -0.5312854263356224, -0.8575558118742936, -0.29560495805036346, -0.21291865568535684, -0.4031722756891618, -0.5110056567435519, -0.5330257352102787, -0.28319065796075105, -0.4342059873718789, -0.2791864755578377, -0.5135232940989543, 0.4797944028794916, -0.11533832734649341, 0.18963524221534822, 2.2791842609732385, 0.12024810700218147, -0.4669025205210362, 0.8707370952122142, 0.27657714973996583, -0.010262266248969532, -0.1378363243030014, 0.06099241079308726, 0.35862932063901276, 0.8689172584952276, 0.19842490667132573, 0.12886468626942893, -0.6399418175702003, 0.09661196042388909, 0.13491852467867016, 0.46800670588808124, 0.26182880456709867, 0.49187914209456274, 0.2553447641253772, 0.2732980410574122, 1.691397596071118, -0.47554477282783836, -0.5523655306917297, 0.3057323636287901, -0.07565643380480847, -0.33350915729699615, 0.41660467136448087, -0.5703655024026721, 0.0015026910617672584, 0.9938760213270443, -0.09863384240278705, -0.2547437887300674, -0.5284531210850592, -0.15679708037862028, -0.12993656045995206, 0.0695155347436236, 0.2597475332304313, 0.9809829427137178, 0.08744478480777396, 0.1801509512494528, 1.9349500190165356, 0.04020119083126694, -0.4600896779808812, -0.1250544857473077, 0.026029537036047934, -0.004052055207287353, -0.25704261886648005, -0.19720133908282822, -0.406360903549948, 0.05132271417512868, -0.4163963601995846, -0.8088785510479832, -0.8793703443232737, -0.060327590374930296, -0.5826355949944864, 0.010560691468139538, -0.5770769369906583, 0.5156995702133491, 0.5719218497681, 0.13442325317276005, 1.4795343601766895, -0.09718222023955647, -0.29812980896112595, 0.03677944408736078, 0.4200697279698362, -0.6745348326903446, 0.01311497127323591, -0.15232161208764694, 0.17450543403149865, 0.0032365275074912603, -0.026570406782087552, 0.020831990687240065, -0.6000141963248016, -0.2664839449216254, -0.29103717070534874, -0.03128501393214478, -0.25070275826406174, 0.6184530541515869, 0.04645821535399275, -0.29901032323831966, 1.9775992833524647, -0.4109574318347695, -0.9157379884941479, -0.14277183178377317, -0.26833598154274885, 0.18466705688795493, -0.9052957635354284, -0.46320955932057384, -0.44721746329507095, 0.20049772121677112, -0.012773128945724466, -0.7254719311318905, -0.7016580441538798, 0.16560676285823117, -0.5086074567861996, -0.5634689233436116, -0.7159228171868346, 0.7474728999948432, -0.16579605042245227, 0.2088597839639048, 2.2608585829907986, -0.08595264814181114, -0.4283145777283499, 0.3384104329689972, 0.23319170397478928, -0.27702095501005264, -0.19597809734008464, -0.4126038907299048, 0.28434529037365763, 0.5224664233499388, 0.19306555387259206, -0.22142578074321229, -0.2131158045146444, -0.1893219674469817, 0.13683051990041495, 0.16909147726837315, 0.14037577558778302, 0.8627115828785152, 0.3932640145762685, 0.5117884541166067, 2.812482740591287, 0.16984918901525958, 0.9683320950202061, 0.3275760908032872, 0.5782072255830606, -0.018758494078054966, 0.4157688028387474, -0.21633256792677802, 0.9722708879772854, 0.9950808407498943, 0.2135519145675886, 0.2977231276536063, -0.2733810372885772, 0.21881228201375258, 0.6338097105870455, 0.4633491342207927, 0.4673915105441626, 2.7705917988821263, 0.9515467620096363, 0.25985995105695014, 2.2926072813053793, 0.21057035794546586, -0.6234391816140699, 0.015704423880392504, 0.2957573725329903, -0.08443321382909108, 0.02530556602359635, -0.2662674394063833, 0.29230182296704044, 0.698698767504009, 0.4002472811006461, -0.20913837658789536, -0.011987105344703153, 0.14851919478157338, -0.042834696846123066, 0.9828618354852141, 0.004139788340288497, 0.06787938376957658, 0.6563926089573675, -0.39889212823543296, 1.689772900344117, -0.2837869060121019, -0.4532860480037648, 0.02121212290199459, -0.6092240560466005, -0.44553778359564067, -0.1557006973960835, -0.8005120159977759, -0.37802716736524555, 0.016846614939833415, -0.63370302828565, -0.58476172709838, -0.23234167509003978, 0.7829774258778581, -0.44081255909242856, -0.06577411175868697, -0.2518356231607708, 0.4401994618092884, 0.10019908895714114, -0.3626984770200412, 1.9231248910767196, -0.42532026771580966, -0.8856544276422138, -0.1523018024581857, -0.2512644092616787, -0.41089203969207944, -0.20638883335792269, -0.9792699948127455, -0.08423249122705201, 0.9711088747135914, -0.17869371642684317, -0.6315087157361994, -0.5934173705977743, -0.8132119992825038, -0.2105109055465597, -0.19073890002552663, -0.4876487987391185, 0.2831243282324536, 0.43398661312012443, 0.3882139879839652, 1.9015539518585172, -0.13935763973408122, -0.2543328423178828, 0.32200297587685434, 0.42514445749120355, 0.05296609991885948, -0.46637578887862396, 0.4897516558162911, 0.2261407326217224, 0.36637406034426956, 0.3375066590058875, -0.4893184549980303, -0.25521777163119364, 1.3221079424272144, 0.22204076525077543, 1.2314716770646736, -0.17816670719198413, 1.2579515625591808, 0.2230423631214551, 0.1315835285550282, 2.4293788126412803, -0.21679488249928092, -0.7731173640686473, 0.41838995642723287, -0.08789326914721851, -0.30670971966036564, -0.29439267003221037, -0.43731499461549994, 0.025641995644305027, 0.49776376954036206, 0.010010671608962518, -0.06112280519738267, -0.6232079837770517, -0.060507698697708, -0.23712292526221024, -0.1335498644104316, -0.27048212442066166, 0.7524592847685527, 0.30650561923395525, 0.09688533385369436, 1.4676193255988172, 0.033254665706303754, -0.1744037055573351, 1.0827025320703392, -0.22264505601066986, -0.31770879204103686, 0.06972534701581701, -0.21007189036285556, 0.4462987112323553, 0.9075421766402174, -0.02262633862982799, -0.02919820239785853, -0.26608553854185324, 0.2453040591775535, 0.07922493041392897, 0.0502389381938988, 0.3183243931040858, 0.9321978895423877, 0.4479342094767844, 0.4037122973838114, 1.835613888795657, -0.11777550944413662, -0.46008228388285133, 0.31984551668843153, -0.26721573757716666, -0.5830871642051401, -0.2016738002887628, -0.5587239273871853, 0.1299446165366298, 0.3357036753999454, -0.04894051319963921, -0.1246659262448313, -0.6427509929034283, -0.09130946412794838, -0.1943142795008006, -0.07826616232160521, -0.009635638888066289, 0.8418035234235655, 0.2337769057593837, 0.6298569626762198, 2.84807368569399, 0.8536691867000652, 0.4567878222832933, 0.9236005847692823, 0.5481569462247675, 0.45273923375729197, 0.9325221262380134, -0.04853680529235049, 1.1551208398756039, 1.1464926586522655, 0.9857682563296555, 1.1278679088145598, 0.5417474734843406, 0.9153627557054067, 0.7614537239189536, 0.5693581745078689, 0.19064140504705662, 0.9337118943736034, 1.802644660747717, 0.2647519394276605, 2.5658412180097745, 0.24630008389840174, -0.0675636803470443, 0.6704483492699671, 0.47248321174698393, -0.07961912154266049, -0.007189324685382914, 0.22771657133580803, 0.3860722974848949, 0.5730046869069314, 0.6245516126134572, 0.15336257357971128, 0.09052081593905204, 0.7479947485785097, 0.20758951699810702, 0.7015694071908234, 0.2812680069271609, 1.2549849112614186, 0.5233853279791925]}, 'log_odds_source': {'aa_anchor': 'anchor', 'aa_anchor10': 'anchor@10', 'aa_anchor11': 'anchor@11', 'aa_anchor8': 'anchor@8', 'aa_anchor9': 'anchor@9', 'aa_tcr': 'tcr', 'aa_tcr10': 'tcr@10', 'aa_tcr11': 'tcr@11', 'aa_tcr8': 'tcr@8', 'aa_tcr9': 'tcr@9', 'aa_tcr_c': 'tcr_c', 'aa_tcr_m': 'tcr_m', 'aa_tcr_n': 'tcr_n', 'kmer_llr': 'pair'}, 'logistic': {'coef': [0.11949579886625111, -0.026388547388491116, 0.03688791951418523, 0.12315135708771073, 0.00793113416429338, 0.032785084028538934, -0.0491770492059215, 0.1283027809168922, 0.0010358114328155197, 0.03544359835457312, -0.09272877328007892, -0.029929037150132728, 0.005234358233192701, 0.054646426854836824, 0.05481812119007732, 0.004802849974804372, 0.031791737568215295, -0.8795515088055536, 0.13246571365233376, 0.11209166951408028, 0.41632421759666804, 0.3710860928169145, 0.21945608459259616, 0.2831765986455691, 0.09310214697766131, 0.12584078767223586, 0.2711009144487395, 0.10677716086875681, 0.2388095572730888, 0.6543935078687705], 'intercept': -3.8119421736439913, 'tau': 4.0}, 'n': 464161, 'n_immunogenic': 14712, 'paratope': {'A': [0.0468, 0.3668], 'C': [0.1994, 0.5024], 'D': [0.1813, 0.5419], 'E': [0.1014, 0.313], 'F': [0.0754, 0.3616], 'G': [-0.0, 0.2369], 'H': [0.2189, 0.5621], 'I': [0.0486, 0.416], 'K': [0.1234, 0.4516], 'L': [-0.0251, 0.1864], 'M': [0.0345, 0.3921], 'N': [0.1193, 0.3291], 'P': [0.0469, 0.2501], 'Q': [0.0384, 0.3411], 'R': [0.0515, 0.421], 'S': [0.0124, 0.2607], 'T': [0.1307, 0.3778], 'V': [0.0143, 0.2991], 'W': [0.0847, 0.4427], 'Y': [0.0172, 0.2474]}, 'prevalence': 0.03169589862138353, 'seed': 20260817, 'standardizer': {'mean': [3.585166758896088, 6.808608481871727, 9.314074211318918, 4.7808659475268716, 1.3706172407505959, -1.1956991886307702, 5.437991241121321, -0.17470653501694436, 0.36020443337546376, 13.744965561518644, 10.890707254593467, 0.06078909482557044, 0.3367987089608995, 1.993329900616381, 1.320692604505764, 0.5490802401177581, -0.1541268070176281, -0.14749849482590252, -0.02042398402213005, -0.014461619053264312, -0.10887835324113007, -0.08397399955209743, -0.02768170726510107, -0.041256419055064476, -0.011095587661907503, -0.023404604320161237, -0.042033855105112226, -0.05205149786415519, -0.07093132420268783, -0.3009136507757868], 'std': [18.338465868430596, 15.272311296957708, 0.7785394170589226, 14.28094041797769, 11.073130188216279, 13.196624288892432, 9.981009433871744, 2.103026551096585, 1.8789707926397332, 2.040641126152638, 2.5470821127038956, 0.029406827409666877, 0.04828902479536415, 1.064927514941948, 0.5583973810425537, 0.23698125064685832, 0.5361092115927888, 0.5265313599296632, 0.20078005110769423, 0.16370954872505514, 0.4592756393360549, 0.40885275689301376, 0.24535545130466027, 0.295520132855347, 0.14842445509414, 0.2117117554672323, 0.27496149723350805, 0.31126309236162925, 0.37329783588644283, 0.9119874447264713]}}#
The human class-I table, for callers that predate the species split.
- mhcmatch.complement.TABLES: dict = {'human': {'alphabet': 'ACDEFGHIKLMNPQRSTVWY', 'anchors': [0, 1, 2, -2, -1], 'arm': 'chowell_rebuilt/human', 'features': ['pc1', 'pc2', 'length', 'pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr', 'mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr', 'kd_run_max', 'kd_run_n', 'kd_run_frac', 'aa_anchor', 'aa_tcr', 'aa_anchor8', 'aa_tcr8', 'aa_anchor9', 'aa_tcr9', 'aa_anchor10', 'aa_tcr10', 'aa_anchor11', 'aa_tcr11', 'aa_tcr_n', 'aa_tcr_m', 'aa_tcr_c', 'kmer_llr'], 'fits': {'em': {'immunogenic': {'mean': [-0.06616284744384984, 0.033257527759931516, 0.7425823626840402, 0.009944481875500932, -0.09744510738436167, -0.10270366444392376, 0.15899611033358801, -0.021620240473703266, 0.0453647892974406, 0.017723454003337618, 0.5022583696379227, 0.029067682576747078, -0.011969435994234565, 0.15024378096183383, 0.23894006374608343, -0.01841692959247472, 0.03140174804271794, 0.03149765470224604, 0.10795577805722431, 0.09403711588635565, 0.285996840399637, 0.25304989654414084, -0.4283587039003903, -0.5300476240404369, 0.08119225093371836, 0.1213505372650956, 0.044784546090369255, -0.0738356488326727, 0.03588980019953098, 0.026272257168881406], 'var': [1.1516692663151664, 1.1084522262605425, 0.19860447709908008, 1.0869453773015285, 1.0310605727504196, 1.1253315125231618, 1.1444420360343173, 1.0927980035842606, 1.1519899538076086, 1.1071525977980514, 0.6586179938029688, 0.8947099943206341, 0.8880246294471223, 1.1369114207966735, 1.0532738268222466, 0.8999572184564513, 1.1610496236862218, 1.3275002839607049, 0.09581536907339404, 0.10353666897503208, 0.2807555549028624, 0.34450507409073905, 4.564916481275358, 4.441787944835681, 0.10721622172558731, 0.1657772376791387, 1.0902456488673937, 1.559667039606054, 1.156856098160668, 1.4032918413480189]}, 'mixing': 0.20847508289499742, 'non_immunogenic': {'mean': [0.017426242443356672, -0.008759504226350764, -0.19558439193947458, -0.0026192184712624652, 0.025665492520388408, 0.027050512871895235, -0.04187704842323269, 0.005694427714860107, -0.011948364485953862, -0.004668076093047893, -0.13228687181118653, -0.007655965597128482, 0.003152559202088637, -0.03957182397370027, -0.06293301514573582, 0.004850726541647888, -0.00827072134405988, -0.008295981633823916, -0.028433836122047553, -0.0247678817254575, -0.07532702220934928, -0.06664932084712008, 0.11282287439673347, 0.1396061197470374, -0.021384748449546936, -0.03196180280491581, -0.011795537647336141, 0.01944713637972825, -0.009452802950278383, -0.006919694974540168], 'var': [0.9585960947283462, 0.9710673715761963, 1.0275840819740554, 0.9770670750613483, 0.9886594572583193, 0.9634797671746005, 0.9535443010666477, 0.9754029519903206, 0.9592834760347848, 0.9716731960596869, 1.0059726022242426, 1.0274505716934175, 1.029444872547767, 0.9564283707210262, 0.9669707503174307, 1.0262368262064476, 0.9572539814923111, 0.9134115157580363, 1.234269793626622, 1.2331716577819047, 1.1622201173805702, 1.1513392221513465, 1.000785366389977e-08, 1e-08, 1.232951486845345, 1.214820890534065, 0.9755633505553665, 0.850778526544332, 0.9582579947100118, 0.8935496732434686]}}, 'supervised': {'immunogenic': {'mean': [0.48458438715948515, 0.06298948453314321, -0.029654583164330322, 0.39128355525065656, 0.1014682115471637, 0.24996067426166946, -0.016188312542218187, -0.3491597579256158, -0.2717712962243248, 0.3675560892923596, 0.18135148896737566, -0.042275395210148434, -0.006615296723916411, 0.20049069236962921, -0.019114653444797746, 0.19690361209248805, 0.6315466359237181, 0.6564536683522727, 0.14271720635612214, 0.12582837414290413, 0.5589047848821895, 0.5188691769230301, 0.26888146312369027, 0.3491639056385894, 0.11709064950644237, 0.18159236648548993, 0.354008713848338, 0.3918615009934163, 0.4387206547144678, 0.7714373541380991], 'var': [1.3930230496478904, 1.1602006424759057, 0.6030177894148014, 1.2263787688544587, 0.9778518154042607, 1.0985013400029648, 1.1708403457939764, 1.1354769271622067, 1.095528758069639, 1.3172639009536289, 0.9862524597940981, 1.054503535191982, 1.0096275333834448, 1.1582998494436165, 0.9218909598402559, 1.0523029114404265, 1.5847444990351371, 2.0242241415857265, 0.6287863446403338, 0.6798051670099481, 1.7587954787066233, 2.182601004693286, 1.5692679535176246, 2.000733627789036, 0.7036790214229164, 1.0860954313051818, 1.986962383550572, 1.9224728597493752, 1.669564234269702, 2.1575359637691927]}, 'mixing': 0.03169589862138353, 'non_immunogenic': {'mean': [-0.01586210115917164, -0.00206186084840703, 0.0009706957352459654, -0.012808046440973456, -0.003321400933777193, -0.008182066129269654, 0.0005298987295771788, 0.011429190761580657, 0.008896002238425973, -0.01203136548461819, -0.005936253291110097, 0.001383817995791796, 0.00021654124361396656, -0.006562744752223986, 0.0006256878566463744, -0.006445327369643187, -0.020672677228585953, -0.021487969422078916, -0.004671621340607238, -0.004118792210880399, -0.018294861475243836, -0.016984359362011706, -0.008801408136358138, -0.011429326530386547, -0.0038327766566152344, -0.005944138035086989, -0.011587913641138271, -0.012826964578101358, -0.014360824636687601, -0.02525177796385514], 'var': [0.9791968914743827, 0.994621969619256, 1.0129648653826153, 0.9874142364597183, 1.0003769476971038, 0.9946635842540551, 0.9943989636734838, 0.9914441481583227, 0.9943762110655474, 0.9850479244412838, 0.9993382267703844, 0.998155507163318, 0.9996833889205654, 0.9934594770280398, 1.0025444345477688, 0.9969773070126338, 0.9673762074590005, 0.9519060687198755, 1.011462555639168, 1.0099458548589053, 0.9646022565242619, 0.9521883135121277, 0.9789219266960645, 0.9631212447947094, 1.0092361349801686, 0.9960670700841824, 0.9634568602941078, 0.9646134000808131, 0.9715762636565185, 0.9419920898218387]}}}, 'kd_threshold': -0.8500000000000001, 'log_odds': {'aa_anchor': [-0.17195912826031323, 1.8451940731045058, -0.3383326314979689, -0.5309006197745685, -0.10548669598002158, 0.14203178413119355, -0.4438454689097804, 0.08188954910316015, -0.12307214034493219, 0.2531804504988795, 0.5553004601198759, 0.1305669975254693, -0.035223985277909264, -0.2042807334750738, -0.3067910455939242, -0.08833926936175907, 0.060135842421909835, 0.15719944510128148, 0.1361199476438575, 0.06362187397922714], 'aa_anchor10': [-0.2022520811905535, 1.6582365394881133, -0.21191815390528346, -0.4273384044275228, -0.19388742686109328, 0.15314775491206678, -0.5778472050268788, 0.0421627706836043, 0.012321574897647736, 0.19777070487777615, 0.5231004866997857, 0.15504047994184544, -0.03311038630277352, -0.07875792042935048, -0.3743890548054498, 0.10555445424948529, 0.1198990535270954, 0.2992389719106292, -0.15750866458101953, -0.04342257958925888], 'aa_anchor11': [-0.16065601260790485, 1.6174140397801606, -0.27465433156603636, -0.14863832184342884, 0.06167147632298109, -0.07367551716753917, -0.1010235917043052, -0.028861400128323833, -0.230528091540513, 0.03736191683968837, 0.36971437324133927, -0.07392299797558932, 0.21726648668495496, -0.4227304106649723, -0.24945682828088422, 0.09779000750694689, 0.15277128278115182, 0.06643388308785925, -0.058436472170595, 0.2639780408845138], 'aa_anchor8': [-0.13629873025390715, 2.6257263950162546, -0.7974534382634779, -0.4950730959299041, -0.08115130790852332, 0.20165648504369704, -0.3868574556412123, -0.028843247661880955, -0.09865958906304018, 0.04648481655892489, 0.4944428841186088, -0.03053823305143455, 0.22466831742077442, -0.16959297760528402, 0.40538983996537015, 0.039470325723848454, 0.11370982206898317, -0.021221858621291112, 0.4424026125901088, 0.16730566685241888], 'aa_anchor9': [-0.17862695719588118, 1.9054169060709576, -0.27658399902294883, -0.5987303966114133, -0.09066545816289695, 0.17642552217525376, -0.43654794822798815, 0.10186616378673508, -0.1596994472827098, 0.2892885908046945, 0.5735795678820441, 0.12129354135612713, -0.06062207561426547, -0.23424807832612737, -0.31999365078425024, -0.18613321931811377, 0.021565292739611497, 0.12556288364885804, 0.3091340996907128, 0.06728014303186924], 'aa_tcr': [0.1269650827576445, 2.0508651169873318, -0.1278672175910449, -0.3965923638827822, 0.312048362043968, 0.05456623245180037, -0.28960607601331123, -0.08651718328776914, -0.3822461012648932, 0.07643556676200847, 0.5557280465342154, 0.13994712385397712, -0.16182712771524965, -0.39178104663270386, 0.13865507523536413, -0.09380087552214711, 0.1356944452714206, -0.06006938642577664, 0.8066684257701158, 0.3563580715841401], 'aa_tcr10': [0.11041695229878457, 1.9695372867631002, -0.27353380791584314, -0.5717224209998402, 0.47369453970111497, 0.004229081709119953, -0.31082693732559763, 0.13940711867842825, -0.33756958013949534, 0.13552090933857208, 0.5602298437823778, -0.003674095757914664, -0.2440546440131821, -0.467221239478429, 0.12067295968179748, 0.030798063215855098, 0.08170025394600078, -0.006905284398639466, 0.9126222570779312, 0.31409292051353965], 'aa_tcr11': [0.034307092582310794, 2.1276411057690137, -0.32117676334623724, -0.22987957712244578, 0.2692742510626802, -0.09214137696067315, -0.1046682419251912, 0.038192708748248094, -0.4901140111953084, -0.00674018569624657, 0.4210891741674745, -0.17233903117264138, 0.08571050978170502, -0.3319401982231489, 0.3854749883302704, -0.06271559344998678, 0.13828374712029357, 0.002301475026118993, 0.6788054088668662, 0.5037665937803593], 'aa_tcr8': [-0.29697431523976325, 2.442364901686199, 0.21768694002602063, -0.06801892407850607, 0.6132789072842479, 0.2080022184566439, -0.6029684606121299, -0.18840420706895777, 0.17071023514109074, -0.18143216168115472, 0.6777731326144099, 0.040877527639617384, -0.004906500768449895, -0.31094769397583377, 0.07652626579107435, -0.25894400383096405, -0.009474325818346951, -0.22749187598759413, 1.0455387632473006, 0.6328266631195407], 'aa_tcr9': [0.17945770071163913, 2.0619637052103297, -0.029981044514179267, -0.3654405343298275, 0.20675501336828495, 0.12656639409957382, -0.2782317593896546, -0.1917114470680965, -0.4077037365276799, 0.07032137075823908, 0.5592548998295923, 0.23170647822326806, -0.17681564687095896, -0.3842853433911029, 0.12292084043406026, -0.12364457611197066, 0.15333852825101246, -0.07875146626659957, 0.7355366135993782, 0.31904688155671534], 'aa_tcr_c': [0.16333550161931853, 1.8707098333760452, -0.1148505690154531, -0.4670323791682849, 0.3314358608017822, 0.0754790845130513, -0.3641240402505592, -0.12193367623157014, -0.43474646569214226, 0.08468515767806473, 0.5125890659921324, -0.016602970922443117, -0.0327462463587751, -0.47481451107669725, -0.20126018609099328, -0.058127301973135204, 0.1656247749868056, -0.05152218711154033, 0.797024826894055, 0.2857599761897811], 'aa_tcr_m': [0.008107412078389054, 2.177840329799901, -0.04579142184387397, -0.3610209167881262, 0.28258237212748005, -0.017691991710642174, -0.2506611603667275, -0.06943184967596983, -0.37064317866220176, 0.14042794973508954, 0.4553026314937103, 0.07890375752309797, -0.4131976989638213, -0.29321347265154074, 0.42957503508856787, -0.156967566477761, 0.1484842222490097, -0.10365523055072456, 0.797835467602722, 0.34723132268913437], 'aa_tcr_n': [0.1954528096969046, 2.2986892206804814, -0.1450903656778295, -0.3338366929044301, 0.26025426122320683, 0.14542760252377196, -0.17975415714020437, -0.0486790082582389, -0.3081890476685425, -0.10720632699825039, 0.7943039763213995, 0.42803131629758884, -0.10253553972469698, -0.3429217934402682, 0.27063702790624644, -0.07733311107678142, 0.0014274289169473597, -0.06758518949137082, 0.8294865346693872, 0.5232115791339824], 'kmer_llr': [-0.017874478710091957, 2.4432388959174407, 0.06144394763864014, -0.6114595687515676, 0.35940076483606376, 0.15473638788561495, -0.24590417844759394, 0.0811799423628008, -0.6759640849584203, 0.07704089966355099, 0.6876447031574555, 0.3596030991356125, 0.2937716689416785, -0.4560810125888013, -0.05588165663171196, 0.18203482103036883, -0.03050596545520179, -0.00279635952314905, 0.9071851400401734, 0.8170942241911412, 2.039704290339335, 3.9193938320740926, 2.1948431125394867, 1.4795343601766895, 2.914371891493655, 2.0803334022280726, 1.4968059466853498, 1.8137364482633638, 1.5286670487543343, 2.1504166725145826, 2.725471363601658, 2.170793857306818, 2.0533775922395456, 1.4943698918874695, 1.7214750513083974, 2.0980755606044923, 2.3600115901071916, 2.1948431125394867, 2.725471363601658, 2.4070176324831225, -0.06015697214592741, 1.5358872967278216, -0.020396323000508865, -0.3722189044643205, 0.0833748258927951, 0.16644368827812972, -0.6500520212923888, -0.1094628095243877, -0.42434483187710015, -0.09896229280251223, -0.02967895563648515, -0.301205713801755, -0.1342283435993199, -0.4390103892007451, 0.12164621760607641, -0.4682697866954566, -0.2576821277454737, -0.23685898356808455, 0.4260161636165396, 0.025040878014332968, -0.20391580887762562, 1.155254164320839, -0.16269973320831443, -0.5903989533073615, -0.01164859093226589, -0.3190510741217656, -0.6669814276717467, -0.5312854263356224, -0.8575558118742936, -0.29560495805036346, -0.21291865568535684, -0.4031722756891618, -0.5110056567435519, -0.5330257352102787, -0.28319065796075105, -0.4342059873718789, -0.2791864755578377, -0.5135232940989543, 0.4797944028794916, -0.11533832734649341, 0.18963524221534822, 2.2791842609732385, 0.12024810700218147, -0.4669025205210362, 0.8707370952122142, 0.27657714973996583, -0.010262266248969532, -0.1378363243030014, 0.06099241079308726, 0.35862932063901276, 0.8689172584952276, 0.19842490667132573, 0.12886468626942893, -0.6399418175702003, 0.09661196042388909, 0.13491852467867016, 0.46800670588808124, 0.26182880456709867, 0.49187914209456274, 0.2553447641253772, 0.2732980410574122, 1.691397596071118, -0.47554477282783836, -0.5523655306917297, 0.3057323636287901, -0.07565643380480847, -0.33350915729699615, 0.41660467136448087, -0.5703655024026721, 0.0015026910617672584, 0.9938760213270443, -0.09863384240278705, -0.2547437887300674, -0.5284531210850592, -0.15679708037862028, -0.12993656045995206, 0.0695155347436236, 0.2597475332304313, 0.9809829427137178, 0.08744478480777396, 0.1801509512494528, 1.9349500190165356, 0.04020119083126694, -0.4600896779808812, -0.1250544857473077, 0.026029537036047934, -0.004052055207287353, -0.25704261886648005, -0.19720133908282822, -0.406360903549948, 0.05132271417512868, -0.4163963601995846, -0.8088785510479832, -0.8793703443232737, -0.060327590374930296, -0.5826355949944864, 0.010560691468139538, -0.5770769369906583, 0.5156995702133491, 0.5719218497681, 0.13442325317276005, 1.4795343601766895, -0.09718222023955647, -0.29812980896112595, 0.03677944408736078, 0.4200697279698362, -0.6745348326903446, 0.01311497127323591, -0.15232161208764694, 0.17450543403149865, 0.0032365275074912603, -0.026570406782087552, 0.020831990687240065, -0.6000141963248016, -0.2664839449216254, -0.29103717070534874, -0.03128501393214478, -0.25070275826406174, 0.6184530541515869, 0.04645821535399275, -0.29901032323831966, 1.9775992833524647, -0.4109574318347695, -0.9157379884941479, -0.14277183178377317, -0.26833598154274885, 0.18466705688795493, -0.9052957635354284, -0.46320955932057384, -0.44721746329507095, 0.20049772121677112, -0.012773128945724466, -0.7254719311318905, -0.7016580441538798, 0.16560676285823117, -0.5086074567861996, -0.5634689233436116, -0.7159228171868346, 0.7474728999948432, -0.16579605042245227, 0.2088597839639048, 2.2608585829907986, -0.08595264814181114, -0.4283145777283499, 0.3384104329689972, 0.23319170397478928, -0.27702095501005264, -0.19597809734008464, -0.4126038907299048, 0.28434529037365763, 0.5224664233499388, 0.19306555387259206, -0.22142578074321229, -0.2131158045146444, -0.1893219674469817, 0.13683051990041495, 0.16909147726837315, 0.14037577558778302, 0.8627115828785152, 0.3932640145762685, 0.5117884541166067, 2.812482740591287, 0.16984918901525958, 0.9683320950202061, 0.3275760908032872, 0.5782072255830606, -0.018758494078054966, 0.4157688028387474, -0.21633256792677802, 0.9722708879772854, 0.9950808407498943, 0.2135519145675886, 0.2977231276536063, -0.2733810372885772, 0.21881228201375258, 0.6338097105870455, 0.4633491342207927, 0.4673915105441626, 2.7705917988821263, 0.9515467620096363, 0.25985995105695014, 2.2926072813053793, 0.21057035794546586, -0.6234391816140699, 0.015704423880392504, 0.2957573725329903, -0.08443321382909108, 0.02530556602359635, -0.2662674394063833, 0.29230182296704044, 0.698698767504009, 0.4002472811006461, -0.20913837658789536, -0.011987105344703153, 0.14851919478157338, -0.042834696846123066, 0.9828618354852141, 0.004139788340288497, 0.06787938376957658, 0.6563926089573675, -0.39889212823543296, 1.689772900344117, -0.2837869060121019, -0.4532860480037648, 0.02121212290199459, -0.6092240560466005, -0.44553778359564067, -0.1557006973960835, -0.8005120159977759, -0.37802716736524555, 0.016846614939833415, -0.63370302828565, -0.58476172709838, -0.23234167509003978, 0.7829774258778581, -0.44081255909242856, -0.06577411175868697, -0.2518356231607708, 0.4401994618092884, 0.10019908895714114, -0.3626984770200412, 1.9231248910767196, -0.42532026771580966, -0.8856544276422138, -0.1523018024581857, -0.2512644092616787, -0.41089203969207944, -0.20638883335792269, -0.9792699948127455, -0.08423249122705201, 0.9711088747135914, -0.17869371642684317, -0.6315087157361994, -0.5934173705977743, -0.8132119992825038, -0.2105109055465597, -0.19073890002552663, -0.4876487987391185, 0.2831243282324536, 0.43398661312012443, 0.3882139879839652, 1.9015539518585172, -0.13935763973408122, -0.2543328423178828, 0.32200297587685434, 0.42514445749120355, 0.05296609991885948, -0.46637578887862396, 0.4897516558162911, 0.2261407326217224, 0.36637406034426956, 0.3375066590058875, -0.4893184549980303, -0.25521777163119364, 1.3221079424272144, 0.22204076525077543, 1.2314716770646736, -0.17816670719198413, 1.2579515625591808, 0.2230423631214551, 0.1315835285550282, 2.4293788126412803, -0.21679488249928092, -0.7731173640686473, 0.41838995642723287, -0.08789326914721851, -0.30670971966036564, -0.29439267003221037, -0.43731499461549994, 0.025641995644305027, 0.49776376954036206, 0.010010671608962518, -0.06112280519738267, -0.6232079837770517, -0.060507698697708, -0.23712292526221024, -0.1335498644104316, -0.27048212442066166, 0.7524592847685527, 0.30650561923395525, 0.09688533385369436, 1.4676193255988172, 0.033254665706303754, -0.1744037055573351, 1.0827025320703392, -0.22264505601066986, -0.31770879204103686, 0.06972534701581701, -0.21007189036285556, 0.4462987112323553, 0.9075421766402174, -0.02262633862982799, -0.02919820239785853, -0.26608553854185324, 0.2453040591775535, 0.07922493041392897, 0.0502389381938988, 0.3183243931040858, 0.9321978895423877, 0.4479342094767844, 0.4037122973838114, 1.835613888795657, -0.11777550944413662, -0.46008228388285133, 0.31984551668843153, -0.26721573757716666, -0.5830871642051401, -0.2016738002887628, -0.5587239273871853, 0.1299446165366298, 0.3357036753999454, -0.04894051319963921, -0.1246659262448313, -0.6427509929034283, -0.09130946412794838, -0.1943142795008006, -0.07826616232160521, -0.009635638888066289, 0.8418035234235655, 0.2337769057593837, 0.6298569626762198, 2.84807368569399, 0.8536691867000652, 0.4567878222832933, 0.9236005847692823, 0.5481569462247675, 0.45273923375729197, 0.9325221262380134, -0.04853680529235049, 1.1551208398756039, 1.1464926586522655, 0.9857682563296555, 1.1278679088145598, 0.5417474734843406, 0.9153627557054067, 0.7614537239189536, 0.5693581745078689, 0.19064140504705662, 0.9337118943736034, 1.802644660747717, 0.2647519394276605, 2.5658412180097745, 0.24630008389840174, -0.0675636803470443, 0.6704483492699671, 0.47248321174698393, -0.07961912154266049, -0.007189324685382914, 0.22771657133580803, 0.3860722974848949, 0.5730046869069314, 0.6245516126134572, 0.15336257357971128, 0.09052081593905204, 0.7479947485785097, 0.20758951699810702, 0.7015694071908234, 0.2812680069271609, 1.2549849112614186, 0.5233853279791925]}, 'log_odds_source': {'aa_anchor': 'anchor', 'aa_anchor10': 'anchor@10', 'aa_anchor11': 'anchor@11', 'aa_anchor8': 'anchor@8', 'aa_anchor9': 'anchor@9', 'aa_tcr': 'tcr', 'aa_tcr10': 'tcr@10', 'aa_tcr11': 'tcr@11', 'aa_tcr8': 'tcr@8', 'aa_tcr9': 'tcr@9', 'aa_tcr_c': 'tcr_c', 'aa_tcr_m': 'tcr_m', 'aa_tcr_n': 'tcr_n', 'kmer_llr': 'pair'}, 'logistic': {'coef': [0.11949579886625111, -0.026388547388491116, 0.03688791951418523, 0.12315135708771073, 0.00793113416429338, 0.032785084028538934, -0.0491770492059215, 0.1283027809168922, 0.0010358114328155197, 0.03544359835457312, -0.09272877328007892, -0.029929037150132728, 0.005234358233192701, 0.054646426854836824, 0.05481812119007732, 0.004802849974804372, 0.031791737568215295, -0.8795515088055536, 0.13246571365233376, 0.11209166951408028, 0.41632421759666804, 0.3710860928169145, 0.21945608459259616, 0.2831765986455691, 0.09310214697766131, 0.12584078767223586, 0.2711009144487395, 0.10677716086875681, 0.2388095572730888, 0.6543935078687705], 'intercept': -3.8119421736439913, 'tau': 4.0}, 'n': 464161, 'n_immunogenic': 14712, 'paratope': {'A': [0.0468, 0.3668], 'C': [0.1994, 0.5024], 'D': [0.1813, 0.5419], 'E': [0.1014, 0.313], 'F': [0.0754, 0.3616], 'G': [-0.0, 0.2369], 'H': [0.2189, 0.5621], 'I': [0.0486, 0.416], 'K': [0.1234, 0.4516], 'L': [-0.0251, 0.1864], 'M': [0.0345, 0.3921], 'N': [0.1193, 0.3291], 'P': [0.0469, 0.2501], 'Q': [0.0384, 0.3411], 'R': [0.0515, 0.421], 'S': [0.0124, 0.2607], 'T': [0.1307, 0.3778], 'V': [0.0143, 0.2991], 'W': [0.0847, 0.4427], 'Y': [0.0172, 0.2474]}, 'prevalence': 0.03169589862138353, 'seed': 20260817, 'standardizer': {'mean': [3.585166758896088, 6.808608481871727, 9.314074211318918, 4.7808659475268716, 1.3706172407505959, -1.1956991886307702, 5.437991241121321, -0.17470653501694436, 0.36020443337546376, 13.744965561518644, 10.890707254593467, 0.06078909482557044, 0.3367987089608995, 1.993329900616381, 1.320692604505764, 0.5490802401177581, -0.1541268070176281, -0.14749849482590252, -0.02042398402213005, -0.014461619053264312, -0.10887835324113007, -0.08397399955209743, -0.02768170726510107, -0.041256419055064476, -0.011095587661907503, -0.023404604320161237, -0.042033855105112226, -0.05205149786415519, -0.07093132420268783, -0.3009136507757868], 'std': [18.338465868430596, 15.272311296957708, 0.7785394170589226, 14.28094041797769, 11.073130188216279, 13.196624288892432, 9.981009433871744, 2.103026551096585, 1.8789707926397332, 2.040641126152638, 2.5470821127038956, 0.029406827409666877, 0.04828902479536415, 1.064927514941948, 0.5583973810425537, 0.23698125064685832, 0.5361092115927888, 0.5265313599296632, 0.20078005110769423, 0.16370954872505514, 0.4592756393360549, 0.40885275689301376, 0.24535545130466027, 0.295520132855347, 0.14842445509414, 0.2117117554672323, 0.27496149723350805, 0.31126309236162925, 0.37329783588644283, 0.9119874447264713]}}, 'mouse': {'alphabet': 'ACDEFGHIKLMNPQRSTVWY', 'anchors': [0, 1, 2, -2, -1], 'arm': 'chowell_rebuilt/mouse', 'features': ['pc1', 'pc2', 'length', 'pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr', 'mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr', 'kd_run_max', 'kd_run_n', 'kd_run_frac', 'aa_anchor', 'aa_tcr', 'aa_anchor8', 'aa_tcr8', 'aa_anchor9', 'aa_tcr9', 'aa_anchor10', 'aa_tcr10', 'aa_anchor11', 'aa_tcr11', 'aa_tcr_n', 'aa_tcr_m', 'aa_tcr_c', 'kmer_llr'], 'fits': {'em': {'immunogenic': {'mean': [0.14312370634486676, 0.16557843346884654, 0.5128065164897488, 0.12339035765818519, 0.07706062311119909, 0.06952353629098441, 0.1574738613288356, -0.16171673327480834, -0.1194230737471546, 0.08672479156214237, 0.48834034342956, -0.043463111475816586, -0.056797325534043576, 0.22450608570720082, 0.21658895480452678, 0.08411004076461961, 0.23662124927264536, 0.2148995074580367, 0.24544731929071378, 0.2476361813953243, 0.3036928644233879, 0.3494929528372943, -0.3806064515522984, -0.36322719184900176, 0.1812460089008381, 0.20797141806254732, 0.15787959100380852, -0.0636343821800119, 0.23362136475552556, 0.18486420078401353], 'var': [1.3366292007778482, 1.1260127995542892, 0.5397194970856732, 1.1464691638520583, 1.036263008956585, 1.2249919507241267, 1.1616885291656642, 1.1251966167126373, 1.2196691562094661, 1.2003553544848022, 0.7856801376368169, 0.920876791408655, 0.9052971398126473, 1.2140570461105102, 1.0359819895084563, 0.9535086599206211, 1.1940143145796225, 1.688603480109305, 0.209819228402907, 0.3471671255512018, 1.0409426960311055, 1.3726619114252188, 3.9374253702654425, 3.954484631726769, 0.15515853974288568, 0.39535124691330403, 1.6292043790268873, 1.5446271786490802, 1.4819835281237557, 1.6470518813210646]}, 'mixing': 0.24221349849914206, 'non_immunogenic': {'mean': [-0.045747045590407244, -0.05292431518252884, -0.16390983498140319, -0.039439618085368305, -0.024631110587640812, -0.02222200965538523, -0.050333827270721665, 0.05168998874322181, 0.03817154361621429, -0.027720096794128773, -0.156089641087793, 0.013892240446834448, 0.018154293980941857, -0.07175953166992306, -0.06922895614473894, -0.026884336408952023, -0.07563193655738568, -0.06868895318676156, -0.07845303892968884, -0.07915267127605743, -0.09707023154348245, -0.11170944671072623, 0.12165434459341934, 0.11609936138149216, -0.0579321877836696, -0.06647450786862545, -0.050463511824108666, 0.020339642237095754, -0.07467307476381334, -0.05908868095494192], 'var': [0.8837619659862849, 0.9481580680597558, 1.0362003405333446, 0.946761697258523, 0.9859043820731401, 0.9260464126732396, 0.9378592831108543, 0.9489520642289404, 0.9237708935277261, 0.9327874264472823, 0.9679148477497845, 1.0244935996266071, 1.028909469481716, 0.9103204564222419, 0.9687121098501856, 1.0118761649627572, 0.9143703416756251, 0.7604202900895842, 1.2271568188380395, 1.182800668684306, 0.9480112139326297, 0.8293643401427054, 1e-08, 1.0000604388067371e-08, 1.256183024584032, 1.1750219909154196, 0.7883718670036672, 0.8242112460149019, 0.8229209257001397, 0.778766088689481]}}, 'supervised': {'immunogenic': {'mean': [0.6210075688060877, -0.01945745406543652, -0.0456845593184066, 0.49096696827720937, 0.009284831870431044, 0.34980986567909433, -0.03776810252455127, -0.5040323448856722, -0.3778243849799156, 0.4491687706693304, 0.22726993717797078, -0.03666951952972699, -0.012015389299309941, 0.2530790343049985, -0.03606705495612654, 0.23105006537598585, 0.6826210046999133, 0.6971941201822071, 0.31680177469563386, 0.33797595438939265, 0.5000992519576556, 0.5171573396090823, 0.3041438640133869, 0.32543117539887206, 0.2156688853186513, 0.25415039199985634, 0.3250055477442498, 0.39990527349316524, 0.520995487946946, 0.7740105907252691], 'var': [1.2813623788649668, 1.0944436080305044, 0.609586325629402, 0.9601960052068509, 1.0038183317945322, 1.2134318049465322, 1.134762221728634, 0.9696211866091288, 1.1479447936088318, 1.1216244835317215, 1.0438402680967365, 1.047034053169049, 0.9965362764094995, 1.293213682273709, 0.8470129928809526, 1.1029711277894307, 1.1194383746445848, 1.8580693977184501, 0.4555433654718886, 0.7542204950732135, 2.205020371283901, 2.4855523448589607, 1.21588981654603, 2.0986188692468333, 0.34107164739227974, 0.8404477191426631, 2.1634449925929142, 1.2683586944065994, 1.7429859405292798, 1.6264339438163766]}, 'mixing': 0.10933389902418328, 'non_immunogenic': {'mean': [-0.0762319108661606, 0.0023885037453731045, 0.005608017404064402, -0.0602687504049052, -0.0011397614314362387, -0.04294098146310941, 0.004636231134459426, 0.06187259337733515, 0.0463799094980813, -0.055137804125906825, -0.02789856752774337, 0.004501374354694335, 0.0014749515660115414, -0.03106676851350549, 0.0044274186929887695, -0.02836259793617407, -0.08379528076556317, -0.0855842065312017, -0.03888906651696485, -0.0414883072672556, -0.061389785752149, -0.06348375478362453, -0.03733524210748737, -0.039948370361688024, -0.026474478038686373, -0.03119828324601642, -0.0398961223520579, -0.049090453474608615, -0.06395490746626635, -0.09501382805216343], 'var': [0.9123094171127833, 0.9883543876943038, 1.0476376718507163, 0.9716638707894784, 0.9995194095793454, 0.9569350395998009, 0.9832606509002214, 0.9687151625307648, 0.9621644611036755, 0.9572636157856593, 0.9874995544093047, 0.9940410108245211, 1.0004053037245377, 0.9551790139305536, 1.0186006765232563, 0.9800021523476821, 0.921116261344951, 0.8276740651836324, 1.0530024117472894, 1.014427394638681, 0.8176077978599909, 0.7807794530560427, 0.9607492076296906, 0.8505425102712357, 1.074476276494231, 1.0106834952367647, 0.842622922007879, 0.9450161810747627, 0.8713841720317269, 0.8405326478905414]}}}, 'kd_threshold': -0.8500000000000001, 'log_odds': {'aa_anchor': [-0.00996606926865784, 2.5951294287960796, -0.7103218600532224, -1.0686311703318228, 0.10898980233447686, -0.0054052232940566824, -0.35002509664809267, 0.1299308601996776, -0.25575430871011307, 0.15914014219243922, 0.8438475469129401, -0.2154052812184819, 0.31073983106715586, -0.42513769705864757, -0.24503836658740097, -0.014759415233200457, -0.024221716914019975, -0.030625644094542714, 0.6050752830524884, 0.2338794791469785], 'aa_anchor10': [-0.17541957143717246, 2.7326276976177146, -1.0432479542571085, -0.9809849679193827, 0.17146525425650072, -0.1838256991781848, -0.3155083646078545, 0.2018746597042953, -0.16829576570932137, 0.23424666492180268, 1.1463362685451473, -0.3809655812624979, 0.12689690853304647, -0.538607351649234, -0.09942709206265876, -0.010023793194412267, -0.27514906857225396, 0.06614841191258813, 0.5991189346676027, 0.3723911710484873], 'aa_anchor11': [-0.3168130445829238, 3.0735274921799496, -1.6538603265323912, -1.8563577734433112, 0.21830068058655527, -0.23502282709736422, -0.5955136197143225, 0.3966240200745763, -0.16122168184454067, 0.30187204112316723, 1.2017253152783587, -0.3825464997080399, -0.053606247740738855, -0.27308169919780223, -0.3357182226650339, 0.04941392608773443, -0.22768212012676958, -0.019387620901657687, -0.37115500142194424, 0.7909691895357787], 'aa_anchor8': [0.02237782061173199, 2.0225972923790074, -1.0656591519964262, -1.4188199223425557, -0.1212825504875088, 0.21822242052087804, -0.2954493358225605, 0.2936764011221187, -0.385223984219631, 0.12734530110734155, 0.714744471202049, -0.04117427452429556, 0.23875121723027437, -0.47595188101171226, -0.20474429774975977, 0.15361262608581772, 0.17148384953597295, -0.06863251022860561, -0.09641546191943107, 0.41637177099455247], 'aa_anchor9': [0.05434809011842612, 2.7801617864888852, -0.35797680422744804, -0.8190052835219146, 0.11857350539615208, 0.0269631609070089, -0.3355168189680158, -0.023599697355983995, -0.20320776641489546, 0.1287692400630729, 0.6980107405386833, -0.21220174658533786, 0.4049254944493561, -0.41469409792973, -0.2473629738604517, -0.09902236803167508, 0.011755231096659546, -0.02067569534130609, 0.8729848450087081, 0.004205353352335628], 'aa_tcr': [0.11158462317236717, 2.773959531831438, -0.0993814377936384, -0.29671810806874266, 0.4950107076279453, -0.00016270415284225237, -0.295647663191398, 0.12410365697561065, -0.6520732243895435, -0.0091045818055, 0.8711861233484659, 0.09271524735876246, -0.2901417838675875, -0.5028664472546183, -0.0702567107737444, -0.05764223762239684, 0.021267736887371047, -0.05961947430965431, 0.8414969403179011, 0.44252305607409204], 'aa_tcr10': [0.14452336323282955, 2.7647428509620076, -0.27303124822977676, -0.30798896302550416, 0.6143544448118448, -0.2331943765628468, -0.04703302854604585, 0.1971390549584613, -0.5847531563232575, 0.18293207436529757, 1.2027446635994674, 0.0816790776425278, -0.5113814737783633, -0.6585624691458074, 0.049693229926436544, -0.005135993202697087, 0.028414528536732764, 0.015028342421871344, 0.7667575271402747, 0.30589199181565396], 'aa_tcr11': [0.042926325241647856, 2.9003697945156524, -0.7895690061266629, -0.7915433129800777, 0.6393630303776776, -0.3833691741191667, -0.2225349870647575, 0.24237121287574004, -0.47042485007288226, 0.38893639039422423, 1.4626645383965529, 0.15788930858279437, -0.5669489563014309, -0.5750469202268622, 0.10258790549005825, -0.10261002927518481, 0.1098189843463917, -0.006438385693241955, 1.1720230083245715, 0.7555479833927148], 'aa_tcr8': [-0.3626440402975324, 2.5690093765719366, 0.025152763677318113, -0.16113505155131636, 0.5716458918736516, -0.2940218366346805, -0.3270023489716065, 0.15256964736345102, -0.7693940827296135, -0.4177094648848714, 0.14329159817426618, 0.4257479627254508, -0.33763705409399103, -0.5355484143287876, -0.11204150782514999, 0.016121861455391517, 0.10389869076731628, -0.34296188767169866, 0.44423758739279595, 0.727599998700061], 'aa_tcr9': [0.25611875334301626, 2.7565035137077114, 0.0315792749221524, -0.30293376920372683, 0.4833436678602476, 0.2763942471895562, -0.39608284166471996, 0.06641716904970618, -0.6421198031734301, -0.08097270631773501, 0.7710438091180425, -0.1006770393040215, -0.053726795463244326, -0.4737200041343188, -0.10586162477797378, -0.09686090704837946, -0.05122900072975867, -0.011995701167633932, 0.9028414672177298, 0.3915866895934963], 'aa_tcr_c': [0.1567402196952874, 2.7264368854118315, -0.18612278745969535, -0.26665195917354545, 0.7469810008938964, 0.05105625357002408, -0.30566025168897903, 0.12538129624249628, -0.9114980709696425, 0.0758398085986891, 0.9547759452427305, -0.05394852113794668, -0.05793638638900056, -0.5162769946096821, -0.2905931903804695, -0.04758468295787388, 0.04809361857382566, 0.004081676160733405, 0.929273848867064, 0.4292549334284228], 'aa_tcr_m': [0.06040864281957914, 2.944928961060338, -0.39027352332045373, -0.5272985419894227, 0.3926480736280371, -0.08845857572090399, -0.3938329496610251, 0.07786352302480992, -0.9161797421571483, 0.02228140375683152, 0.8838121918738211, 0.14101175178008019, -0.7282021360854594, -0.5074887196245412, 0.12285779355736359, -0.16303490498628337, 0.02092417847246919, -0.13251087115944982, 0.5824957217012088, 0.48492599355688126], 'aa_tcr_n': [0.08560626373448521, 2.6587189318351836, 0.16515666836662835, -0.22547288708232927, 0.3320831434654905, 0.04564772848685372, -0.19907053579036305, 0.15704050294825533, -0.1325443306694063, -0.23930737878525132, 0.7037971757516512, 0.25022020515325893, -0.24378592296794555, -0.49363813072452123, 0.10065796644980018, 0.02027761719567467, -0.08179738019077876, -0.11926920006029063, 0.970448592894039, 0.41046372086635596], 'kmer_llr': [0.15700106006952907, 2.6020064931053426, -0.5148701638077364, -0.18437555126677374, 0.45976405018647615, 0.698371827198268, -0.8381063048064856, 0.14787150198409638, -0.8799659287420241, 0.19629552225488744, 0.87478554501486, 0.1514119883061218, -0.3684079724643574, -0.8460551975330297, -0.1334426821849819, -0.011634023525791903, 0.4300882820403604, -0.14396294285916067, 1.119320286310387, 0.2153382829007553, 4.106083889881617, 4.656130226800889, 2.8643707575728348, 2.091180869339352, 2.091180869339352, 3.035642478180204, 2.6789675342414725, 3.7006187817734526, 1.1748901374651979, 2.000209091133626, 2.091180869339352, 2.1601737408263038, 3.007471601213507, 2.6507966572747748, 2.378862941791134, 2.2988202341175974, 4.211444405539444, 2.879638229703623, 1.8034987968875704, 4.083611034029559, 0.20225765165764997, 2.245331549166611, -0.05952369236197441, -0.521208213193745, 0.39490341598443557, -0.24821819677740908, -0.361149610014035, -0.07923245054620942, -0.8194520736966346, -0.09136454533381766, 0.924745984332481, 0.7098005230218911, -0.9473714013975663, -0.6442683059509715, 0.03279273685734907, -0.20513461064109784, -0.2216959096912401, -0.41179078784628675, -0.05725354382743486, -0.20782640230680904, 0.06447562374045912, 2.879638229703623, -1.2299738046492665, -0.6597463561192498, -0.1472259979632824, -0.15531475700364972, -0.738103196576974, -0.362693671200919, -1.1978992142020966, -0.7102434025403159, 0.252901384476405, -0.04549890419540059, -0.43614649633259983, -0.5934078450838598, -0.5963708102145171, -0.5453221804708095, -0.11434251081156166, -0.31961780829492614, -0.1945971053383122, -0.2615540073401643, 0.7316646273008551, 2.6890178700949727, 0.045640879335198115, 0.2620338680396772, 1.2058736829851648, -0.1710161627494058, -0.12802261471564158, 0.7307277394033491, -0.4058471409809927, 0.5904761570417181, 1.01582544283551, 0.38410448734133595, 0.06701852338429148, -0.5103858513604074, 0.0004397724055831276, 0.5003548040998513, 0.2838204596688172, 0.3440535541320662, 1.2203525115419556, 0.9825182448177419, 0.21553991613059598, 2.0111381616658166, -0.22103511671565457, -0.29100727425366024, 1.1293668824411132, 0.1245348218048612, -0.912850207029333, -0.15907840465295298, -1.3992476460507453, -0.12698688184211004, 0.8316382138586063, 0.25216851526075157, -0.31892456736594976, -0.12679134019111515, -0.25715024544289733, -0.5733842198231187, -0.011075350106655613, 0.5441548892957826, 1.161644910715177, 0.38845502248169517, 0.007120381803221498, 2.378862941791134, -1.3298191396189818, -0.5829677800871762, 0.17425825715729193, -0.9651760260310729, 0.07116274813031787, -0.7504007243873803, -0.6902968003176753, -0.6408225727853507, 0.7048865082194613, 0.04597249930175895, -0.9915627812042684, -0.7652893368811311, -0.3067144034590177, -0.11342381529448886, -0.26638635807204647, -0.38908540089214494, -0.029082666860739792, 0.2193786924377612, 0.0015869561954984235, 2.6665450142429137, -0.3067144034590177, -0.1979812033225521, 0.5574547786409481, -0.3039557810199378, 0.10704950746384156, 0.5762084231580111, -1.291088502437284, 0.13511834882002116, 0.7972598283504766, 0.09663877144600352, 0.1752729705878382, -0.3603465583748129, 0.06422951028091362, 0.003957187217461744, -0.04352335101553262, 0.045315623193121546, 1.3492435246099754, 0.7474461226382578, -0.6633893473977501, 2.9666496066932533, -0.754212893273734, -0.7898884958944992, 0.026777205024057338, -0.968379445748611, -0.7155408592698871, -0.6423484170638423, -1.0020237470234585, -0.8258591356389733, 0.5713551155949395, -0.41255327358662175, -1.1006662831409288, -0.4860010565578188, -0.10176104620486548, -0.4345477749689026, -0.3049302816655164, -0.7972047905851829, 0.08584729981323846, 0.041431117466919076, -0.06742887186908675, 2.3424952976202587, -0.20922794423209812, -0.20482525255665074, 0.46090887000241043, -0.33021237472632237, -0.6053085093617812, 0.3292743632609785, -0.8339338370006848, 0.3597360501326827, 0.6830674608248222, -0.3215633342004178, -0.8004527183772341, -0.5564113627257434, -0.07270728772927537, -0.1216082245379786, -0.032235585920354026, 0.21617868970708987, 1.778806184297201, 0.6464981899813074, 0.814887403433791, 3.5952582661156267, 1.481415297718458, 1.197362993317256, 1.5605526182771827, 0.822669543875846, 0.2193786924377621, 0.9040151833297978, -0.133442682184981, 1.3343178743932995, 1.3439664675091318, 0.7141889008817754, 0.580588791541885, 0.14527072028403953, 0.3301930587780513, 0.9639952082271863, 0.9242493366937659, 0.992568580671243, 3.7959289615777774, 0.8102470238772881, 0.12260192047132623, 2.7843280498992975, 0.18750348390731197, -0.6013651452268922, 0.9381333741162265, 0.23379335643994636, -0.5580288317399242, -0.003764858876448507, -0.3937257804486478, 0.16297029738344015, 1.2087916891408783, 0.13778004855488124, -0.3098442964679453, -0.46936319465426735, 0.42767573563497674, -0.08389090335573357, 0.06303262204706694, -0.3188521114734417, 1.120401952181128, 0.7235786412316143, -0.5605509011726335, 1.8680373180251433, -0.5506199462216568, -0.7504007243873803, 0.6027812852823082, -0.4666204948568975, -0.17750267197901248, -0.17825201633779297, -1.5516546462731764, -0.3271110313666634, 0.7048865082194622, -0.677088191755912, -0.3435759967631764, -1.3648931225486374, -0.6024805943107578, -0.1833837113225183, -0.582083215677442, -0.703762498456233, -0.5976379698349703, -0.09314030316095856, -0.532987847782155, 3.631625910286502, -0.19385991400378355, -1.2193621440546716, -0.47856466038567724, -0.07977578117625672, -1.2184493187973127, -1.0370405872607176, -1.5490334127933032, -0.5973824440893614, 1.1046858787919485, -0.3645154023488466, -0.9480365161543771, -1.0720578752234529, -0.8995388623910943, -0.6556150983386999, -0.6789047226823106, -0.24968641022570992, 0.3565798139512468, 0.418658981478675, 0.318469595081992, 2.5611844985850887, -0.21763477340532944, -0.22690841019065822, 0.3757374879841793, 0.03567209387114456, -0.07469317811279019, -0.6744391835843349, -0.5780294984465932, -0.1907849364519576, 1.1025694758855717, 0.21019026638335525, -0.3004250745514536, -0.6628544450046814, 1.6493481170603133, -0.12179206496500594, 0.2834210587648567, -0.15868744148228853, 1.323925716625685, 0.2193786924377612, 0.2575189084572882, 2.208963904995736, -0.36669710806072864, -0.7051619389085353, 0.3723538696023292, 0.03557180302761509, -0.37973953847390796, 0.16180002223524959, -0.9596707034357177, 0.1874710747744155, 1.0413587448406751, -0.2258160023159963, -0.22718809135595475, -0.8381063048064856, -0.7495939474174405, -0.27187604540971044, 0.14095105913952288, -0.05864146910228296, 0.5797233652654565, 0.3388802368518453, 0.2486490747378749, 3.14100299383803, -0.2952857076353954, -0.8223133059776666, 0.26752803633534317, -0.1854899348892225, -0.19882544144783232, 0.18312094441513693, -0.5215591519585328, 0.33552640475341367, 1.0957528169064732, -0.22690841019065822, -0.3579066826436952, -0.5898406593749383, -0.1127551425848532, -0.10809920717896304, 0.23728961900429102, -0.03452239015655678, 1.0795799576608731, 0.461264425561712, 0.008273119683029684, 3.295153673665289, -0.33656736660869857, -0.637347863106398, 0.3757374879841793, -0.22310026341788447, -0.275942744792264, -0.0028594717616359233, -1.1993509444352837, -0.07400435425225016, 1.4190870979772399, -0.232431590846768, -0.44559761115548824, -0.33923759516457785, -0.6837199926588085, -0.35037533957503353, -0.14705892056428826, 0.13778004855488124, 0.5042158127573115, 0.6635012920566075, 0.3162285184276783, 2.4966459774475167, 0.14527072028404042, 1.5521843686066656, 1.0473768171662377, 0.3181135331234506, 0.40478191576912437, 0.7694250293570333, -0.023351992151753542, 1.4991298056507762, 1.6857157612311875, 1.2847050034724043, 0.27880211290856227, 1.1596226653344095, 0.3235189516903585, 0.11401817677993531, 0.9254292778336151, 0.3864327771009277, 2.091180869339353, 1.302723508975082, 0.5322268623372439, 2.9384787297265564, 0.0702355343411254, 0.3864327771009277, 0.8577941798071231, -0.09487040739874164, 0.600107556986285, 1.0278145197329174, -0.14952881993660494, 0.2680557408330779, 1.146719260498501, 0.8464759621737468, 0.06830967914791142, 0.5318673440878205, 0.1728586768028686, 0.3522638814856771, -0.026975184521174533, 0.12506801296652004, 0.9672507726869535, 1.0351281950900386]}, 'log_odds_source': {'aa_anchor': 'anchor', 'aa_anchor10': 'anchor@10', 'aa_anchor11': 'anchor@11', 'aa_anchor8': 'anchor@8', 'aa_anchor9': 'anchor@9', 'aa_tcr': 'tcr', 'aa_tcr10': 'tcr@10', 'aa_tcr11': 'tcr@11', 'aa_tcr8': 'tcr@8', 'aa_tcr9': 'tcr@9', 'aa_tcr_c': 'tcr_c', 'aa_tcr_m': 'tcr_m', 'aa_tcr_n': 'tcr_n', 'kmer_llr': 'pair'}, 'logistic': {'coef': [0.12551420796218662, 0.011746965690129688, -0.22796120873496709, 0.07389043313873397, -0.10648502804956149, 0.0981711722426051, 0.12875607543264733, -0.07270088127999416, 0.22781527923445063, -0.01464490734396946, 0.19703550119972588, 0.05942800972836916, -0.04333446425231443, 0.032323813691251826, 0.06151924942123461, -0.10680728245257694, 0.08415955976711734, -1.0316874083306022, 0.3513886251185817, 0.38937808996351464, 0.37747313651651576, 0.44008083790377955, 0.31284697323611715, 0.3089903633919363, 0.2950373758403601, 0.2913008156208962, 0.22854742491454486, 0.0651838617708407, 0.2907564651418195, 0.893787772998956], 'intercept': -2.6803856922142644, 'tau': 4.0}, 'n': 47140, 'n_immunogenic': 5154, 'paratope': {'A': [0.0468, 0.3668], 'C': [0.1994, 0.5024], 'D': [0.1813, 0.5419], 'E': [0.1014, 0.313], 'F': [0.0754, 0.3616], 'G': [-0.0, 0.2369], 'H': [0.2189, 0.5621], 'I': [0.0486, 0.416], 'K': [0.1234, 0.4516], 'L': [-0.0251, 0.1864], 'M': [0.0345, 0.3921], 'N': [0.1193, 0.3291], 'P': [0.0469, 0.2501], 'Q': [0.0384, 0.3411], 'R': [0.0515, 0.421], 'S': [0.0124, 0.2607], 'T': [0.1307, 0.3778], 'V': [0.0143, 0.2991], 'W': [0.0847, 0.4427], 'Y': [0.0172, 0.2474]}, 'prevalence': 0.10933389902418328, 'seed': 20260817, 'standardizer': {'mean': [3.649061829380585, 8.026884001018251, 9.072910479422996, 5.918703917394879, 4.0730983529274685, -2.2696420880144097, 3.9537856480908196, -0.3019153585065778, 0.5219495120916435, 13.831227195587637, 10.148506788290465, 0.06310424664121059, 0.33805213484655444, 1.8381417055579126, 1.277895630038184, 0.5352962805826635, -0.252232403137967, -0.1718184419660689, -0.12433160035260404, -0.0722116039466836, -0.06933721421291336, -0.0646821815674958, -0.049932359407905566, -0.04023770735202519, -0.08874299682438072, -0.0755211620108445, -0.03413589022035551, -0.07822360968723212, -0.0987321163119205, -0.3785270097660761], 'std': [18.32119223144693, 14.18739320828701, 0.9419106801415313, 13.983141047815161, 10.34894494135274, 12.899403706872747, 9.853255541060484, 1.9410970099623657, 1.8385842630913234, 2.0065068003403557, 2.7392906053643205, 0.02962738503853292, 0.0469905931559011, 1.0282499024909733, 0.5717053374996469, 0.24464656713775987, 0.8864423156372201, 0.6707030040588131, 0.6658127889389893, 0.416676136943885, 0.494762791203886, 0.4544240965858623, 0.41044452275654136, 0.34657991976206737, 0.5802170482251391, 0.45662314333473913, 0.28677994099002807, 0.46936947578760807, 0.5137388926023074, 1.2583479310885775]}}}#
standardizer, linear head, the fitted log-odds tables, and the EM / supervised Gaussian parameters kept for comparison. Class-I artifacts are loaded eagerly because every caller that predates the class split expects them present at import.
- Type:
The frozen class-I models, one per species
- mhcmatch.complement.SPECIES = ('human', 'mouse')#
Species with a fitted table. Both are the
chowell_rebuiltarm of their own host – human 464,161 rows / 14,712 immunogenic, mouse 47,140 / 5,154 – never pooled, because the two hosts have different MHC and different thymic repertoires and a fit across them is fitting a mixture.
- mhcmatch.complement.CLASSES = ('mhc1', 'mhc2')#
The two MHC classes, which are two different constructions and not one with a parameter.
- mhcmatch.complement.BLOCKS = {'aa': ['aa_anchor', 'aa_tcr', 'aa_anchor8', 'aa_tcr8', 'aa_anchor9', 'aa_tcr9', 'aa_anchor10', 'aa_tcr10', 'aa_anchor11', 'aa_tcr11', 'aa_tcr_n', 'aa_tcr_m', 'aa_tcr_c'], 'kmer': ['kmer_llr'], 'motif': ['kd_run_max', 'kd_run_n', 'kd_run_frac'], 'phys': ['pc1', 'pc2', 'length'], 'pot': ['mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr'], 'role': ['pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr']}#
it is the class-I layout and every vendored class-I artifact is written against it.
- Type:
Class-I feature blocks.
BLOCKSkeeps its name and its contents
- mhcmatch.complement.BLOCKS_MHC2 = {'aa': ['aa_anchor', 'aa_tcr', 'aa_nflank', 'aa_core', 'aa_cflank', 'aa_anchorL0', 'aa_tcrL0', 'aa_anchorL1', 'aa_tcrL1', 'aa_anchorL2', 'aa_tcrL2', 'aa_anchorL3', 'aa_tcrL3'], 'kmer': ['kmer_llr'], 'motif': ['kd_run_max', 'kd_run_n', 'kd_run_frac'], 'phys': ['pc1', 'pc2', 'length'], 'pot': ['mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr'], 'role': ['pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr']}#
Class-II feature blocks. The
aablock has the same shape as class I’s – the pooled role pair, a position key, and a length key – and differs only in what the position key is: relative thirds of the TCR face at class I, register zones at class II. Seeencode().
- mhcmatch.complement.FITTED = {'aa_anchor': 'anchor', 'aa_anchor10': 'anchor@10', 'aa_anchor11': 'anchor@11', 'aa_anchor8': 'anchor@8', 'aa_anchor9': 'anchor@9', 'aa_tcr': 'tcr', 'aa_tcr10': 'tcr@10', 'aa_tcr11': 'tcr@11', 'aa_tcr8': 'tcr@8', 'aa_tcr9': 'tcr@9', 'aa_tcr_c': 'tcr_c', 'aa_tcr_m': 'tcr_m', 'aa_tcr_n': 'tcr_n', 'kmer_llr': 'pair'}#
Columns computed from a fitted log-odds table rather than from the peptide alone:
feature -> which count matrix it weights.<matrix>@<bin>is that matrix restricted to the rows of one length bin, so a per-length table costs a mask rather than another count matrix.
- mhcmatch.complement.FITTED_MHC2 = {'aa_anchor': 'anchor', 'aa_anchorL0': 'anchor@0', 'aa_anchorL1': 'anchor@1', 'aa_anchorL2': 'anchor@2', 'aa_anchorL3': 'anchor@3', 'aa_cflank': 'cflank', 'aa_core': 'core', 'aa_nflank': 'nflank', 'aa_tcr': 'tcr', 'aa_tcrL0': 'tcr@0', 'aa_tcrL1': 'tcr@1', 'aa_tcrL2': 'tcr@2', 'aa_tcrL3': 'tcr@3', 'kmer_llr': 'pair'}#
The class-II equivalent.
<matrix>@<bin>here indexesmhc2_length_bin(), notlength_bin()–encode()writes the right one intocounts["bin"]per class.
- mhcmatch.complement.ZONES = {'mhc1': ('tcr_n', 'tcr_m', 'tcr_c'), 'mhc2': ('nflank', 'core', 'cflank')}#
Which zone matrices a class partitions its TCR-facing face into. The pooled
tcrmatrix is their sum in both cases, so the split costs no extra pass andposbayesstays recoverable.
- mhcmatch.complement.MHC2_ZONES = ('nflank', 'core', 'cflank')#
the residues before the 9-mer core, the non-anchor residues inside it, and the residues after it.
This was built as the class-II replacement for
LENGTH_BINS, on the reasoning that a floating core makes total length uninformative. The measurement says join, not replace (bench/results/complementarity_mhc2.md): the zones earn +0.0029 human / +0.0034 mouse AUROC over the pooled pair, total length earns +0.0070 / +0.0159, and carrying both beats either on both hosts and both metrics. The reasoning was right about the core and wrong about the ligand – an 18-mer and a 13-mer with the same core do present the same residues, but a class-II ligand’s length is the length of its flanks, which is its own covariate and not a register question.- Type:
Register-relative zones for class II
- mhcmatch.complement.TCR_THIRDS = ('tcr_n', 'tcr_m', 'tcr_c')#
Relative thirds of the class-I TCR-facing face. The same cell means the same fraction along the peptide at every length, which is what the contact profile already does for its per-position weights.
- mhcmatch.complement.LENGTH_BINS = (8, 9, 10, 11)#
8, 9, 10 and 11+. Closed at both ends, which is a shippability property and not only a variance one – a model with one table per observed length cannot score a 12-mer at all, and real inputs contain those.
- Type:
Length bins for the class-I
aablock
- mhcmatch.complement.MHC2_LEN_EDGES = (14, 16, 19)#
Class-II total-length quartile edges, from the corpus’s own distribution (11-25, median 15).
LENGTH_BINScannot be reused: it clamps to 11 and would put every class-II ligand in one bin.
- mhcmatch.complement.mhc2_length_bin(L)[source]#
Which class-II length quartile (0-3) a ligand of length
Lfalls in. Closed both ends.- Return type:
int
- mhcmatch.complement.length_bin(L)[source]#
Which of
LENGTH_BINSa peptide of lengthLfalls in.- Return type:
int
- mhcmatch.complement.KD_THRESHOLD = -0.8500000000000001#
the median of the Kyte-Doolittle scale itself over the 20 standard residues (-0.85), so it is a property of the scale rather than a constant tuned on any corpus. It admits
ACFGILMSTVand excludes the other ten. Same rule asmhcmatch.immuno._aggregate().Taken from the plain dict with the stdlib rather than from
BASISwithnumpy, because sphinx mocksnumpyat doc-build time and a module-levelfloat(np.median(...))then raises on a Mock – the whole module fails to import and its page renders empty.- Type:
The “hydrophobic” cut for the
motifrun features
- mhcmatch.complement.PHYS_SCALE = 'Rose'#
The chemistry half of the recognition axis, as a scale name.
"Rose"is the shipped choice – the average fractional area a residue buries on folding – selected out of 576 candidates by the BIC change it produced inside the general model, not by standalone AUROC. Any other name here is exploratory: it re-parameterises a published result and should be reported as a comparison, never substituted silently. The shipped aggregate records the same name asphys_scale; this constant is whatburial()actually reads.
- mhcmatch.complement.blocks(cls='mhc1')[source]#
The feature blocks of one class:
BLOCKSorBLOCKS_MHC2.- Parameters:
cls (str)
- Return type:
dict
- mhcmatch.complement.fitted(cls='mhc1')[source]#
The fitted-column map of one class:
FITTEDorFITTED_MHC2.- Parameters:
cls (str)
- Return type:
dict
- mhcmatch.complement.mhc2_anchors(peptide, register=None)[source]#
0-based P1/P4/P6/P9 of a class-II ligand’s 9-mer core, memoised.
Delegates to
mhcmatch.store.anchor_indices(), so the register is the same onedecompose, the logos and the signatures use.registerpins the frame – passmhcmatch.diffusion.AnchorModel.best_register()to annotate with the frame the model actually scored with, rather than the allele-agnostic heuristic.The heuristic register is an argmax over up to seventeen offsets per peptide and is the only non-vectorised step in
encode(), so it is cached: a corpus repeats its peptides across alleles and folds, and the register does not depend on either.- Parameters:
peptide (str)
register (int | None)
- Return type:
tuple
- mhcmatch.complement.table(species='human', cls='mhc1')[source]#
The fitted parameters for one
(species, cls).- Parameters:
species (str)
cls (str)
- Return type:
dict
- mhcmatch.complement.phys_scale(scale='Rose')[source]#
Resolve a residue scale by name to
{residue: value}.Accepts a key of
mhcmatch.data.aa_tables.HYDROPHOBICITY(45 scales, including the shipped"Rose"), a"FAMILY:COMPONENT"key intoDESCRIPTORS(e.g."KIDERA:KF4"), or a ready dict overAA. Raises rather than guessing on an unknown name.- Return type:
dict
- mhcmatch.complement.burial(peptides, cls='mhc1', scale='Rose', registers=None, per_residue=True)[source]#
C_phys: a residue scale averaged over the TCR-facing positions.With the default
scale="Rose"this is the shipped chemistry term – the Rose burial propensity, i.e. the average fraction of a residue’s surface buried on folding, over the face a receptor reads. It has no fitted residue parameters: the basis is imported, which is why it cannot memorise the corpus’s cysteine gradient (correlation with per-peptide cysteine count +0.108, against +0.688 for the full fittedscore()).``per_residue=True``, and the old sum was a length detector. The TCR face is
L - 5residues wide and the Rose scale is strictly positive (0.52 to 0.91), so summing it gives roughly0.75 (L - 5): on 60,000 fit-corpus peptides the summed column correlates with peptide length at Pearson +0.954, against +0.052 for a centred scale like Kidera KF4. A chemistry term that is 91 % length variance is not measuring chemistry, and it made the two scales incomparable – one carrying length, one not. Dividing by the face width is the same correctionmhcmatch.mimicry.corpus_R()makes with its per-window divisor, for the same reason. Passper_residue=Falseto restore the summed form.scale=is for exploration, not for scoring. Passing another basis re-parameterises a result that was selected by BIC inside the general model over 576 candidates, so a number produced with a different scale is a comparison and must be reported as one. The obvious comparisons – Kidera factors, against which the Chowell-family literature is usually written – are reachable as"KIDERA:KF4","KIDERA:KF2"and so on; on the neoantigen corpus they lose to"Rose".>>> burial(["GILGFVFTL"])[0] > 0 True >>> abs(burial(["GILGFVFTL"], scale="KIDERA:KF4")[0]) >= 0 True
- Parameters:
cls (str)
per_residue (bool)
- Return type:
list[float]
- mhcmatch.complement.encode(peptides, cls='mhc1', registers=None, positions='mask', paratope='loop')[source]#
(features, counts)for an iterable of peptides – the vectorised feature builder.Everything additive is a matrix product against a residue vector, so the peptide-side feature set is two (n, 20) count matrices (one per role) times seven property vectors. Run statistics need position order and use a loop over positions (at most 25 iterations), never over peptides. Peptides are grouped by residue layout so each group is one batch of array operations.
countscarriesanchorandtcr(n, 20) – the matrices theaalog-odds tables weight – plus that class’sZONESmatrices and the adjacent TCR-facing residue pairs as a sparse pair list,pair_code(which of 400 pairs) andpair_row(which peptide). A 9-mer has at most 3 such pairs, so a dense (n, 400) matrix would be 99% zeros and would cost 1.5 GB of temporaries per pass on a 500k-peptide corpus;apply_log_odds()sums the sparse form instead.The two classes differ in one place and it is deliberate. Class I is anchored at fixed peptide positions (
ANCHORS), so its geometry is a function of the peptide’s length: theaablock carries one table perLENGTH_BINSbin, and the TCR face is split into relative thirds. A class-II ligand is anchored by a 9-mer core that floats inside an 11–25-mer, so its total length is the length of its flanking regions and carries no register information – an 18-mer and a 13-mer with the same core present the same residues to the receptor. Binning a class-II table on total length would therefore split it on a variable that says nothing about the object being modelled. The class-II construction bins on the register instead (MHC2_ZONES): before the core, inside it, after it. That both classes reach the sameaa_anchor/aa_tcrpooled pair on top of it is what keepsmhcmatch.posbayes.llr()a strict special case of either.registers(class II only) pins each peptide’s core to an explicit frame, element for element withpeptides;Noneuses the allele-agnostic heuristic register.positionsselects which peptide positions the TCR-facing chemistry is read over – the binary anchor mask ("mask", the shipped default) or the positional contact profile ("profile").paratopeselects which marginalisation of the TCRen potential thepotblock projects onto –PARATOPE("loop", the shipped default) orPARATOPE_CONTACT("contact"). The two compose, and both are explained in full onscore(), which is where the measured consequences of changing either are recorded.The hydropathy stretch, exactly. A residue enters the
motifblock when all three hold: it is TCR-facing (an anchor is buried in the groove and is not part of any stretch the receptor reads), it is one of the 20 standard residues, and its Kyte-Doolittle value exceedsKD_THRESHOLD.kd_run_maxis then the longest run of consecutive such positions,kd_run_nthe number of runs (counted as rising edges, soIIDIis two runs andIIDDis one), andkd_run_fracthe above-threshold count divided by the number of TCR-facing positions. An anchor breaks a run rather than bridging it: two stretches either side of a buried residue are two stretches.Non-standard residues (
Xmasks,B/J/O/U/Z) contribute to no count matrix and to no property sum, but they still count towardlength– and in themotifblock they behave exactly like a below-threshold residue rather than like a gap. A mask breaks a run (AAAIIXIAAgiveskd_run_max = 2, the same asAAAIIDIAA, not the 3 it would give if the mask were transparent) and it sits inkd_run_frac’s denominator while never entering its numerator. That is the conservative reading of an unknown residue – it is not evidence of a continued hydrophobic stretch – and it is stated here because the opposite was once claimed.- Parameters:
cls (str)
positions (str)
paratope (str)
- Return type:
tuple[dict[str, numpy.ndarray], dict[str, numpy.ndarray]]
- mhcmatch.complement.apply_log_odds(counts, source, weights)[source]#
A fitted log-odds table applied to one count structure – one value per peptide.
anchor/tcr/ the threeTCR_THIRDSare dense (n, 20) and this is a matrix product.pairis the sparse pair list, and summingweights[code]per row withnumpy.bincount()costs O(pairs) rather than O(n x 400) – the difference between seconds and a gigabyte of temporaries on a corpus."<matrix>@<bin>"is that matrix restricted to oneLENGTH_BINSbin. The restriction is applied to the result, not by materialising a masked copy: a per-length table is then a length-n mask rather than another (n, 20) matrix, which is what makes eight of them affordable.- Parameters:
counts (dict)
source (str)
- Return type:
numpy.ndarray
- mhcmatch.complement.design(peptides, species='human', cls='mhc1', registers=None, mask_cys=False, positions='mask', paratope='loop')[source]#
The (n, k) design matrix in
feature_names()order, before standardization.mask_cys,positionsandparatopeare forwarded unchanged; seescore()for what each one does and what it costs.- Parameters:
species (str)
cls (str)
mask_cys (bool)
positions (str)
paratope (str)
- Return type:
numpy.ndarray
- mhcmatch.complement.feature_names(species='human', cls='mhc1')[source]#
Column order of the design matrix, matching the vendored coefficients.
- Parameters:
species (str)
cls (str)
- Return type:
list[str]
- mhcmatch.complement.features(peptide, species='human', cls='mhc1')[source]#
The feature vector for one peptide, as a name -> value dict. For inspection; use
score()(which takes a list) for anything at scale.- Parameters:
peptide (str)
species (str)
cls (str)
- Return type:
dict[str, float]
- mhcmatch.complement.score(peptides, species='human', cls='mhc1', registers=None, blocks=None, mask_cys=False, positions='mask', paratope='loop')[source]#
Log-odds of immunogenic vs not, one per peptide. Carries no prior.
Larger is more immunogenic. The training corpus’s own base rate is divided out, so this is directly comparable across settings and composes with any prevalence via
posterior().Accepts a single string as well, and still returns an array of length 1 – so a caller never has to branch on whether they passed one peptide or a million.
blocksrestricts the score to a subset ofBLOCKS. The head is linear over a fixed block partition, so this is an exact partial sum of the shipped score, not a refit: the parts add back to the whole, up to the constant offset that belongs to no block. Withoff = logistic.intercept - log(prev / (1 - prev))fromtable(),- score(p, blocks=(“phys”, “role”, “pot”, “motif”)) + score(p, blocks=(“aa”, “kmer”)) + off
== score(p)
exactly. The offset is dropped from a block score deliberately – an intercept is not attributable to any block, and a constant cannot change a ranking. Whole blocks only:
aa_tcris fitted at-0.8796while every per-lengthaa_tcr{8,9,10,11}and everyaa_tcr_{n,m,c}is positive, so the pooled column is collinear with its own decompositions and the fit pushed it negative to compensate. Dropping columns within a block without refitting inverts the dominant term; dropping a whole block cannot.mask_cysscores with cysteine zeroed out of the fitted log-odds tables, asmhcmatch.posbayesdoes by construction. See_mask_cys()for the size of the artifact and why the default is off.positions="profile"reads the TCR-facing chemistry over the positional contact profile (mhcmatch.immuno.contact_profile(): per-position TCR-peptide contact frequency from 8,062 contacts over 370 crystals, keyed by class and length) instead of the binary anchor mask. It touches therole,potandmotifblocks only –aaandkmerare count matrices keyed on a role and a weighted count is a different object, so they come back bit-identical. The anchor face is untouched for the same reason the TCR face is re-weighted: the profile measures contact with a receptor and carries no information about burial in the groove.paratope="contact"reads thepotblock’s TCRen projection offPARATOPE_CONTACT– the potential marginalised over the receptor residues that actually contact peptide – instead ofPARATOPE’s flat whole-loop composition. Independent of ``positions``: one changes which peptide positions are read, the other which receptor residues the potential was averaged over, and they compose.The coefficients are unchanged under either, so both are the shipped head reading a different encoding, not a refit. Whether either should be refit is the question a benchmark answers; the defaults stay
"mask"and"loop".- Parameters:
species (str)
cls (str)
mask_cys (bool)
positions (str)
paratope (str)
- Return type:
numpy.ndarray
- mhcmatch.complement.kidera_design(peptides, anchors=None, roles=None)[source]#
All ten Kidera factors, each summed over the anchors, the TCR face and the whole peptide.
The fitted model uses one of the ten –
kf4_anchorandkf4_tcr, the hydropathy axis – because that is the factor the role split was established on, not because the other nine were compared and rejected. This builds the full set so they can be. Label-free: it reads only the vendored Kidera table andANCHORS, so it needs no fitted artifact and no per-fold refit.>>> kidera_design(["GILGFVFTL"]).shape (1, 30)
- Return type:
numpy.ndarray
- mhcmatch.complement.kidera_names()[source]#
Column names of
kidera_design(), in order.- Return type:
list[str]
- mhcmatch.complement.posterior(peptides, prior, species='human', cls='mhc1')[source]#
P(immunogenic | peptide)at an explicitprior. Exact, becausescore()has none.priorhas no default on purpose. The training corpus runs at ~3.2% positives, a viral proteome scan nearer 3.0e-3 and the NCI screen 4.2e-4 – a default would silently pick one setting’s base rate for every caller and overstate the rest by up to 75x.- Parameters:
prior (float)
species (str)
cls (str)
- Return type:
numpy.ndarray
mhcmatch.recognition module#
The head dispatcher over the recognition axis: complement (the six-block model above, and the
default), plus posbayes, physchem_glm and esm64_glm, each fitted alone so their BIC is
comparable. See Complementarity — the I and C of EPIC.
Recognition score for MHC epitopes: how immunogenic a peptide looks, given how it is presented.
Four heads. default_head() – what score() uses when none is named – is complement,
the six-block, 30-feature model of mhcmatch.complement, which is served by its own module and
has no vendored table here. The other three are each fitted alone so that their fit criteria are
comparable and each score is readable on its own terms; lowest_bic_head() reports the
parsimony winner among those three, currently posbayes for both species. The two answer
different questions – BIC asks which head buys its parameters on one training arm, the default asks
which recognition term to score with – so they do not have to agree, and they do not.
complementSix feature blocks over roles, contact potentials, hydropathy runs, residue identity and adjacent residue pairs; 30 features, linear head.
mhcmatch.complement, and the source of the shippedC_physterm. The default.posbayesNaive Bayes over amino-acid identity conditioned on face – MHC-facing or TCR-facing – scored as a summed log-likelihood ratio. Two 20-cell tables, three parameters. Conditioning on the face rather than on absolute position is what keeps it well defined when peptide length varies, and it is why the two tables can disagree in sign without anything being told to flip. Pure numpy, no optional dependency, and the whole model prints in forty numbers.
physchem_glmRaw sums of the Kidera factors over each face. Length is not a separate feature: summing a constant-1 factor over a face gives that face’s size, so it enters as \(KF_0\) and the two face sizes add to the peptide length. 22 features, all interpretable.
esm64_glm64 principal components of a whole-peptide ESM2 pool. The most accurate head on mouse and the least explainable; needs
pip install 'mhcmatch[esm]'.
All three take what the caller knows, at whatever resolution they know it: peptide with both masks given, peptide with an MHC (masks from the allele’s layout), or peptide alone (the class-I default).
Coefficients come from chowell_iedb_full_matched – the rebuilt Chowell corpus with negatives
resampled so the allele group carries no signal about the label, so no coefficient can be paid for
recognising which allele happened to be typed. Selection, transfer and the full metric matrix are in
bench/results/shipped_models.md.
- mhcmatch.recognition.HEADS = ('complement', 'posbayes', 'physchem_glm', 'esm64_glm')#
complementis the six-block, 30-feature model ofmhcmatch.complementand is the default. The other three are fitted alone on one arm so their BIC is comparable to each other;posbayeswins that comparison at three parameters. That is a statement about parsimony on one corpus, not about which term to score with: in the integrated neoantigen fit it is the six-block form that carries the recognition signal, and a 3-parameter head is a different claim that must not be substituted for it silently. BIC across the three is still reported bytable().
- mhcmatch.recognition.FITTED_HEADS = ('posbayes', 'physchem_glm', 'esm64_glm')#
The three heads with their own vendored artifacts.
complementis served by its own module.
- mhcmatch.recognition.table(head=None, species='human')[source]#
Fitted parameters for one head.
head=Noneresolves tolowest_bic_head(), not todefault_head(): this returns a vendored table, and the default head iscomplement, which has none – it ismhcmatch.complement, whose parameters come fromcomplement.parameters(species).- Parameters:
head (str)
species (str)
- Return type:
dict
- mhcmatch.recognition.default_head(species='human')[source]#
What
score()uses when no head is named: the six-blockcomplementmodel.lowest_bic_head()still reports the parsimony winner among the three separately fitted heads. They answer different questions – BIC asks which head buys its parameters on one training arm, and the default asks which recognition term to score with.- Parameters:
species (str)
- Return type:
str
- mhcmatch.recognition.lowest_bic_head(species='human')[source]#
The lowest-BIC head among the three with their own fitted tables.
- Parameters:
species (str)
- Return type:
str
- mhcmatch.recognition.feature_names(head=None, species='human')[source]#
Column names of
design(), in order.head=Nonefollowstable().- Parameters:
head (str)
species (str)
- Return type:
list[str]
- mhcmatch.recognition.roles_for(peptides, mhc=None, anchors=None, store=None, cls='mhc1')[source]#
Per-peptide boolean masks,
Truewhere the residue faces the MHC.Resolution order is explicit
anchors, then the store’s layout viamhcmatch.store.Store.decompose(), then the class-I default. Everything downstream reads this, so it is the one thing worth getting right.The two non-explicit routes report different anchor sets, deliberately. The default is the five-pocket set the shipped
posbayestables were fitted on (P1–P3, PΩ-1, PΩ). The store route is the store’s own layout: class-I P2/PΩ, or – the reason to take this route – class-II’s P1/P4/P6/P9 within the floating 9-mer core, which no length-only default can give. Passcls="mhc2"with astorefor a class-II ligand.mhcis forwarded todecomposeas itsallele, which accepts it for forward-compat; the layout it returns is the class default and is not allele-conditioned today.- Parameters:
cls (str)
- mhcmatch.recognition.design(peptides, species='human', head=None, *, mhc=None, anchors=None, store=None, cls='mhc1', roles=None)[source]#
The design matrix for one head, in
feature_names()order, before standardization.head=Noneresolves the same waytable()does – tolowest_bic_head()– because only the three separately fitted heads have a design matrix here. Thecomplementdefault is encoded bymhcmatch.complement.design().- Parameters:
species (str)
head (str)
cls (str)
- Return type:
numpy.ndarray
- mhcmatch.recognition.score(peptides, species='human', head=None, **kw)[source]#
Recognition log-odds. Higher is more likely to be immunogenic.
Not a probability – pass it through
posterior()with a prior for the set being scored.- Parameters:
species (str)
head (str)
- Return type:
numpy.ndarray
- mhcmatch.recognition.posterior(peptides, prior, species='human', head=None, **kw)[source]#
sigmoid(score + logit(prior)).prioris the immunogenic rate expected of the set.- Parameters:
prior (float)
species (str)
head (str)
- mhcmatch.recognition.embed(peptides, batch=256)[source]#
Whole-peptide mean-pooled ESM2 embeddings,
(n, 1280).Batched within length groups: slicing batches out of a length-sorted list and keeping only what matches the batch’s first length silently leaves the rest unembedded.
- Parameters:
batch (int)
- Return type:
numpy.ndarray
- mhcmatch.recognition.mhc2_core(peptides, register_start=None)[source]#
The register-anchored 9-mer core of each class-II peptide, and where it starts.
- mhcmatch.recognition.score_mhc2(peptides, species='human', head=None, *, register_start=None, warn=True)[source]#
Class-II score on an MHC-I-fitted model. Read this before using the number.
There is no fitted class-II recognition model here, and this is not one: no class-II immunogenicity corpus with a usable negative set exists to fit against. This applies the class-I coefficients to the class-II binding core, with the groove-facing positions redefined as P1/P4/P6/P9 of the register-anchored 9-mer.
Scoring the core rather than the whole peptide keeps every feature in the range the model was fitted on – nine residues, face sizes of four and five. Scoring a 15-mer directly would put the face sizes far outside it. But the coefficients were fitted where the groove-facing residues are the termini and the TCR-facing ones a contiguous middle, and in class II the two faces interleave, so a coefficient learned on one geometry is being read on another. The register is also a heuristic unless supplied, and an error in the frame moves every residue between faces.
Rank class-II peptides against each other with it. Do not compare the values with class-I scores, do not read them as calibrated, and say which model produced any number you report.
nanwhere no 9-mer core can be assigned.- Parameters:
species (str)
head (str)
warn (bool)
- mhcmatch.recognition.MHC2_ANCHORS = (0, 3, 5, 8)#
0-based P1/P4/P6/P9 within the register-anchored 9-mer core.
mhcmatch.immuno module#
Physicochemical featurization of epitopes for immunogenicity prediction.
The signal this targets is the Chowell/Calis one: immunogenic epitopes differ from presented-but- non-immunogenic ones in the physicochemistry of their TCR-facing residues, not their anchors. Two things follow, and both are parametrized here rather than hardcoded:
Which positions are TCR-facing is not agreed even within this toolchain –
mhcmatch.store.anchor_indices() masks class-I P2+PΩ, while
seqtree.layout.presentation_features() and mhcmatch.diffusion.MHC1_ANCHORS mask
P1-P3+PΩ-1,PΩ. ANCHOR_SCHEMES keeps all of them selectable so the choice is an ablation
axis with a reported number, not a silent constant. "contact" is the third option: a continuous
per-position weight from observed TCR-peptide contact frequency, which needs no anchor call at all.
How residues are aggregated matters as much as which scale is used. Summed/averaged descriptors
are the established positive result (Chowell 2015 on hydrophobicity; Pogorelyy 2018 associates
epitope length and summed Kidera factors 6 and 10 with precursor frequency), so sum/mean are
always emitted. But a contiguous hydrophobic stretch is a different object from the same residues
scattered across the peptide, and no sum can express it – hence run_max/run_n/run_frac.
length is emitted as a feature, deliberately. Ligand length distribution is allele-specific and
is part of what distinguishes a real ligand set; it is signal here, not a nuisance to regress out.
All scales are plain dict[str, float] (see mhcmatch.data.aa_tables), so any AAindex-style
table can be passed through the same call.
- mhcmatch.immuno.ANCHOR_SCHEMES: dict[str, tuple] = {'full': (), 'p2_pomega': (2, -1), 'pockets': (1, 2, 3, -2, -1)}#
Class-I anchor position sets, 1-based (negatives count from the C-terminus), matching the three definitions that coexist in the toolchain.
fullmasks nothing (whole-peptide baseline);contactis continuous and ignores this table entirely.
- mhcmatch.immuno.DEFAULT_SCALES = ('KIDERA', 'VHSE', 'MJ', 'KyteDoolittle')#
VHSE, MJ potential, Kidera – plus Kyte-Doolittle, the axis Chowell used, so the reproduction gate runs on the same scale the paper did.
- Type:
The scales named in the original brief
- mhcmatch.immuno.scales(names=('KIDERA', 'VHSE', 'MJ', 'KyteDoolittle'))[source]#
Resolve scale names to residue->value tables.
A name may be a descriptor family (
"KIDERA"-> its 10 components, each its own scale), a hydrophobicity scale ("KyteDoolittle"), or"MJ"for the Miyazawa-Jernigan partition energy. Unknown names raise rather than being silently skipped.- Return type:
dict[str, dict[str, float]]
- mhcmatch.immuno.position_weights(peptide, cls='mhc1', scheme='p2_pomega', register_start=None, contact_profile=None)[source]#
Per-position weights: 0 at anchors, 1 at TCR-facing positions.
schemeis a key ofANCHOR_SCHEMES, or"contact"– which requirescontact_profile, a callable(length) -> list[float]of continuous per-position weights (observed TCR-peptide contact frequency, normalized per length).Class II ignores
schemeand always masks the register-anchored core P1/P4/P6/P9, because that definition is agreed across the toolchain. Passregister_startfrommhcmatch.diffusion.AnchorModel.best_register()so the annotated frame matches the scored one;Noneuses the allele-agnostic heuristic register.- Parameters:
peptide (str)
cls (str)
scheme (str)
register_start (int | None)
- Return type:
list[float]
- mhcmatch.immuno.features(peptide, cls='mhc1', scheme='p2_pomega', scale_names=('KIDERA', 'VHSE', 'MJ', 'KyteDoolittle'), register_start=None, contact_profile=None)[source]#
Physicochemical feature vector for one peptide.
Returns
lengthplus, per scale, the seven statistics in_STATSkeyed"{scale}_{stat}". Non-standard residues (Xmasks,B/J/O/U/Z) are dropped from the aggregation rather than scored as zero – a zero is a real value on a centred scale like Kidera and would be a silent bias.- Parameters:
peptide (str)
cls (str)
scheme (str)
register_start (int | None)
- Return type:
dict[str, float]
- mhcmatch.immuno.feature_names(scale_names=('KIDERA', 'VHSE', 'MJ', 'KyteDoolittle'))[source]#
Column order matching
features(), for building a matrix without a dict round-trip.- Return type:
list[str]
- mhcmatch.immuno.contact_profile(cls='mhc1')[source]#
Structure-derived per-position weights:
(length) -> list[float], forscheme="contact".Backed by
mhcmatch.data.contact_profile– 8,062 TCR<->peptide residue contacts over 370 crystal structures. Two derived steps turn a contact frequency into a weight:Zeroing. A position contacted less than half as often as a uniform footprint would predict (
frac < 1/(2L)) is treated as not TCR-facing. The threshold is derived from the profile’s own uniform expectation, not tuned. On class-I 9-mers this zeroes exactly P1, P2, P3 and PΩ – the anchor set the contact data supports, recovered without being told anchors exist.Rescaling. Surviving positions are scaled to mean 1, so the weighted statistics in
features()are on the same scale as the binary schemes (where a kept position weighs 1) and the two are directly comparable.Lengths with fewer than
MIN_STRUCTURESobserved structures fall back to the class’s pooled relative-position profile, interpolated onto the requested length – the TCR footprint scales with the peptide, so relative position transfers where absolute position does not.- Parameters:
cls (str)
mhcmatch.posbayes module#
Position-role naive Bayes over amino-acid identity: anchor and TCR-facing residues get separate conditional distributions, because for several amino acids their contributions carry opposite signs and pooling averages that away. Emits a log-likelihood ratio, so the caller supplies the prior.
Grouped 5-fold CV 0.712 human / 0.758 mouse; size-matched transfer 0.731 human→mouse, 0.692 mouse→human. Cysteine is masked — see the module warning.
Position-role naive Bayes: amino-acid evidence for immunogenicity, split by anchor / TCR-facing.
A whole-peptide physicochemical predictor scored a peptide from pooled physicochemical descriptors and did not distinguish where a residue sits. But anchor and TCR-facing positions are different channels – an anchor residue is buried in the MHC groove and a TCR-facing one is contacted by the receptor – and the benchmark finds their contributions carry opposite signs for several amino acids. Pooling averages that away.
This model keeps them apart in the simplest form that can: for each role separately, the conditional amino-acid distribution given the class, Laplace-smoothed, scored as a summed log-likelihood ratio.
score(p) = sum_i log [ P(p_i | role(i), immunogenic) / P(p_i | role(i), non-immunogenic) ]
Each peptide reduces to two 20-vectors of counts, which is what lets 8-11mers pool in one model.
It emits a log-likelihood ratio, not a probability. That is deliberate: the LLR carries no prior, so a caller can supply whatever base rate their setting actually has –
logit P(immunogenic) = llr(peptide) + log(prior / (1 - prior))
– and posterior() does exactly that. The distinction matters because the training corpus runs
at ~3.2% positives while a viral proteome scan runs at ~3.0e-3 (counted from distinct 9-mers against
known epitopes) and the NCI screen at 4.8e-4. Reading a corpus-prevalence probability as an
operational one overstates it by 11-66x.
Measured performance, peptide-grouped 5-fold cross-validation (no peptide in both train and test), on the IEDB positive-T-cell-assay vs self-eluted-ligand corpus:
metric |
human |
mouse |
|---|---|---|
rows |
464,310 |
47,203 |
immunogenic |
14,712 |
5,154 |
AUROC (this model) |
0.712 |
0.758 |
Size-matched cross-species transfer, mean over 10 matched subsamples:
human -> mouse: 0.731 (sd 0.003)
mouse -> human: 0.692 (sd 0.000)
is not like-for-like. It is quoted because an in-sample baseline that still loses is the conservative direction, not because it is a fair contest. The module is gone; the measurement is not, and neither is this row.
Warning
Cysteine is masked, and the reason is an assay artefact worth knowing about. The negatives in this corpus are mass-spectrometry-eluted ligands, and cysteine is systematically under-detected in immunopeptidomics unless alkylated. Measured on the training corpus: Cys-containing peptides are 11.59% of the T-cell-assayed positives, 1.73% of the IEDB eluted negatives and 0.17% of the thymus MS negatives – a 6.5x depletion driven by platform, not biology. Fitted freely, Cys took the single largest coefficient in the model (+1.84 anchor / +2.05 TCR-facing). It is therefore zeroed here. The cost is small and measured: grouped-CV AUROC 0.712 -> 0.690.
Any model trained on MS-eluted negatives against assayed positives inherits this, including any model whose training corpus was built the same way.
- mhcmatch.posbayes.ANCHORS = (0, 1, 2, -2, -1)#
Anchor positions, signed, matching
mhcmatch.immuno.ANCHOR_SCHEMES"pockets".
- mhcmatch.posbayes.HUMAN = {'anchor': (-0.17202, 0.0, -0.338211, -0.530797, -0.105505, 0.141961, -0.443731, 0.08185, -0.122986, 0.253121, 0.554954, 0.130571, -0.035569, -0.204233, -0.306762, -0.088272, 0.060235, 0.157231, 0.136212, 0.063724), 'n': 464310, 'n_immunogenic': 14712, 'prevalence': 0.031685727208115265, 'tcrface': (0.126897, 0.0, -0.127745, -0.396513, 0.312013, 0.054715, -0.289776, -0.086445, -0.38218, 0.076293, 0.555381, 0.140134, -0.161836, -0.391744, 0.13859, -0.093884, 0.135715, -0.059949, 0.806245, 0.356368)}#
Fitted on 464,310 rows / 14,712 immunogenic (human host, MHC-I, 8-11mers).
- mhcmatch.posbayes.MOUSE = {'anchor': (-0.009639, 0.0, -0.711281, -1.067674, 0.107452, -0.005274, -0.350006, 0.130257, -0.255666, 0.159464, 0.843716, -0.215143, 0.305866, -0.425195, -0.243769, -0.014672, -0.024961, -0.03003, 0.604519, 0.234526), 'n': 47203, 'n_immunogenic': 5154, 'prevalence': 0.10918797534055039, 'tcrface': (0.11129, 0.0, -0.099216, -0.295732, 0.495405, 6.9e-05, -0.295119, 0.1245, -0.652293, -0.009727, 0.870243, 0.093684, -0.29161, -0.502684, -0.07, -0.05781, 0.021781, -0.060342, 0.841119, 0.443017)}#
Fitted on 47,203 rows / 5,154 immunogenic (mouse host, MHC-I, 8-11mers).
- mhcmatch.posbayes.llr(peptide, species='human')[source]#
Log-likelihood ratio of immunogenic vs non-immunogenic. Carries no prior.
Larger is more immunogenic. Non-standard residues are skipped; cysteine contributes 0 by construction (see the module warning).
- Parameters:
peptide (str)
species (str)
- Return type:
float
- mhcmatch.posbayes.posterior(peptide, prior, species='human')[source]#
P(immunogenic | peptide)at an explicitprior. Exact, becausellr()has none.prioris not optional and has no default on purpose. The training corpus runs at ~3.2% positives; a viral proteome scan is nearer 3.0e-3 and the NCI screen 4.8e-4, so a default would silently pick one setting’s base rate for every caller.- Parameters:
peptide (str)
prior (float)
species (str)
- Return type:
float
mhcmatch.luksza module#
The Łuksza recognition term \(R = Z/(1+Z)\) – a soft partition function over near-matches,
replacing a hard distance cut, so how many near-matches a candidate has and how near they
are both enter. viral_R is not a term of the shipped aggregate — the sum
lived only in the benchmark repo, which made mhcmatch.rank.aggregate_score() callable with a
feature no installed user could supply. EPIC does not score with it; the quantity is still the
published recognition term and is
still computed here. k and a0 are read from the shipped artifact when one carries them, so
a refit needs no code change.
The Łuksza recognition term \(R = Z/(1+Z)\), computable without the benchmark repo.
The shipped EPIC aggregate does not score with this term (see SHAPE). It ships
because the quantity is the published recognition term, and a refit or a head-to-head against the
Łuksza model needs it computable from an installed library rather than only inside the benchmark
repository.
The model. A soft partition function over near-matches, replacing a hard distance cut
(Balachandran, Łuksza et al., Nature 551:512–516, 2017, doi:10.1038/nature24462; Łuksza et al.,
Nature 606:389–395, 2022, doi:10.1038/s41586-022-04735-9):
a_e is the alignment score against reference epitope e, k an inverse temperature and
a0 a horizontal offset. R saturates at 1 when some reference is close and decays smoothly
otherwise, so how many near-matches there are and how near they are both enter — which a
distance cut discards.
The alignment score is identities, not Smith–Waterman, matching the fit: for a peptide of
length L at d substitutions a_e = L - d, so Z collapses to a sum over distances
weighted by how many reference windows sit at each. That is cruder than a gapped alignment and
monotone in the same direction, and the masked-Hamming search makes it exact rather than
approximate.
``k`` and ``a0`` are read from the shipped artifact, not hardcoded. They were fitted by profile likelihood on our own corpus rather than carried over from the papers, which estimated them on a different reference set, a different alignment and a different cohort.
Warning
``R`` is only comparable against the reference set it was fitted with. The shipped
standardizer has sigma = 3.8e-8 because the sum saturates near zero for almost every peptide,
so an R computed against a different viral ligandome is not on the same axis as the fitted
coefficient. viral_r() therefore defaults to the same reference and radius the fit used.
Supply a different one only if you are refitting.
Where the time goes, measured rather than assumed. End to end against the 57,331-peptide viral
ligandome at radius 4, on 20,000 peptides: 57,000 peptides/s, of which the seqtree neighbour
search is 98.6 %, counts_by_distance() 1.3 % and r_term() 0.1 %. The two functions
here are already an order of magnitude faster than the search that feeds them, so vectorising them
further optimises 1.4 % of the run – do not bother. If this path ever needs to be faster, the
search is the thing to attack.
- mhcmatch.luksza.r_term(counts, lengths, k=None, a0=None)[source]#
R = Z/(1+Z)from per-distance reference counts.counts[i, d]is how many reference windows sit at exactlydsubstitutions from peptidei;lengths[i]is its length.k/a0default to the shipped fitted shape.The exponent is clipped to
[-60, 60]:exp(-k(a0 - a))overflows for smalla0and largekon a peptide with thousands of near-matches, and at 60 the term is already far past saturatingR.>>> import numpy as np >>> r = r_term(np.array([[0, 0], [5, 0]]), np.array([9, 9]), k=1.0, a0=9.0) >>> bool(r[1] > r[0]) True
- Parameters:
k (float | None)
a0 (float | None)
- Return type:
numpy.ndarray
- mhcmatch.luksza.counts_by_distance(peptides, hits, category, max_subs=4)[source]#
(counts, lengths)from amhcmatch.mimics.neighbours()result.hitsis{peptide: {category: [(n_subs, ref_peptide), ...]}}. Distances abovemax_subsare dropped rather than folded into the last bin, becauseZweights them byexpand a fold-in would silently inflate the tail.- Parameters:
hits (dict)
category (str)
max_subs (int)
- mhcmatch.luksza.viral_r(peptides, ref_sets=None, *, max_subs=4, k=None, a0=None, threads=1, category='viral')[source]#
viral_Rend to end – the published recognition term, computed standalone.Searches
peptidesagainst the viral ligandome withmhcmatch.mimics.neighbours()and turns the per-distance counts intoR. Withref_sets=Noneit loads the same default reference the coefficient was fitted against. The shipped aggregate does not consume this column –mhcmatch.rank.aggregate_score()scoresEPIC, which has noviral_Rfeature.>>> from mhcmatch import luksza >>> luksza.viral_r(["NLVPMVATV", "GILGFVFTL"]) array([...])
- Parameters:
max_subs (int)
k (float | None)
a0 (float | None)
threads (int)
category (str)
- Return type:
numpy.ndarray
mhcmatch.mimicry module#
Mimicry as a signed, per-component immune-response risk. Three references — viral (priming),
self (tolerance, and the autoimmunity read-out) and thymus (negative selection) — each split
into an anchor and a TCR-facing channel, because a whole-peptide distance averages two
different measurements. Scores are log-odds; mhcmatch.mimicry.probability() is a separate
step that requires a named corpus, because the screens behind any calibration run from 0.048 % to
46.8 % positive and an unqualified probability is mostly a statement about which intercept was used.
The tested-neoantigen database is exposed as mhcmatch.mimicry.annotate() — prior evidence,
never a fitted term. Every labelled screen we hold sits inside that database, so a coefficient on
it would be memorisation; held out of the fit, fuzzy matching at two substitutions still recovers
0.08–0.34 of a fresh screen’s positives against 0.00–0.26 for exact lookup, which is what makes it
worth reporting. Command line: mhcmatch neoag.
Mimicry as a signed, per-component immune-response risk.
Three reference sets, each answering a different question, and never summed into one “similarity”
number (mhcmatch.mimics.KINDS makes the same point about the raw scan):
virala foreign presented ligandome. A hit says a pre-existing anti-pathogen repertoire maycross-react, which raises expected immunogenicity.
thymusthe thymic self-immunopeptidome. A hit says reactive precursors met the peptide duringnegative selection.
selfthe host proteome. The same tolerance argument without the presentation guarantee, andsimultaneously the autoimmunity read-out: whatever cross-reactive clones survived selection are the ones that would attack the tissue displaying the mimic.
Every component is split into two channels, because a whole-peptide distance averages two
different measurements. The anchor channel counts substitutions only over
mhcmatch.complement.ANCHORS; the tcr channel counts them over the complement. The two
partition the peptide, so no position is weighted twice.
There are two conditionings with two different sign patterns, and they must not be confused. The shipped coefficients are the first one.
Standalone – this artifact, bench/results/mimicry_model.md: the six columns plus screen
indicators and nothing else – the sign follows the reference, the way the design predicts.
viral is positive on both channels (+0.60 anchor, +0.44 tcr), self is negative on both
(-0.30, -0.46): priming and tolerance respectively. thymus is positive on the anchor channel
(+0.37) and unresolved on the TCR channel (+0.08, |z| = 1.1).
Residual to a model that already contains a whole-peptide physicochemical term and a foreignness term –
bench/results/mimicry_residual.md – a different pattern appears: across all four references
tried, anchor-restricted similarity is positive and TCR-face-restricted similarity is negative, with
whole-peptide similarity between them and near zero. That is a statement about what mimicry adds to
those terms, not about mimicry on its own, and quoting the second pattern as though it were the
first is a mistake this paragraph exists to prevent.
Mechanistically the channels are different questions either way, which is why they are kept apart.
Anchor similarity to a presented reference is largely presentation – the peptide carries an anchor
motif that reference’s alleles present – and it correlates with the binder score (r = +0.25 to
+0.33). TCR-face similarity correlates with nothing in the binding stack (|r| < 0.11 against
presentation and affinity) but strongly with the whole-peptide physicochemical log-odds
(r = +0.73 to +0.82; the row count behind that range was not recorded alongside it), which is
precisely why its sign moves once that term enters the model.
That earlier arrangement keeps its recorded coefficients: the letter
V names the generation, not the module.
Scores are log-odds, calibration is separate and explicit. score() returns signed
contributions and their sum on the log-odds scale, which is corpus-free. probability() maps
that sum to a risk of immune response against a named fitted corpus, because an absolute
probability is a property of the corpus’s prevalence and candidate generation, not of the peptide.
Callers who want a number in [0, 1] should say which corpus they mean.
The tested-neoantigen database is an annotation, never a fitted term. annotate() reports
the nearest validated-immunogenic neoantigen and its distance, and that is all it does. Every
labelled screen we hold is inside that database – retrieval recall at exact match is 1.000 on all
seven – so a fitted coefficient on it would be memorisation. Held out properly it still earns its
place as prior evidence: rebuilt without the test screen, fuzzy matching at two substitutions
recovers 0.08-0.34 of a screen’s positives where exact lookup recovers 0.00-0.26.
from mhcmatch import mimicry s = mimicry.score([“GILGFVFTL”], refs) # per-component log-odds + aggregate p = mimicry.probability(s, corpus=”screens”) # optional, named a = mimicry.annotate([“GILGFVFTL”], refs) # prior evidence, outside the model
- mhcmatch.mimicry.COMPONENTS = ('viral', 'self', 'thymus')#
Reference categories entering the fitted aggregate, in feature order.
neoagis deliberately absent – seeannotate().
- mhcmatch.mimicry.CHANNELS = ('anchor', 'tcr')#
The two signed channels every component is split into.
- mhcmatch.mimicry.params(cls='mhc1')[source]#
The frozen model: standardizer, coefficients with posterior sds, the radius per channel, the reference window totals behind the density normalisation, and the fit’s provenance.
Loaded on first use rather than at import, because
annotate()andmasks()are useful without a fitted aggregate and must not be blocked by its absence.- Parameters:
cls (str)
- Return type:
dict
- class mhcmatch.mimicry.MimicryScore(peptide, components, logodds, autoimmune, density=<factory>, nearest=<factory>)[source]#
Bases:
objectOne peptide’s mimicry read-out.
componentsis{component: {channel: signed log-odds contribution}};logoddsis their sum.autoimmuneis theselfcomponent’s total, reported separately because a self mimic is simultaneously a tolerance argument and a cross-reactivity liability for a vaccine, and those two license different decisions.- Parameters:
peptide (str)
components (dict[str, dict[str, float]])
logodds (float)
autoimmune (float)
density (dict[str, float])
nearest (dict)
- peptide: str#
- components: dict[str, dict[str, float]]#
- logodds: float#
- autoimmune: float#
- density: dict[str, float]#
- mhcmatch.mimicry.masks(length, cls='mhc1', peptide=None, register=None)[source]#
Positions each channel counts substitutions over, for one peptide.
anchoris the MHC-facing set andtcris its complement, so the two channels partition the peptide and no position is counted twice.The two classes take their anchors from different places, and the class is never inferred. Class I is anchored at fixed peptide positions (
mhcmatch.complement.ANCHORS, the same five the shipped role model calls MHC-facing), solengthalone determines the split. A class-II ligand is anchored by a 9-mer core that floats inside an 11–25-mer, so its split is a function of the register, not the length: passcls="mhc2"with thepeptide, and optionally aregisterto pin the frame instead of using the allele-agnostic heuristic (mhcmatch.complement.mhc2_anchors()). Reading a class-II ligand on the class-I layout labels the wrong residues as anchors and returns a confident, wrong face.>>> masks(9)["anchor"] [0, 1, 2, 7, 8] >>> masks(15, "mhc2", "AAAKFVAAWTLKAAA")["anchor"] # P1/P4/P6/P9 of the floating core [4, 7, 9, 12]
- Parameters:
length (int)
cls (str)
peptide (str | None)
register (int | None)
- Return type:
dict[str, list[int]]
- mhcmatch.mimicry.features(peptides, refs, cls='mhc1', *, threads=1)[source]#
Per-(component, channel) mimic density: hits per million same-length reference windows.
Density and not a raw count because the window totals span three orders of magnitude across components, channels and lengths, so a count standardized across that mix is largely measuring which length the peptide is.
log1pbecause the counts are heavy-tailed by construction.The radius per channel is the fitted one (
params()): the anchor channel is searched wider than the TCR channel because it projects onto more positions and so saturates later.Batched per (component, channel, length), not per peptide. The indices are already built per that key by
load_references(), so a per-peptidesearch()asked the same index six questions per peptide – one C++ entry per question, GIL held between them – wheresearch_batchasks it once for every peptide of that length with the GIL released. Same index, same radius, same hits; the loop that was there is the one rule 2 forbids.- Parameters:
refs (dict)
cls (str)
threads (int)
- Return type:
list[dict]
- mhcmatch.mimicry.corpus_R(peptides, spectrum, cls='mhc1', registers=None)[source]#
The Luksza corpus density per component, exactly – one table lookup per query window.
spectrumcomes fromcorpus_spectrum(). Returns{component: rho}per peptide, where\[\rho_k(q) \;=\; \frac{1}{m_k(q)\,N_k}\sum_{i=0}^{m_k(q)-1}\; \sum_{x} T_k[x]\,\beta^{\,d_H(f(q)[i:i+k],\,x)}\]with \(f(q)\) the TCR face, \(m_k(q)\) its sliding
k-mer count, \(T_k\) the reference k-mer frequency table, \(N_k=\sum_x T_k[x]\) its total mass and \(\beta=e^{-\kappa}\). The inner sum runs over the whole reference set – there is no radius and no k-nearest cutoff, because \(\beta^{d}\) is the threshold. It is evaluated as a table lookup because the weight factorises over positions (corpus_spectrum()).What the two divisors are for. \(N_k\) makes the value a density, so
thymus(26,513 peptides) andself(12 M proteome windows) are on one scale and the standing claim that thymus makes the other two redundant is a comparison rather than a size effect. \(m_k(q)\) makes it per query window, which is what removes the length artefact: the old fixed-face column varied 17x in mean across lengths 8-11 and correlated with length at Spearman -0.502, against +0.036 here (3,600 real epitopes, 900 per length, none in the reference).So \(\rho\in[0,1]\) is the expected mismatch weight between a uniformly chosen query window and a uniformly chosen reference window. The Luksza \(Z/(1+Z)\) saturation is dropped as redundant: it existed to bound an unbounded count, \(\rho\) is already bounded, and the old column never left its linear regime anyway (
Zstayed below 1.32e-3, soRwasZto three figures). \(a_0\) is gone with it – it was a scale the standardizer absorbed, and the length compensationexp(kappa*(L - a0))it carried is now the explicit \(m_k\) divisor.A peptide whose face cannot supply a
k-mer yields no key, exactly as an unindexed length did before. With the shipped ``k = 3`` this cannot happen for a canonical class-I ligand: the face isL - 5wide and the shortest ligand is an 8-mer. It fires only for a non-standard residue.>>> spec = corpus_spectrum(components=("thymus",)) >>> corpus_R(["GILGFVFTL"], spec)[0]["thymus"]
- Parameters:
spectrum (dict)
cls (str)
- Return type:
list[dict]
- mhcmatch.mimicry.corpus_counts(pmhc_dir=None, cls='mhc1', comp='thymus', k=3, self_species='human', weights=None, mask='slice')[source]#
(T, N): the sliding-k-mer count table over one reference component’s TCR faces.maskmust match the one the queries will use (CORPUS_MASKS); it selects the same face construction on the reference side, so the two are comparable by construction.Tis a flatA**karray of window counts with multiplicity – one increment per reference peptide per window, the published Luksza form, not per distinct face – andN = T.sum()the total reference window mass. Memoised per(cls, comp, k, species, weights)for the process; see_COUNTS.weights="locus"divides each reference peptide’s contribution by the size of itslocus_weights()component, so a protein region that appears many times over – a recurrent hotspot, a family of overlapping registers – counts once instead of many times. The density stays in [0, 1] becauseNis the same weighted mass. Off by default: it changes what the column means, so it is an arm to measure rather than a correction to assume.It applies to the assayed deposits (``thymus``, ``viral``) and is a no-op for ``self``, and that is the semantics rather than an optimisation. Locus weighting corrects a sampling bias: a peptide deposit over-represents whatever was detected or tested most, so one region can enter it many times. A proteome is not a sample – every window appears exactly once by construction, there is no “assayed more often” to undo, and sliding windows overlap their neighbours by definition, so a sequence-overlap grouping over 122 M of them would collapse each protein (and every homolog of it) into one degenerate component at ruinous cost.
Every length the class admits contributes, which is the point of sliding: a 9-mer reference informs an 11-mer query because the table is keyed on the k-mer, not on the length. A reference whose face is narrower than
kcontributes nothing and is skipped rather than padded.- Parameters:
cls (str)
comp (str)
k (int)
self_species (str)
weights (str | None)
mask (str)
- mhcmatch.mimicry.contract(T, kappa, k=3, kernel=None)[source]#
Apply the mismatch kernel along every axis of an
A**kcount table. One-time, ~1 ms.Ais inferred fromlen(T)(alphabet()), so a wildcard-masked21**ktable works unchanged.kerneldefaults to the Hamming formK = (1-beta)I + beta*Jwithbeta = exp(-kappa), kept only so older results stay reproducible; passblosum62_kernel()for the graded score. Any position-additive, ungapped score factorises this way; a gapped alignment does not, which is the one real limit.- Parameters:
kappa (float)
k (int)
- mhcmatch.mimicry.corpus_spectrum(pmhc_dir=None, cls='mhc1', components=None, k=3, shapes=None, self_species='human', weights=None, mask='slice', kernel=None)[source]#
Contracted sliding-k-mer tables over the TCR face, one per component. No search.
Returns
{component: (table, n_kmers, k, mask)}wheretableis a flatA**karray andn_kmersthe total reference window count. Feed it tocorpus_R(), which reads the mask back off the tuple so a query cannot be posed against a table built the other way.kernelis passed tocontract().Nonekeeps the original Hamming form; pass a callablekappa -> (A, A)array –blosum62_kernel()partially applied, orfunctools.partial(blosum62_kernel, mask=mask)– to grade the substitutions, since each component carries its ownkappa.Why there is no index here. The Luksza sum
Z = sum_r exp(-kappa*(a0 - s(q,r)))weights every reference by its similarity, and with an ungapped position-additive score the weight factorises:beta**hamming(u, x) = prod_p K[u_p, x_p], K = (1-beta)*I + beta*J, beta = exp(-kappa)
so the sum over the whole reference set is one 20x20 matrix applied along each axis of the k-mer frequency table – a tensor contraction, computed once. Every query is then a table lookup. That is exact where a radius-capped neighbour search is a truncation: measured against a brute-force all-vs-all over every reference k-mer, the contraction agrees to 5.6e-16, while the radius-2 search it replaced captured a median 0.4999 of the true sum (IQR 0.4115-0.5556, min 0.1539 over 600 real 9-mers). Cost went from ~46 s to 2.3 ms for 340,876 queries, and the ~7.5 GB proteome index the
selfchannel needed became a 64 KB table.The two halves are split on purpose:
corpus_counts()builds the table (expensive, and memoised) andcontract()applieskappa(~1 ms), so profilingkappacosts one build, not one per grid point.shapessupplieskappaper component (corpus_shapes());a0is not used and does not need to be, because the length compensation it stood in for is now done explicitly by normalizing per query window (seecorpus_R()).Species.
self_specieskeys every component, not justself, and this function honours it literally – pass"mouse"and you get the mouse tables, which is what makes the substitution measurable. But the SCORER does not call it that way: it resolves each component throughreference_species()first, and that map sends every mouse component to"human", in both classes. So a mouse run is scored against the human tables throughout, not the mouse ones, and this docstring once said the opposite.The mouse tables remain reachable and are the arm that measures the difference: they stand on 25,264 and 40,244 reference windows against 140,482 and 136,618 for class I. See
reference_species()for the per-component transfer measurements and for why thinness is not the explanation.- Parameters:
cls (str)
k (int)
shapes (dict | None)
self_species (str)
weights (str | None)
mask (str)
- Return type:
dict
- mhcmatch.mimicry.face_kmers(peptide, cls='mhc1', k=3, register=None, mask='slice')[source]#
Packed sliding
k-mers of the peptide’s TCR face; empty when it cannot carry one.The class-I anchor set is
{P1, P2, P3, POmega-1, POmega}, so undermask="slice"the TCR face is contiguous –peptide[3:L-2], widthL - 5– at every length; class II gathers its face from around the floating core instead, and the k-mers slide over that projection. Sliding rather than taking the whole face is what lets a query of one length be compared against references of another: the table is keyed on the k-mer, not on the length.Under
mask="wildcard"the anchors are replaced byWILDCARDin place rather than deleted, so the window count isL - k + 1and a k-mer may span a pocket. SeeCORPUS_MASKSfor why that changes whichkare admissible.The packing base follows the mask (
alphabet()), so a table built under one mask cannot be indexed with the other – the sizes differ andcorpus_R()checks it.>>> face_kmers("SIINFEKL").size # W = 3, exactly one 3-mer 1 >>> face_kmers("SIINFEKL", mask="wildcard").size # 8 - 3 + 1 6
- Parameters:
peptide (str)
cls (str)
k (int)
register (int | None)
mask (str)
- mhcmatch.mimicry.SHAPES: dict = {'self': 5.0, 'thymus': 3.0, 'viral': 8.0}#
Fitted decay
kappaper component, by profile deviance on the neoantigen corpus (bench/immuno/corpus_exact.py). Passed asshapes=only to re-measure it. The shipped aggregate records the same values ascorpus_shapes;corpus_shapes()reads that copy so a re-vendored refit moves the scored column with it, and this one is the fallback.``a0`` is gone, and the value is one number rather than a pair. It was a scale the standardizer absorbed, and the length compensation
exp(kappa*(L - a0))it carried is now the explicit per-window divisor incorpus_R().The three differ, and they differ for a reason. Writing
gamma = (1-beta)/(1+19*beta)withbeta = exp(-kappa)for the fraction of order-1 sequence structure the contraction keeps (contract()),thymussits at an interior optimum withgamma = 0.49– graded tolerance, near-misses count – whileviralruns togamma = 0.99, near-exact 3-mer matching.selfcarries no shape information at all: its profile moves 0.12 deviance units across a 24-fold range ofkappa, because the human proteome occupies essentially every cell of the 3-mer table, so smoothing cannot change the ordering. That is the mechanism behindselfreading as how many rather than whether.
- mhcmatch.mimicry.CORPUS_K: int = 3#
Width of the sliding k-mer over the TCR face. 3 is not a tuning choice, it is the only width every class-I length can supply: the face is
L - 5residues wide (the anchors are P1-P3 and POmega-1/POmega) and the shortest ligand is an 8-mer, soW_min = 3. Atk = 4an 8-mer has no window at all and atk = 5neither an 8- nor a 9-mer does – which reads as a low score and is really a structural zero, the exact failure mode this module removed. Measured on 3,600 real epitopes balanced 900 per length, none of them in the thymic reference, the normalized density correlates with peptide length at Spearman +0.036 for k=3 against +0.587 (k=4) and +0.830 (k=5) – and against -0.502 for the fixed-face column this replaced.
- mhcmatch.mimicry.LOCUS_W: int = 7#
Substring width at which two peptides are called the same locus by
locus_weights(). Seven residues links every register of one mutation –VVVGAVGVGK,VVGAVGVGKandGADGVGKSALare one KRAS G12 locus at 7 and three separate peptides at 11 – while being long enough that an incidental match between unrelated proteins is rare: there are 20**7 = 1.28e9 7-mers against ~1.2e7 windows in a whole proteome.
- mhcmatch.mimicry.locus_weights(peptides, w=7)[source]#
Per-peptide weights that make each locus count once, not each peptide.
Two peptides are the same locus when they share a substring of
wresidues; the relation is closed transitively, and every peptide in a component of sizemgets weight1/m, so the component contributes unit mass however many peptides it supplied.Why this exists. A peptide set is not a set of independent observations. One recurrent hotspot – KRAS G12, which is both the most-tested and the most-validated public neoantigen – contributes many overlapping registers of the same eleven residues, and an unweighted k-mer table reads that as evidence about immunogenicity when it is evidence about what got assayed. Measured on the immunogenic arm of the neoantigen fit corpus, the effect is large enough to dominate the set’s most distinctive 3-mers (
VGA,GAV,AVG– all fragments ofVVVGA[VDC]GVGKSA). Seebench/results/kmer_spectrum.md.Grouping is by sequence rather than by gene symbol because the corpora that need it do not all carry one, and because the register-level overlap is what actually repeats. Where a real source annotation exists, pass its own weights to
corpus_counts()instead.>>> [round(x, 3) for x in locus_weights(["VVVGAVGVGK", "VVGAVGVGK", "SIINFEKLAA"])] [0.5, 0.5, 1.0]
- Parameters:
w (int)
- Return type:
list
- mhcmatch.mimicry.corpus_shapes(artifact=None)[source]#
{component: kappa}– the decayC_corpuswas fitted with.Reads the shipped aggregate’s
corpus_shapeswhen it carries one, so a refit that is re-vendored moves the scored column rather than leaving it on a stale module constant; otherwise returnsSHAPES. Same convention asmhcmatch.luksza.shape().Accepts the older
(kappa, a0)pair form for reading an old artifact and keeps onlykappa; nothing writes that form any more.- Parameters:
artifact (dict | None)
- Return type:
dict
- mhcmatch.mimicry.score(peptides, refs, cls='mhc1', allow_missing=False, *, threads=1)[source]#
Signed per-component log-odds contributions and their sum, one per peptide.
Raises if
refsis missing a component. A missing feature standardizes to zero, so the aggregate would silently be a different, smaller model rather than an error – and the usual way to get here isload_references(with_self=False), which drops the component carrying the largest coefficients. Passallow_missing=Trueto accept that deliberately — its channels then come back NaN, and so doMimicryScore.logoddsandMimicryScore.autoimmune, because the fitted coefficients describe the full set. A zero onautoimmunereads as “no self-similarity found”; the truth underwith_self=Falseis “never looked”, and those license different decisions.- Parameters:
refs (dict)
cls (str)
allow_missing (bool)
threads (int)
- Return type:
list[MimicryScore]
- mhcmatch.mimicry.probability(scores, corpus='screens', cls='mhc1')[source]#
Map the aggregate log-odds to a risk of immune response against a named corpus.
The intercept is the corpus’s own base rate, so this number is not transferable: the seven neoantigen screens behind
"screens"run from 0.048 % positive to 46.8 %, and a probability quoted without naming the corpus is quoting one of those prevalences by accident. UseMimicryScore.logoddsto rank; use this only to report.- Parameters:
corpus (str)
cls (str)
- Return type:
list[float]
- mhcmatch.mimicry.annotate(peptides, pmhc_dir=None, cls='mhc1', max_subs=2, *, threads=1)[source]#
Nearest validated-immunogenic neoantigen and its distance. Prior evidence, not a score.
This is kept out of
score()on purpose. Every labelled screen we hold is contained in the tested-neoantigen database, so retrieval recall at distance 0 is 1.000 on all seven and a fitted coefficient would be measuring the answer key. Held out honestly the channel is still useful – with the test screen removed from the database, matching at two substitutions recovers 0.08-0.34 of its positives against 0.00-0.26 for exact lookup – which is why it is reported at all.- Parameters:
cls (str)
max_subs (int)
threads (int)
- Return type:
list[dict]
- mhcmatch.mimicry.NEOAG_COLUMNS: tuple = ('neoag_distance', 'neoag_nearest', 'neoag_n_within', 'known')#
The columns
annotate()appends, in order, afterpeptide. It lives here rather than inline in the CLI or in a pipeline stub for the same reasonmhcmatch.rank.BASE_COLUMNSdoes: a consumer has to be able to name the schema without running the command, and a schema typed out a second time is a schema that drifts.
- mhcmatch.mimicry.load_references(pmhc_dir=None, cls='mhc1', with_self=True, self_species='human', lengths=None)[source]#
Reference window sets per (component, channel, length), ready for
features().with_self=Falseskips the host proteome, which dominates the cost. The aggregate is not defined withoutself– it carries the largest coefficients in the fit – soscore()raises unless the caller passesallow_missing, andmhcmatch rank --score aggregaterefuses the combination outright.``lengths`` is the difference between usable and not at class II. The index is per-length and the cost is per-length: one proteome pass is ~11 s for
window_arrayplus ~1.0 min to resolve a class-II register for each of its 12,685,964 windows. The class admits fifteen lengths (11-25), so building all of them is ~19 min; class I admits four, at ~65 s. A run queries the lengths its own peptides have – usually one or two – and building the other thirteen is work nobody asked for. Pass the queried lengths and pay for those: ~75 s for one class-II length, ~17 s for one class-I length.None(the default) keeps every length the class admits, so an existing caller is unchanged. A length outsidemimics._LEN[cls]was never indexed and still is not;features()skips a query it has no index for.The build itself is vectorized:
mhcmatch.proteome.Proteome.window_array()replaces a per-window Python loop (2.7x), and the per-channel projection is onenp.uniqueover a fixed-width byte view rather than asetdefaultover 12 M strings.- Parameters:
cls (str)
with_self (bool)
self_species (str)
- Return type:
dict
- mhcmatch.mimicry.safety(scores, top=5, symbols=None)[source]#
Where the self/thymus mimics are expressed – the autoimmunity read-out, made actionable.
There is no tumour argument, and there used to be one that did nothing.
tumorsat at positional #2 and was never read –mhcmatch.expression.safety_profile()conditions on no context at all – sosafety(scores, "SKCM")returned the pooled profile while reading as if it had been conditioned. On a read-out whose job is to say which tissue you cannot afford to damage, a caller believing they narrowed the question is the dangerous direction. Removed rather than accepted-and-ignored; add it back only whensafety_profilecan actually take one.A
selforthymushit says the candidate resembles a peptide the body presents; the decision it feeds is whether the tissue presenting it is one you can afford to damage. That needs the mimic’s source protein, which is whyload_references()carries it andfeatures()keeps it rather than collapsing everything to a density.Returns, per peptide, the tolerance-side hits with
mhcmatch.expression.safety_profile()resolved for each source that maps to a gene.The ``thymus`` deposit names its sources as UniProt accessions (
Q8WZ42) whilemhcmatch.expression.safety_profile()is keyed on HGNC symbols (TTN), so the join needs a map: passsymbols=frommhcmatch.proteome.gene_symbols(path, key="accession")(). Without it an accession simply fails to resolve andprofilecomes back empty – which reads as no risk found and is the dangerous direction to be wrong in, somhcmatch.vector.self_origin_risk()refuses to run without the map rather than defaulting it.One gap remains and it returns an empty ``profile`` rather than a guess. The
selfcomponent is built from proteome windows, which carry no source column at all, so it is not resolvable to a gene however good the map is. The mimic peptide and its raw source are always returned.geneis the resolved symbol andsourcestays the deposit’s own identifier, because a withdrawal decision has to be traceable back to the record it came from.- Parameters:
top (int)
- Return type:
list[dict]
- mhcmatch.mimicry.CORPUS_REFERENCE: dict = {('mouse', 'self'): 'human', ('mouse', 'thymus'): 'human', ('mouse', 'viral'): 'human'}#
(query species, component) -> reference species, for the components where the two differ. Read byreference_species(), which is the only thing that should consult it, and applied bycli._aggregate_channelswhen it builds the scoring channels.One entry: a mouse ``thymus`` query is scored against the HUMAN thymic table. The mouse thymic deposit is one haplotype. Of its 6,661 class-I peptides, every one of the 2,663 that carries an allele annotation is
H-2Db(1,574) orH-2Kb(1,089) – no other haplotype appears at all – so its k-mer table is the H-2b groove’s motif rather than a measure of what the thymus presents. That column is then applied to a fit spanning six H-2 allotypes, and it is collinear withbinder, which already scores the groove.The mouse class-I corpus block reads human tables throughout – see
reference_species()for the per-component measurement behind one uniform rule.
- mhcmatch.mimicry.NATIVE_CORPUS_COMPONENTS: tuple = ('self', 'thymus')#
The host compartments, and the two components
native=Truewill route back to the query species’ own table.viralis deliberately not one of them: it is not a host compartment, so “the mouse’s own” makes no sense for it in the way it does for a proteome or a thymus. What a mouse viral table samples is 9 H-2 allotypes against the human table’s 129 – a thinner sample of the same pathogen ligandome, not a different compartment – so there is nothing to recover by switching, andCORPUS_REFERENCEkeeps it human even undernative.
- mhcmatch.mimicry.reference_species(species, comp, native=False)[source]#
Which species’ reference table a
speciesquery of componentcompis scored against.native=Trueoverrides the redirect below for the host components named inNATIVE_CORPUS_COMPONENTSand returnsspeciesitself, so a mouse query is matched against the mouse tables. It is not the default, and the reason is measured, not assumed – see the per-component transfer figures below, and in particularthymusatr= 0.3245 with the matched-mass arm that rules out thinness as the explanation. mhcmatch rank –native-corpus warns on every run that uses it. All twelve tables ({mhc1,mhc2}|{self,thymus,viral}|{human,mouse}|3) ship, so this is a routing switch and nothing is fetched or rebuilt.A coefficient fitted under one routing does not transfer to the other. Every shipped mouse artifact was fitted with the human tables, so
--native-corpusscores those coefficients against a different feature. Use it to measure the mouse tables, not to rank against a fit that never saw them.Usually
speciesitself. The exception is mouse class I, whose whole corpus block – ``thymus``, ``self`` and ``viral`` alike – is matched against the human tables, because the mouse reference deposits are too small and too groove-skewed to be a reference.Nothing is trained on human data by this. A corpus channel is a k-mer density lookup: the query peptide’s TCR-facing windows are matched against a counted reference table, and the table is the only thing that is human. Every coefficient in
aggregate_mhc1_mouse.jsonis fitted on mouse neoantigens, and presentation, expression and physicochemistry all read mouse sources. What crosses the species line is the reference corpus being matched to, nothing else.The corpus channel transfers to the extent that its reference deposit is not one groove’s motif, and the three components sit at three points on that axis. Measured on the 921-row mouse class-I fit population (bench/epic/corpus_transfer.py in the benchmark repo),
rbeing the Pearson correlation between the same peptide’s density under the two species’ tables:self– no groove at all (an unpresented proteome window is not presented by anything), 112,565,681 mouse against 121,968,158 human reference windows,r= 0.9990. The same table twice over: the substitution is free in either direction and taking human keeps one reference source.viral– 9 mouse allotypes (H-2Kb50.2 %) against 129 human,r= 0.8382. A 9-allotype sample of a 129-allotype space.thymus– 2 mouse allotypes (H-2Db1,574,H-2Kb1,089, nothing else) against a pooled human donor panel, 25,264 against 140,482 windows,r= 0.3245. The H-2b motif, and nothing else.
It is not a sample-size effect, and that was measured rather than assumed. The mouse thymic table stands on 25,264 reference windows against human’s 140,482, so thinness is the obvious explanation and it is the wrong one: thinning the human deposit at the peptide level to the mouse table’s window count, 40 draws, still reproduces the full human column at r = 0.8933 (range 0.8728-0.9109) and still disagrees with the mouse table at 0.2903 (0.2467-0.3310). A human table cut to mouse’s size does not become the mouse table. What differs is which grooves each deposit sampled – so depositing more mouse thymic peptides from the same two allotypes would not close it, and the substitution is the fix rather than a stopgap.
This is the same failure
background="ligand"had: a pooled statistic over a pool dominated by one allotype measures that allotype. Mouse class II is the extreme case – its thymic deposit is 1,490 peptides, allI-Ab.The redirect is not keyed on class, and that matters. This function takes
(species, comp)only, so a mouse class-II query is routed to the human class-II tables by the same rule –clsis carried separately and selects which of the twelve tables is read. The point was once moot: both class-II artifacts were the six-term neoantigen fits, which declare no corpus block, so nothing read a class-II corpus table. The two class-II pathogen fits do declare one, and they were fitted under this redirect –C_corpus_selfresolves in both (p = 1.1e-03 human, 3.5e-10 mouse) andC_corpus_thymusin neither mouse cell (p = 0.24), which is what theI-Ab-only thymic deposit above predicts.Expression is NOT covered by this and must not be. Human and mouse organs and tumours are different tissues, so a human expression level is not a stand-in for a mouse one at any sample size;
mhcmatch.expressionstays species-keyed throughout. Mapping a gene to its mouse orthologue is a different operation – it fixes gene identity and still reads a mouse transcriptome.- Parameters:
species (str)
comp (str)
native (bool)
- Return type:
str
mhcmatch.mimics module#
Cross-reactivity by category, never summed: a hit in the thymic immunopeptidome argues tolerance
and autoimmune risk, a hit in the host proteome argues peripheral cross-reactivity without implying
presentation, and a viral or bacterial hit argues a pre-existing repertoire may cross-react.
KINDS states which is which; neighbours() runs the
whole query set through one threaded C++ index per (category, length).
Molecular-mimicry annotation for strong binders.
For each strong-binding neoantigen, search reference peptide sets for mimics — near-identical
presented peptides — and report the presentation-aware E-value (mhcmatch.search.find_mimics(),
lower = more significant mimicry) per category:
thymus — the thymic self-immunopeptidome (HLA Ligand Atlas). A significant thymic mimic means the neoantigen resembles a self-peptide presented during negative selection: reactive T cells were likely deleted (reduced immunogenicity) and it flags cross-reactivity / autoimmune risk for a cancer vaccine.
self — a window of the host proteome with no evidence of thymic presentation. Encoded is not presented, so the tolerance argument is weaker than
thymus; the peripheral cross-reactivity risk is the same. Kept as its own category rather than merged intothymusbecause those two hits license different conclusions.viral / bacterial — foreign presented peptides / pathogen and commensal proteomes. A foreign mimic can raise immunogenicity (a pre-existing anti-pathogen repertoire cross-reacts) — molecular mimicry.
neoag — the tested-neoantigen database: has this (or a near-identical) neoantigen been reported.
KINDS states what each category argues; PROTEOME_REFS names the proteomes behind
self and bacterial, which are opt-in (load_reference_sets(..., proteomes=("bacterial",)))
because building their window sets is the expensive part.
This scores cross-reactivity, not presentation or immunogenicity directly; compose it with the
presentation / affinity scores from mhcmatch.predict. Reference data: the isalgo/pmhc_data
compendium (thymus/, ligandome/, immunogenicity/, proteome/).
- mhcmatch.mimics.KINDS = {'bacterial': 'foreign', 'neoag': 'database', 'self': 'self', 'self_mouse': 'self', 'thymus': 'self', 'viral': 'foreign'}#
What a hit in each category argues. Not the same question per category, which is why they are never summed into one “mimicry score”:
category
a hit means
thymuspresented in the thymus, so reactive clones met it during negative selection. Lowers expected immunogenicity; raises autoimmune risk for a vaccine.
selfa window of the host proteome not known to be thymically presented. Weaker tolerance argument – being encoded does not imply being presented – but the same peripheral cross-reactivity risk. Kept distinct from
thymuson purpose.virala foreign presented peptide. Raises expected immunogenicity: a pre-existing anti-pathogen repertoire may cross-react.
bacterialthe same, from pathogen and commensal proteomes.
neoagalready tested as a neoantigen somewhere. Prior evidence, not biology.
- mhcmatch.mimics.DEFAULT_REFS = {'neoag': ('neoantigens/neoag_tested.tsv.gz', 'database'), 'thymus': ('thymus/thymus_immunopeptidome.tsv.gz', 'self'), 'viral': ('ligandome/viral_foreign_iedb.tsv.gz', 'foreign')}#
(folder/file under pmhc_data, kind).
selfis the tolerance reference passed asfind_mimics’self_set; the rest are foreign/database sets.- Type:
Default reference categories
- mhcmatch.mimics.SPECIES_REFS = {('thymus', 'mouse'): 'thymus/thymus_immunopeptidome_mmu.tsv.gz'}#
Per-species overrides of a
DEFAULT_REFSpath.The two deposits carry species differently, and the difference is in the data, not a style choice.
viralis one file whose ownmhc_speciescolumn holds both – 2,773 distinct mouse peptides among 44,993 – so it needs no entry here and is selected withspecies=at load time.thymusis one file per species, because the human deposit is a single multi-donor atlas and the mouse one is assembled from three studies with different antibodies and search pipelines; pooling them into one table would bury that in a column nobody filters on.Resolve through
ref_path()rather than indexing this directly, so a category with no override falls back toDEFAULT_REFSinstead of raising.
- mhcmatch.mimics.CLS_REFS = {('self', 'mhc2', 'human'): 'ligandome/self_ligandome_mhc2_hsa.tsv.gz', ('self', 'mhc2', 'mouse'): 'ligandome/self_ligandome_mhc2_mmu.tsv.gz', ('thymus', 'mhc2', 'human'): 'ligandome/hla_ligand_atlas_mhc2_presented.tsv.gz', ('viral', 'mhc2', 'human'): 'ligandome/viral_ligandome_mhc2_pan.tsv.gz', ('viral', 'mhc2', 'mouse'): 'ligandome/viral_ligandome_mhc2_pan.tsv.gz'}#
Class-II overrides, keyed ``(category, cls, species)``, consulted before
SPECIES_REFS.Two of the three class-II references are a different object from their class-I namesakes, and both differences are forced by what the reference has to argue:
selfat class I is the host proteome and stays one –C_corpus_selfis a fitted term of two shipped class-I artifacts, so redefining it would be a model change. At class II the proteome saturates the table (8,000 of 8,000 3-mer cells, ~111 M window mass), and a near-uniform density cannot discriminate. The class-II channel therefore reads a self ligandome – IEDB class-II ligands whose source protein is the host’s own, 421,174 human and 25,314 mouse peptides – which is what “presented” means and what tolerance is about.viralis pan-species at class II. The deposit carries nomhc_speciescolumn, soload_peptides()keeps every row whichever species asks. Splitting it left the mouse cell on 464 peptides against human’s 14,468; a viral k-mer density is not species-specific the way a proteome is, and pooling deepens the mouse reference 31x while moving human 14,468 -> 14,645.thymusat human class II is the whole displayed proteome, not the thymus – every benign tissue the HLA Ligand Atlas sampled on a class-II molecule, 132,818 peptides against the thymic fraction’s 27,987. The channel was never thymic: measured across the atlas’s 29 tissues on 7,948 human class-II T-cell rows of which 5,150 are immunogenic, lung leads at contrast AUROC 0.5496 (95% CI 0.5365-0.5627) and thymus is second at 0.5475 (0.5345-0.5606) – the same result class I gives, where thymus is 6th of 29. The slot keeps the namethymusbecauseC_corpus_thymusis a released column name and renaming it would misname the model a reader downloads; every reader-facing label says displayed self. Mouse class II is untouched by this and still reads its own thymic deposit.
- mhcmatch.mimics.ref_path(category, species='human', refs=None, cls=None)[source]#
The compendium path for
categoryinspecies, forclswhen it matters.Resolution order is most specific first:
CLS_REFS(category,cls,species), thenSPECIES_REFS(category,species), then the species-agnosticDEFAULT_REFS. Only a deposit that is physically split needs an entry above the last.cls=Noneskips the class layer, which is what a caller with no class in hand wants; it is not the same ascls="mhc1", because a class-I-only override would then be reachable from a call that never named a class.- Parameters:
category (str)
species (str)
refs (dict | None)
cls (str | None)
- Return type:
str
- mhcmatch.mimics.PROTEOME_REFS = {'bacterial': ('ecoli_K12_UP000000625', 'saureus_UP000008816', 'lreuteri_UP000001991', 'mgnavus_UP000018690', 'egallinarum_UP000254807'), 'self': ('human',), 'self_mouse': ('mouse',)}#
Reference proteomes per category, as
mhcmatch.store.fetch_proteome()stems. Gut commensals (L. reuteri, M. gnavus, E. gallinarum), a gut/lab strain (E. coli K12) and a skin/nasal pathogen (S. aureus) – the exposures a human repertoire has plausibly seen.
- mhcmatch.mimics.CANONICAL_LEN = {'mhc1': range(8, 12), 'mhc2': range(11, 26)}#
Canonical presented lengths per class, and what every shipped number rests on. The class-I groove is closed at both ends, so 8-11 is the range the corpus tables, the anchor models and every fitted artifact were built over.
- mhcmatch.mimics.EXTENDED_LEN = {'mhc1': range(12, 15)}#
Extended class-I lengths, 12-14 – real ligands, and deliberately NOT implemented.
MHC I does present past 11 by bulging the peptide out of the groove, and the deposits carry those rows: filtered to
mhc_class = MHCI, viral_foreign_iedb.tsv.gz holds 157 peptides of length 12-13 (0.3 % of its 55,084) and thymus_immunopeptidome.tsv.gz 195 (0.8 % of 25,891); the human neoantigen deposit holds 84,622 rows carrying 276 immunogenic peptides outside 8-11. Mouse holds 1.They are excluded on purpose and the exclusion is held, not fixed: admitting them would rebuild every corpus table, move the human fit population that artifact version 11 was fitted on, and make no published number reproduce.
length_range()therefore refuses the extended range rather than returning it – a range no downstream path was fitted for is worse than an error, because it would score silently and wrongly.
- mhcmatch.mimics.length_range(cls, extended=False)[source]#
The presented lengths
clsadmits.extended=Trueis not implemented and raises.- Parameters:
cls (str) –
"mhc1"or"mhc2".extended (bool) – ask for the extended class-I range (12-14) instead of the canonical one.
- Returns:
a
rangeof peptide lengths.- Raises:
NotImplementedError – when
extendedis true, naming what would have to change.ValueError – when
clsis not a known class.
The two ranges are
CANONICAL_LENandEXTENDED_LEN; see the latter for how many real ligands the canonical cut leaves out and why the cut is held anyway.
- class mhcmatch.mimics.MimicResult(binder, allele, category, n_exact, n_near, top_mimic, top_subs, e_value, n_hits, significant)[source]#
Bases:
objectPer-(binder, category) mimicry summary.
A mimic is a reference peptide of the same length within
near_subssubstitutions of the binder (T cells cross-react across a few substitutions).n_exact/n_nearcount identical and near-identical mimics;top_mimic/top_subsare the closest one.e_value/n_hitsare the raw presentation-aware search stats, kept for reference.- Parameters:
binder (str)
allele (str)
category (str)
n_exact (int)
n_near (int)
top_mimic (str)
top_subs (int)
e_value (float)
n_hits (int)
significant (bool)
- binder: str#
- allele: str#
- category: str#
- n_exact: int#
- n_near: int#
- top_mimic: str#
- top_subs: int#
- e_value: float#
- n_hits: int#
- significant: bool#
- mhcmatch.mimics.load_peptides(pmhc_dir, rel_path, cls, species='human')[source]#
The
peptidecolumn of a compendium TSV, filtered tocls/speciesand plausible presented lengths. Rows without a class/species field are kept (some sets are unlabelled).pmhc_dir=Nonefetches the file from the public HF dataset instead (cached, and overridable with$MHCMATCH_PMHC_DIR), so a fresh install needs no pre-staged mirror.- Parameters:
rel_path (str)
cls (str)
species (str)
- Return type:
list
- mhcmatch.mimics.proteome_peptides(category, lengths)[source]#
Every distinct
lengths-mer window of the reference proteomes of onePROTEOME_REFScategory, standard residues only.A proteome is a sequence reference, not a ligandome: these windows are what the source organism encodes, with no claim that any of them is presented. That is the whole point of keeping
selfseparate fromthymus– seeKINDS.lengthsis required rather than defaulted to the full class range because the cost is linear in it and dominated by the host: the human proteome has ~11.4 M distinct 9-mers, so asking for all of 8-11 is several GB. For the humanselfcategory prefermhcmatch.proteome.Proteome, which indexes on demand instead of materialising a set.- Parameters:
category (str)
- Return type:
list
- mhcmatch.mimics.proteome_window_array(category, L)[source]#
Sorted
|S{L}array of every distinct standard-AA window of aPROTEOME_REFScategory – the vectorized counterpart ofproteome_peptides()for one length.A consumer that projects or indexes these windows wants the array; materialising a 12 M-element Python set to hand it one is most of what
proteome_peptides()costs.- Parameters:
category (str)
L (int)
- mhcmatch.mimics.load_reference_sets(pmhc_dir=None, cls='mhc1', species='human', refs=None, proteomes=())[source]#
(self_set, foreign_sets)forscan().self_setis the single tolerance reference (theself-kind entry, thymus by default);foreign_setsis{name: [peptides]}for the rest.refsoverridesDEFAULT_REFS.proteomesaddsPROTEOME_REFScategories ("bacterial","self") built from FASTA windows over the class’s plausible presented lengths. They land inforeign_setswhatever theirKINDSentry says, becausefind_mimics()computes its E-value against exactly one background and that slot is already the thymic set –KINDSis how a caller reads a category, not how it is passed.- Parameters:
cls (str)
species (str)
- Return type:
tuple
- mhcmatch.mimics.neighbours(peptides, ref_sets, max_subs=2, threads=1)[source]#
{peptide: {category: [(n_subs, ref_peptide), ...]}}– same-length mimics, in batch.The Hamming half of
scan(), and 4,300x faster than the path :func:`scan` used to take. Oneseqtree.Indexper (category, length), queried withsearch_batchin parallel C++ with the GIL released: measured at 237,000 queries/s against 55 for onefind_mimics()call per peptide, returning identical counts and distances (benchmark repo,bench/results/neighbour_search_speed.md).What it does not give is the per-allele presentation-aware E-value – that genuinely needs the k-mer/allele index
find_mimics()builds, which is why this is a second entry point and not a replacement. Use it when the question is how close is the nearest reference peptide, andscan()withevalue=Truewhen the significance of the match is the question.Hits are nearest-first and exclude the query itself (distance 0); same length only, which is what Hamming distance means.
Reference peptides are deduplicated, and that is a fix. The compendia repeat a peptide once per allele/source-organism it was reported under – the viral IEDB set is 57,331 rows over 26,640 distinct peptides – so counting rows makes
n_neara function of how often a sequence was deposited rather than of the sequence neighbourhood. On 400 Chowell peptides this changes exactly one result (AAAAATMALhadn_near = 3, all three the same peptideEAAAATCAL);top_mimicandtop_subsare unaffected either way.>>> neighbours(["GILGFVFTL"], {"viral": ["GILGFVFTA", "DDDDDDDDD"]}) {'GILGFVFTL': {'viral': [(1, 'GILGFVFTA')]}}
- Parameters:
max_subs (int)
threads (int)
- Return type:
dict
- mhcmatch.mimics.scan(binders, self_set, foreign_sets, cls='mhc1', max_subs=2, near_subs=2, self_name='thymus', exclude_query=False, evalue=False, threads=1)[source]#
Mimic-scan an iterable of
(peptide, allele)binders. Returnslist[MimicResult](one per binder × category with >=1 same-length reference peptide withinnear_subssubstitutions).self_setis the tolerance reference (categoryself_name);foreign_setsis{name: [peptides]}.max_subsis the fuzzy-search radius.find_mimics()excludes the exact query (a neoantigen’s identical peptide is its source, not a mimic), son_exactis a direct set-membership check andn_nearcounts same-length reference peptides 1..``near_subs`` substitutions away (from the fuzzy hits, by exact Hamming distance). Onefind_mimics()call per binder scores every category at once.``exclude_query=True`` is required whenever the output becomes a model feature. By default
n_exactanswers is this peptide in the reference set, which is the right question for a lookup (“is this a known viral ligand?”) and the wrong one for a feature: a reference assembled from the same deposit as the labels will contain the positives, so a known epitope scoresn_exact = 1for free and the model reads self-identity as foreignness.Not hypothetical. 45% of the Gfeller cohort’s immunogenic peptides are exact matches to the viral IEDB ligand set, and a foreignness term built that way scored 0.714 AUROC there against 0.554 once self-matches were excluded (
bench/results/gfeller_contamination.md). Withexclude_query=Truea peptide is never its own mimic and only genuine neighbours at 1..``near_subs`` substitutions contribute.``evalue`` defaults to ``False``, the fast path, from 1.20.0 — it used to default to ``True``. On the fast path the search runs through
neighbours()(one batched, threaded index query per category and length) instead of onefind_mimics()call per binder, which is 4,300x the throughput on measured identical answers: every field excepte_valueandn_hitsis identical, andtests/test_mimics.pypins that equality.The default moved because the expensive path was the one you got without asking, while the docstring recommended the other – and find_mimics builds its index per call, so the cost is quadratic in the obvious usage. What you give up is real and not silent:
e_valuecomes backnanandn_hitsbecomes the near-hit count, so a caller reading significance seesnanrather than a plausible wrong number. Passevalue=Truewhen the per-allele presentation-aware significance is the question; it now honoursthreads, which it previously accepted and dropped.