Complementarity — the I and C of EPIC#

The recognition axis: chemistry, identity and the corpus channels.

mhcmatch.complement module#

The recognition axis as one score: whole-peptide physicochemistry and length, the same components split MHC-facing vs TCR-facing, MJ1996 / repertoire-marginalised TCRen contact potentials, hydrophobic-run and dipeptide motifs, and per-role residue log-odds — pooled, per length bin (8/9/10/11+ at class I, quartiles at 14/16/19 at class II) and per position zone (relative thirds of the TCR face at class I, the nflank/core/cflank register zones at class II, selected by cls=), whose pooled aa_anchor/aa_tcr pair reproduces mhcmatch.posbayes.llr() exactly, so that model is a strict special case. Linear head, because a diagonal Gaussian cannot represent a summed log-odds; the EM Gaussian parameters ship alongside for comparison. Vectorised — pass a list, not a loop. See Complementarity — the I and C of EPIC.

Complementarity: how well a presented peptide complements a T-cell repertoire.

This is the recognition axis. A whole-peptide physicochemical predictor – summed two principal components of the amino-acid property matrix over the whole peptide and added its length: 13 parameters, fitted by EM. That construction is kept and five blocks are added, each answering something the pooled version provably cannot:

block

what it adds

phys

PC1/PC2 of the property matrix summed over the peptide, plus length. The whole-peptide feature set, kept as the floor everything else is measured against.

role

The same components computed separately over MHC-facing and TCR-facing residues, plus Kidera KF4 (hydropathy) per role. The two channels carry opposite-sign contributions for several residues; a pooled sum reports their difference, weighted by the corpus’s composition.

pot

Contact potentials, one per side: MJ1996 on the anchors (burial in a groove – MJ is 96.4% one-body with hydropathy as its dominant mode) and TCRen marginalised over a real CDR3 repertoire on the TCR-facing residues. TCRen is only 3.29% one-body, so no per-residue scale can be extracted from it and the unknown receptor side is integrated out instead.

motif

Contiguity of the hydropathy stretch: longest run, number of runs and above-threshold fraction of hydrophobic TCR-facing residues. A run of 3-4 is a different object from the same residues scattered, and no sum can express the difference. KD_THRESHOLD defines “hydrophobic” and encode() defines what breaks a run.

aa

Residue identity as a log-odds per amino acid per role, and the block that knows the peptide’s geometry: a pooled anchor/TCR pair whose sum is exactly mhcmatch.posbayes.llr(), one pair per length bin, and a position key. Same shape in both classes; the length bins are 8/9/10/11+ at class I and quartiles of 11-25 at class II, and the position key is relative thirds of the TCR face at class I and register zones at class II.

kmer

The same over adjacent TCR-facing residue pairs – a preference for a specific dipeptide that no marginal composition feature can express.

The head is linear, and that is a measured choice. The shipped posbayes score is a sum of the two role log-odds – weights fixed at 1 on two of these columns. A diagonal-covariance Gaussian classifier cannot represent that: it maps each column through its own quadratic and re-weights by inverse class variances, so the additive form is outside its hypothesis space and the extra blocks are paid for out of a worse fit to the term carrying most of the signal. On the training corpus the EM Gaussian reaches 0.657 grouped-CV AUROC on the aa block where the plain sum reaches 0.711. A linear head contains the sum as a special case, so whatever it adds is genuinely an addition. The EM Gaussian parameters are vendored alongside anyway, so the comparison stays re-checkable.

It emits a log-odds, not a probability. Like posbayes, the corpus’s own base rate is divided out, so a caller supplies whatever prevalence their setting actually has:

logit P(immunogenic) = score(peptide) + log(prior / (1 - prior))

and posterior() does exactly that.

Vectorised. score() takes an iterable and returns an array; the whole feature set is two (n, 20) count matrices times a handful of property vectors, so scoring a full deposit is seconds, not minutes. Pass a list, not a loop.

>>> from mhcmatch import complement
>>> complement.score(["GILGFVFTL", "SIINFEKL"])
array([1.79, 0.42])

Parameters are vendored per species in mhcmatch/data/complement_mhc1_{human,mouse}.json and never refitted at import. The two hosts are never pooled: different MHC, different thymic repertoires, so one fit across them is fitting a mixture. score(peps, species="mouse") selects. Provenance, the block-by-block cross-validation, the corpus-transfer matrix and the size-matched cross-species transfer are in the benchmark repo (bench/results/complementarity.md).

Both classes, and they are two fits rather than one with a parameter. score(peps) is class I: the role split is P1-P3, PΩ-1, PΩ at fixed peptide positions. score(peps, cls="mhc2") takes the anchors from the P1/P4/P6/P9 core of the floating 9-mer register (mhcmatch.store.anchor_indices()) and reads its own vendored tables, fitted on 603,781 human and 50,258 mouse class-II peptides from the IEDB export. Passing a class-II ligand to the class-I path labels the wrong residues as anchors and returns a confident, wrong number, so the class is an argument and never inferred from the length.

Parameters per class and species in complement_{mhc1,mhc2}_{human,mouse}.json. Validation is in bench/results/complementarity.md and complementarity_mhc2.md.

The shipped model’s recognition term lives here too, under a different name. score() is the thirty-column fitted head above; burial() is C_phys, one of the two Complementarity factors the fitted aggregate carries (mhcmatch.rank) – a single imported residue scale (PHYS_SCALE) summed over the TCR face, with no fitted residue parameters, which is why it cannot memorise the corpus. What it buys over the fitted head, and how the scale was chosen, is on burial() itself and in docs/burial.rst. Its partner C_corpus is mhcmatch.mimicry.corpus_R().

mhcmatch.complement.ANCHORS = (0, 1, 2, -2, -1)#

MHC-facing positions, signed, matching mhcmatch.immuno.ANCHOR_SCHEMES "pockets" and mhcmatch.posbayes.ANCHORS.

mhcmatch.complement.PARATOPE = {'A': (0.0468, 0.3668), 'C': (0.1994, 0.5024), 'D': (0.1813, 0.5419), 'E': (0.1014, 0.313), 'F': (0.0754, 0.3616), 'G': (-0.0, 0.2369), 'H': (0.2189, 0.5621), 'I': (0.0486, 0.416), 'K': (0.1234, 0.4516), 'L': (-0.0251, 0.1864), 'M': (0.0345, 0.3921), 'N': (0.1193, 0.3291), 'P': (0.0469, 0.2501), 'Q': (0.0384, 0.3411), 'R': (0.0515, 0.421), 'S': (0.0124, 0.2607), 'T': (0.1307, 0.3778), 'V': (0.0143, 0.2991), 'W': (0.0847, 0.4427), 'Y': (0.0172, 0.2474)}#

TCRen marginalised over 28,250,990 TRB IMGT CDR3 loops – paratope(a) = sum_b f(b) TCRen(b,a) – and its spread over the same distribution. A residue can have a mild mean energy and still be highly discriminating across receptors, so both are features. Measured in the benchmark repo (bench/results/paratope_basis.md); the 32M-clonotype repertoire is not needed at runtime.

mhcmatch.complement.PARATOPE_CONTACT = {'A': (0.1288, 0.4503), 'C': (0.1266, 0.5443), 'D': (0.1674, 0.6565), 'E': (0.0623, 0.3612), 'F': (0.0375, 0.3135), 'G': (-0.0177, 0.2136), 'H': (0.3078, 0.5235), 'I': (0.0233, 0.4399), 'K': (0.1341, 0.4223), 'L': (0.0396, 0.1878), 'M': (0.0131, 0.375), 'N': (0.0678, 0.3208), 'P': (0.0365, 0.2677), 'Q': (0.005, 0.2787), 'R': (0.0805, 0.4688), 'S': (0.031, 0.2542), 'T': (0.0965, 0.3768), 'V': (0.0596, 0.2969), 'W': (0.0887, 0.4643), 'Y': (0.037, 0.2607)}#

The same potential marginalised over the residues that actually contact peptide, rather than over the whole loop – f_contact(a) = sum_{clonotypes, i} P(contact | locus, i, L) * 1[cdr3_i = a], geometry from the 370 structures and identity from the same 28M clonotypes.

TCRen is a directed contact potential, so the flat loop composition behind PARATOPE is the wrong measure for it, and wrong in a biasing direction: only 35.4% of TRB loop residues ever contact a peptide, and the germline V-encoded head and J-encoded tail carrying most of the flat mass are exactly the positions the structures put at P(contact) = 0.00 (TRB, length 12: positions 1-2 and 9-12). Conditioning on contact is the germline-flank trim; it needs no separate rule. Composition is never taken from the crystals – 370 complexes are a poor estimate of which residues sit mid-CDR3 – only the geometry is.

It is a materially different vector, not a rescaling: Spearman against PARATOPE is +0.7549, 19 of 20 residues change rank, and the spread widens from 0.2440 to 0.3255. A structure-free second route restricted to the N-D-N insert agrees at Spearman +0.7173.

Generated by bench/immuno/paratope_contact_basis.py; bench/results/paratope_contact.md has the composition and paratope_contact_basis.md this vector. Opt in with score(..., paratope="contact"); the default stays "loop" so no recorded number moves.

mhcmatch.complement.PARAMS: dict = {'alphabet': 'ACDEFGHIKLMNPQRSTVWY', 'anchors': [0, 1, 2, -2, -1], 'arm': 'chowell_rebuilt/human', 'features': ['pc1', 'pc2', 'length', 'pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr', 'mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr', 'kd_run_max', 'kd_run_n', 'kd_run_frac', 'aa_anchor', 'aa_tcr', 'aa_anchor8', 'aa_tcr8', 'aa_anchor9', 'aa_tcr9', 'aa_anchor10', 'aa_tcr10', 'aa_anchor11', 'aa_tcr11', 'aa_tcr_n', 'aa_tcr_m', 'aa_tcr_c', 'kmer_llr'], 'fits': {'em': {'immunogenic': {'mean': [-0.06616284744384984, 0.033257527759931516, 0.7425823626840402, 0.009944481875500932, -0.09744510738436167, -0.10270366444392376, 0.15899611033358801, -0.021620240473703266, 0.0453647892974406, 0.017723454003337618, 0.5022583696379227, 0.029067682576747078, -0.011969435994234565, 0.15024378096183383, 0.23894006374608343, -0.01841692959247472, 0.03140174804271794, 0.03149765470224604, 0.10795577805722431, 0.09403711588635565, 0.285996840399637, 0.25304989654414084, -0.4283587039003903, -0.5300476240404369, 0.08119225093371836, 0.1213505372650956, 0.044784546090369255, -0.0738356488326727, 0.03588980019953098, 0.026272257168881406], 'var': [1.1516692663151664, 1.1084522262605425, 0.19860447709908008, 1.0869453773015285, 1.0310605727504196, 1.1253315125231618, 1.1444420360343173, 1.0927980035842606, 1.1519899538076086, 1.1071525977980514, 0.6586179938029688, 0.8947099943206341, 0.8880246294471223, 1.1369114207966735, 1.0532738268222466, 0.8999572184564513, 1.1610496236862218, 1.3275002839607049, 0.09581536907339404, 0.10353666897503208, 0.2807555549028624, 0.34450507409073905, 4.564916481275358, 4.441787944835681, 0.10721622172558731, 0.1657772376791387, 1.0902456488673937, 1.559667039606054, 1.156856098160668, 1.4032918413480189]}, 'mixing': 0.20847508289499742, 'non_immunogenic': {'mean': [0.017426242443356672, -0.008759504226350764, -0.19558439193947458, -0.0026192184712624652, 0.025665492520388408, 0.027050512871895235, -0.04187704842323269, 0.005694427714860107, -0.011948364485953862, -0.004668076093047893, -0.13228687181118653, -0.007655965597128482, 0.003152559202088637, -0.03957182397370027, -0.06293301514573582, 0.004850726541647888, -0.00827072134405988, -0.008295981633823916, -0.028433836122047553, -0.0247678817254575, -0.07532702220934928, -0.06664932084712008, 0.11282287439673347, 0.1396061197470374, -0.021384748449546936, -0.03196180280491581, -0.011795537647336141, 0.01944713637972825, -0.009452802950278383, -0.006919694974540168], 'var': [0.9585960947283462, 0.9710673715761963, 1.0275840819740554, 0.9770670750613483, 0.9886594572583193, 0.9634797671746005, 0.9535443010666477, 0.9754029519903206, 0.9592834760347848, 0.9716731960596869, 1.0059726022242426, 1.0274505716934175, 1.029444872547767, 0.9564283707210262, 0.9669707503174307, 1.0262368262064476, 0.9572539814923111, 0.9134115157580363, 1.234269793626622, 1.2331716577819047, 1.1622201173805702, 1.1513392221513465, 1.000785366389977e-08, 1e-08, 1.232951486845345, 1.214820890534065, 0.9755633505553665, 0.850778526544332, 0.9582579947100118, 0.8935496732434686]}}, 'supervised': {'immunogenic': {'mean': [0.48458438715948515, 0.06298948453314321, -0.029654583164330322, 0.39128355525065656, 0.1014682115471637, 0.24996067426166946, -0.016188312542218187, -0.3491597579256158, -0.2717712962243248, 0.3675560892923596, 0.18135148896737566, -0.042275395210148434, -0.006615296723916411, 0.20049069236962921, -0.019114653444797746, 0.19690361209248805, 0.6315466359237181, 0.6564536683522727, 0.14271720635612214, 0.12582837414290413, 0.5589047848821895, 0.5188691769230301, 0.26888146312369027, 0.3491639056385894, 0.11709064950644237, 0.18159236648548993, 0.354008713848338, 0.3918615009934163, 0.4387206547144678, 0.7714373541380991], 'var': [1.3930230496478904, 1.1602006424759057, 0.6030177894148014, 1.2263787688544587, 0.9778518154042607, 1.0985013400029648, 1.1708403457939764, 1.1354769271622067, 1.095528758069639, 1.3172639009536289, 0.9862524597940981, 1.054503535191982, 1.0096275333834448, 1.1582998494436165, 0.9218909598402559, 1.0523029114404265, 1.5847444990351371, 2.0242241415857265, 0.6287863446403338, 0.6798051670099481, 1.7587954787066233, 2.182601004693286, 1.5692679535176246, 2.000733627789036, 0.7036790214229164, 1.0860954313051818, 1.986962383550572, 1.9224728597493752, 1.669564234269702, 2.1575359637691927]}, 'mixing': 0.03169589862138353, 'non_immunogenic': {'mean': [-0.01586210115917164, -0.00206186084840703, 0.0009706957352459654, -0.012808046440973456, -0.003321400933777193, -0.008182066129269654, 0.0005298987295771788, 0.011429190761580657, 0.008896002238425973, -0.01203136548461819, -0.005936253291110097, 0.001383817995791796, 0.00021654124361396656, -0.006562744752223986, 0.0006256878566463744, -0.006445327369643187, -0.020672677228585953, -0.021487969422078916, -0.004671621340607238, -0.004118792210880399, -0.018294861475243836, -0.016984359362011706, -0.008801408136358138, -0.011429326530386547, -0.0038327766566152344, -0.005944138035086989, -0.011587913641138271, -0.012826964578101358, -0.014360824636687601, -0.02525177796385514], 'var': [0.9791968914743827, 0.994621969619256, 1.0129648653826153, 0.9874142364597183, 1.0003769476971038, 0.9946635842540551, 0.9943989636734838, 0.9914441481583227, 0.9943762110655474, 0.9850479244412838, 0.9993382267703844, 0.998155507163318, 0.9996833889205654, 0.9934594770280398, 1.0025444345477688, 0.9969773070126338, 0.9673762074590005, 0.9519060687198755, 1.011462555639168, 1.0099458548589053, 0.9646022565242619, 0.9521883135121277, 0.9789219266960645, 0.9631212447947094, 1.0092361349801686, 0.9960670700841824, 0.9634568602941078, 0.9646134000808131, 0.9715762636565185, 0.9419920898218387]}}}, 'kd_threshold': -0.8500000000000001, 'log_odds': {'aa_anchor': [-0.17195912826031323, 1.8451940731045058, -0.3383326314979689, -0.5309006197745685, -0.10548669598002158, 0.14203178413119355, -0.4438454689097804, 0.08188954910316015, -0.12307214034493219, 0.2531804504988795, 0.5553004601198759, 0.1305669975254693, -0.035223985277909264, -0.2042807334750738, -0.3067910455939242, -0.08833926936175907, 0.060135842421909835, 0.15719944510128148, 0.1361199476438575, 0.06362187397922714], 'aa_anchor10': [-0.2022520811905535, 1.6582365394881133, -0.21191815390528346, -0.4273384044275228, -0.19388742686109328, 0.15314775491206678, -0.5778472050268788, 0.0421627706836043, 0.012321574897647736, 0.19777070487777615, 0.5231004866997857, 0.15504047994184544, -0.03311038630277352, -0.07875792042935048, -0.3743890548054498, 0.10555445424948529, 0.1198990535270954, 0.2992389719106292, -0.15750866458101953, -0.04342257958925888], 'aa_anchor11': [-0.16065601260790485, 1.6174140397801606, -0.27465433156603636, -0.14863832184342884, 0.06167147632298109, -0.07367551716753917, -0.1010235917043052, -0.028861400128323833, -0.230528091540513, 0.03736191683968837, 0.36971437324133927, -0.07392299797558932, 0.21726648668495496, -0.4227304106649723, -0.24945682828088422, 0.09779000750694689, 0.15277128278115182, 0.06643388308785925, -0.058436472170595, 0.2639780408845138], 'aa_anchor8': [-0.13629873025390715, 2.6257263950162546, -0.7974534382634779, -0.4950730959299041, -0.08115130790852332, 0.20165648504369704, -0.3868574556412123, -0.028843247661880955, -0.09865958906304018, 0.04648481655892489, 0.4944428841186088, -0.03053823305143455, 0.22466831742077442, -0.16959297760528402, 0.40538983996537015, 0.039470325723848454, 0.11370982206898317, -0.021221858621291112, 0.4424026125901088, 0.16730566685241888], 'aa_anchor9': [-0.17862695719588118, 1.9054169060709576, -0.27658399902294883, -0.5987303966114133, -0.09066545816289695, 0.17642552217525376, -0.43654794822798815, 0.10186616378673508, -0.1596994472827098, 0.2892885908046945, 0.5735795678820441, 0.12129354135612713, -0.06062207561426547, -0.23424807832612737, -0.31999365078425024, -0.18613321931811377, 0.021565292739611497, 0.12556288364885804, 0.3091340996907128, 0.06728014303186924], 'aa_tcr': [0.1269650827576445, 2.0508651169873318, -0.1278672175910449, -0.3965923638827822, 0.312048362043968, 0.05456623245180037, -0.28960607601331123, -0.08651718328776914, -0.3822461012648932, 0.07643556676200847, 0.5557280465342154, 0.13994712385397712, -0.16182712771524965, -0.39178104663270386, 0.13865507523536413, -0.09380087552214711, 0.1356944452714206, -0.06006938642577664, 0.8066684257701158, 0.3563580715841401], 'aa_tcr10': [0.11041695229878457, 1.9695372867631002, -0.27353380791584314, -0.5717224209998402, 0.47369453970111497, 0.004229081709119953, -0.31082693732559763, 0.13940711867842825, -0.33756958013949534, 0.13552090933857208, 0.5602298437823778, -0.003674095757914664, -0.2440546440131821, -0.467221239478429, 0.12067295968179748, 0.030798063215855098, 0.08170025394600078, -0.006905284398639466, 0.9126222570779312, 0.31409292051353965], 'aa_tcr11': [0.034307092582310794, 2.1276411057690137, -0.32117676334623724, -0.22987957712244578, 0.2692742510626802, -0.09214137696067315, -0.1046682419251912, 0.038192708748248094, -0.4901140111953084, -0.00674018569624657, 0.4210891741674745, -0.17233903117264138, 0.08571050978170502, -0.3319401982231489, 0.3854749883302704, -0.06271559344998678, 0.13828374712029357, 0.002301475026118993, 0.6788054088668662, 0.5037665937803593], 'aa_tcr8': [-0.29697431523976325, 2.442364901686199, 0.21768694002602063, -0.06801892407850607, 0.6132789072842479, 0.2080022184566439, -0.6029684606121299, -0.18840420706895777, 0.17071023514109074, -0.18143216168115472, 0.6777731326144099, 0.040877527639617384, -0.004906500768449895, -0.31094769397583377, 0.07652626579107435, -0.25894400383096405, -0.009474325818346951, -0.22749187598759413, 1.0455387632473006, 0.6328266631195407], 'aa_tcr9': [0.17945770071163913, 2.0619637052103297, -0.029981044514179267, -0.3654405343298275, 0.20675501336828495, 0.12656639409957382, -0.2782317593896546, -0.1917114470680965, -0.4077037365276799, 0.07032137075823908, 0.5592548998295923, 0.23170647822326806, -0.17681564687095896, -0.3842853433911029, 0.12292084043406026, -0.12364457611197066, 0.15333852825101246, -0.07875146626659957, 0.7355366135993782, 0.31904688155671534], 'aa_tcr_c': [0.16333550161931853, 1.8707098333760452, -0.1148505690154531, -0.4670323791682849, 0.3314358608017822, 0.0754790845130513, -0.3641240402505592, -0.12193367623157014, -0.43474646569214226, 0.08468515767806473, 0.5125890659921324, -0.016602970922443117, -0.0327462463587751, -0.47481451107669725, -0.20126018609099328, -0.058127301973135204, 0.1656247749868056, -0.05152218711154033, 0.797024826894055, 0.2857599761897811], 'aa_tcr_m': [0.008107412078389054, 2.177840329799901, -0.04579142184387397, -0.3610209167881262, 0.28258237212748005, -0.017691991710642174, -0.2506611603667275, -0.06943184967596983, -0.37064317866220176, 0.14042794973508954, 0.4553026314937103, 0.07890375752309797, -0.4131976989638213, -0.29321347265154074, 0.42957503508856787, -0.156967566477761, 0.1484842222490097, -0.10365523055072456, 0.797835467602722, 0.34723132268913437], 'aa_tcr_n': [0.1954528096969046, 2.2986892206804814, -0.1450903656778295, -0.3338366929044301, 0.26025426122320683, 0.14542760252377196, -0.17975415714020437, -0.0486790082582389, -0.3081890476685425, -0.10720632699825039, 0.7943039763213995, 0.42803131629758884, -0.10253553972469698, -0.3429217934402682, 0.27063702790624644, -0.07733311107678142, 0.0014274289169473597, -0.06758518949137082, 0.8294865346693872, 0.5232115791339824], 'kmer_llr': [-0.017874478710091957, 2.4432388959174407, 0.06144394763864014, -0.6114595687515676, 0.35940076483606376, 0.15473638788561495, -0.24590417844759394, 0.0811799423628008, -0.6759640849584203, 0.07704089966355099, 0.6876447031574555, 0.3596030991356125, 0.2937716689416785, -0.4560810125888013, -0.05588165663171196, 0.18203482103036883, -0.03050596545520179, -0.00279635952314905, 0.9071851400401734, 0.8170942241911412, 2.039704290339335, 3.9193938320740926, 2.1948431125394867, 1.4795343601766895, 2.914371891493655, 2.0803334022280726, 1.4968059466853498, 1.8137364482633638, 1.5286670487543343, 2.1504166725145826, 2.725471363601658, 2.170793857306818, 2.0533775922395456, 1.4943698918874695, 1.7214750513083974, 2.0980755606044923, 2.3600115901071916, 2.1948431125394867, 2.725471363601658, 2.4070176324831225, -0.06015697214592741, 1.5358872967278216, -0.020396323000508865, -0.3722189044643205, 0.0833748258927951, 0.16644368827812972, -0.6500520212923888, -0.1094628095243877, -0.42434483187710015, -0.09896229280251223, -0.02967895563648515, -0.301205713801755, -0.1342283435993199, -0.4390103892007451, 0.12164621760607641, -0.4682697866954566, -0.2576821277454737, -0.23685898356808455, 0.4260161636165396, 0.025040878014332968, -0.20391580887762562, 1.155254164320839, -0.16269973320831443, -0.5903989533073615, -0.01164859093226589, -0.3190510741217656, -0.6669814276717467, -0.5312854263356224, -0.8575558118742936, -0.29560495805036346, -0.21291865568535684, -0.4031722756891618, -0.5110056567435519, -0.5330257352102787, -0.28319065796075105, -0.4342059873718789, -0.2791864755578377, -0.5135232940989543, 0.4797944028794916, -0.11533832734649341, 0.18963524221534822, 2.2791842609732385, 0.12024810700218147, -0.4669025205210362, 0.8707370952122142, 0.27657714973996583, -0.010262266248969532, -0.1378363243030014, 0.06099241079308726, 0.35862932063901276, 0.8689172584952276, 0.19842490667132573, 0.12886468626942893, -0.6399418175702003, 0.09661196042388909, 0.13491852467867016, 0.46800670588808124, 0.26182880456709867, 0.49187914209456274, 0.2553447641253772, 0.2732980410574122, 1.691397596071118, -0.47554477282783836, -0.5523655306917297, 0.3057323636287901, -0.07565643380480847, -0.33350915729699615, 0.41660467136448087, -0.5703655024026721, 0.0015026910617672584, 0.9938760213270443, -0.09863384240278705, -0.2547437887300674, -0.5284531210850592, -0.15679708037862028, -0.12993656045995206, 0.0695155347436236, 0.2597475332304313, 0.9809829427137178, 0.08744478480777396, 0.1801509512494528, 1.9349500190165356, 0.04020119083126694, -0.4600896779808812, -0.1250544857473077, 0.026029537036047934, -0.004052055207287353, -0.25704261886648005, -0.19720133908282822, -0.406360903549948, 0.05132271417512868, -0.4163963601995846, -0.8088785510479832, -0.8793703443232737, -0.060327590374930296, -0.5826355949944864, 0.010560691468139538, -0.5770769369906583, 0.5156995702133491, 0.5719218497681, 0.13442325317276005, 1.4795343601766895, -0.09718222023955647, -0.29812980896112595, 0.03677944408736078, 0.4200697279698362, -0.6745348326903446, 0.01311497127323591, -0.15232161208764694, 0.17450543403149865, 0.0032365275074912603, -0.026570406782087552, 0.020831990687240065, -0.6000141963248016, -0.2664839449216254, -0.29103717070534874, -0.03128501393214478, -0.25070275826406174, 0.6184530541515869, 0.04645821535399275, -0.29901032323831966, 1.9775992833524647, -0.4109574318347695, -0.9157379884941479, -0.14277183178377317, -0.26833598154274885, 0.18466705688795493, -0.9052957635354284, -0.46320955932057384, -0.44721746329507095, 0.20049772121677112, -0.012773128945724466, -0.7254719311318905, -0.7016580441538798, 0.16560676285823117, -0.5086074567861996, -0.5634689233436116, -0.7159228171868346, 0.7474728999948432, -0.16579605042245227, 0.2088597839639048, 2.2608585829907986, -0.08595264814181114, -0.4283145777283499, 0.3384104329689972, 0.23319170397478928, -0.27702095501005264, -0.19597809734008464, -0.4126038907299048, 0.28434529037365763, 0.5224664233499388, 0.19306555387259206, -0.22142578074321229, -0.2131158045146444, -0.1893219674469817, 0.13683051990041495, 0.16909147726837315, 0.14037577558778302, 0.8627115828785152, 0.3932640145762685, 0.5117884541166067, 2.812482740591287, 0.16984918901525958, 0.9683320950202061, 0.3275760908032872, 0.5782072255830606, -0.018758494078054966, 0.4157688028387474, -0.21633256792677802, 0.9722708879772854, 0.9950808407498943, 0.2135519145675886, 0.2977231276536063, -0.2733810372885772, 0.21881228201375258, 0.6338097105870455, 0.4633491342207927, 0.4673915105441626, 2.7705917988821263, 0.9515467620096363, 0.25985995105695014, 2.2926072813053793, 0.21057035794546586, -0.6234391816140699, 0.015704423880392504, 0.2957573725329903, -0.08443321382909108, 0.02530556602359635, -0.2662674394063833, 0.29230182296704044, 0.698698767504009, 0.4002472811006461, -0.20913837658789536, -0.011987105344703153, 0.14851919478157338, -0.042834696846123066, 0.9828618354852141, 0.004139788340288497, 0.06787938376957658, 0.6563926089573675, -0.39889212823543296, 1.689772900344117, -0.2837869060121019, -0.4532860480037648, 0.02121212290199459, -0.6092240560466005, -0.44553778359564067, -0.1557006973960835, -0.8005120159977759, -0.37802716736524555, 0.016846614939833415, -0.63370302828565, -0.58476172709838, -0.23234167509003978, 0.7829774258778581, -0.44081255909242856, -0.06577411175868697, -0.2518356231607708, 0.4401994618092884, 0.10019908895714114, -0.3626984770200412, 1.9231248910767196, -0.42532026771580966, -0.8856544276422138, -0.1523018024581857, -0.2512644092616787, -0.41089203969207944, -0.20638883335792269, -0.9792699948127455, -0.08423249122705201, 0.9711088747135914, -0.17869371642684317, -0.6315087157361994, -0.5934173705977743, -0.8132119992825038, -0.2105109055465597, -0.19073890002552663, -0.4876487987391185, 0.2831243282324536, 0.43398661312012443, 0.3882139879839652, 1.9015539518585172, -0.13935763973408122, -0.2543328423178828, 0.32200297587685434, 0.42514445749120355, 0.05296609991885948, -0.46637578887862396, 0.4897516558162911, 0.2261407326217224, 0.36637406034426956, 0.3375066590058875, -0.4893184549980303, -0.25521777163119364, 1.3221079424272144, 0.22204076525077543, 1.2314716770646736, -0.17816670719198413, 1.2579515625591808, 0.2230423631214551, 0.1315835285550282, 2.4293788126412803, -0.21679488249928092, -0.7731173640686473, 0.41838995642723287, -0.08789326914721851, -0.30670971966036564, -0.29439267003221037, -0.43731499461549994, 0.025641995644305027, 0.49776376954036206, 0.010010671608962518, -0.06112280519738267, -0.6232079837770517, -0.060507698697708, -0.23712292526221024, -0.1335498644104316, -0.27048212442066166, 0.7524592847685527, 0.30650561923395525, 0.09688533385369436, 1.4676193255988172, 0.033254665706303754, -0.1744037055573351, 1.0827025320703392, -0.22264505601066986, -0.31770879204103686, 0.06972534701581701, -0.21007189036285556, 0.4462987112323553, 0.9075421766402174, -0.02262633862982799, -0.02919820239785853, -0.26608553854185324, 0.2453040591775535, 0.07922493041392897, 0.0502389381938988, 0.3183243931040858, 0.9321978895423877, 0.4479342094767844, 0.4037122973838114, 1.835613888795657, -0.11777550944413662, -0.46008228388285133, 0.31984551668843153, -0.26721573757716666, -0.5830871642051401, -0.2016738002887628, -0.5587239273871853, 0.1299446165366298, 0.3357036753999454, -0.04894051319963921, -0.1246659262448313, -0.6427509929034283, -0.09130946412794838, -0.1943142795008006, -0.07826616232160521, -0.009635638888066289, 0.8418035234235655, 0.2337769057593837, 0.6298569626762198, 2.84807368569399, 0.8536691867000652, 0.4567878222832933, 0.9236005847692823, 0.5481569462247675, 0.45273923375729197, 0.9325221262380134, -0.04853680529235049, 1.1551208398756039, 1.1464926586522655, 0.9857682563296555, 1.1278679088145598, 0.5417474734843406, 0.9153627557054067, 0.7614537239189536, 0.5693581745078689, 0.19064140504705662, 0.9337118943736034, 1.802644660747717, 0.2647519394276605, 2.5658412180097745, 0.24630008389840174, -0.0675636803470443, 0.6704483492699671, 0.47248321174698393, -0.07961912154266049, -0.007189324685382914, 0.22771657133580803, 0.3860722974848949, 0.5730046869069314, 0.6245516126134572, 0.15336257357971128, 0.09052081593905204, 0.7479947485785097, 0.20758951699810702, 0.7015694071908234, 0.2812680069271609, 1.2549849112614186, 0.5233853279791925]}, 'log_odds_source': {'aa_anchor': 'anchor', 'aa_anchor10': 'anchor@10', 'aa_anchor11': 'anchor@11', 'aa_anchor8': 'anchor@8', 'aa_anchor9': 'anchor@9', 'aa_tcr': 'tcr', 'aa_tcr10': 'tcr@10', 'aa_tcr11': 'tcr@11', 'aa_tcr8': 'tcr@8', 'aa_tcr9': 'tcr@9', 'aa_tcr_c': 'tcr_c', 'aa_tcr_m': 'tcr_m', 'aa_tcr_n': 'tcr_n', 'kmer_llr': 'pair'}, 'logistic': {'coef': [0.11949579886625111, -0.026388547388491116, 0.03688791951418523, 0.12315135708771073, 0.00793113416429338, 0.032785084028538934, -0.0491770492059215, 0.1283027809168922, 0.0010358114328155197, 0.03544359835457312, -0.09272877328007892, -0.029929037150132728, 0.005234358233192701, 0.054646426854836824, 0.05481812119007732, 0.004802849974804372, 0.031791737568215295, -0.8795515088055536, 0.13246571365233376, 0.11209166951408028, 0.41632421759666804, 0.3710860928169145, 0.21945608459259616, 0.2831765986455691, 0.09310214697766131, 0.12584078767223586, 0.2711009144487395, 0.10677716086875681, 0.2388095572730888, 0.6543935078687705], 'intercept': -3.8119421736439913, 'tau': 4.0}, 'n': 464161, 'n_immunogenic': 14712, 'paratope': {'A': [0.0468, 0.3668], 'C': [0.1994, 0.5024], 'D': [0.1813, 0.5419], 'E': [0.1014, 0.313], 'F': [0.0754, 0.3616], 'G': [-0.0, 0.2369], 'H': [0.2189, 0.5621], 'I': [0.0486, 0.416], 'K': [0.1234, 0.4516], 'L': [-0.0251, 0.1864], 'M': [0.0345, 0.3921], 'N': [0.1193, 0.3291], 'P': [0.0469, 0.2501], 'Q': [0.0384, 0.3411], 'R': [0.0515, 0.421], 'S': [0.0124, 0.2607], 'T': [0.1307, 0.3778], 'V': [0.0143, 0.2991], 'W': [0.0847, 0.4427], 'Y': [0.0172, 0.2474]}, 'prevalence': 0.03169589862138353, 'seed': 20260817, 'standardizer': {'mean': [3.585166758896088, 6.808608481871727, 9.314074211318918, 4.7808659475268716, 1.3706172407505959, -1.1956991886307702, 5.437991241121321, -0.17470653501694436, 0.36020443337546376, 13.744965561518644, 10.890707254593467, 0.06078909482557044, 0.3367987089608995, 1.993329900616381, 1.320692604505764, 0.5490802401177581, -0.1541268070176281, -0.14749849482590252, -0.02042398402213005, -0.014461619053264312, -0.10887835324113007, -0.08397399955209743, -0.02768170726510107, -0.041256419055064476, -0.011095587661907503, -0.023404604320161237, -0.042033855105112226, -0.05205149786415519, -0.07093132420268783, -0.3009136507757868], 'std': [18.338465868430596, 15.272311296957708, 0.7785394170589226, 14.28094041797769, 11.073130188216279, 13.196624288892432, 9.981009433871744, 2.103026551096585, 1.8789707926397332, 2.040641126152638, 2.5470821127038956, 0.029406827409666877, 0.04828902479536415, 1.064927514941948, 0.5583973810425537, 0.23698125064685832, 0.5361092115927888, 0.5265313599296632, 0.20078005110769423, 0.16370954872505514, 0.4592756393360549, 0.40885275689301376, 0.24535545130466027, 0.295520132855347, 0.14842445509414, 0.2117117554672323, 0.27496149723350805, 0.31126309236162925, 0.37329783588644283, 0.9119874447264713]}}#

The human class-I table, for callers that predate the species split.

mhcmatch.complement.TABLES: dict = {'human': {'alphabet': 'ACDEFGHIKLMNPQRSTVWY', 'anchors': [0, 1, 2, -2, -1], 'arm': 'chowell_rebuilt/human', 'features': ['pc1', 'pc2', 'length', 'pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr', 'mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr', 'kd_run_max', 'kd_run_n', 'kd_run_frac', 'aa_anchor', 'aa_tcr', 'aa_anchor8', 'aa_tcr8', 'aa_anchor9', 'aa_tcr9', 'aa_anchor10', 'aa_tcr10', 'aa_anchor11', 'aa_tcr11', 'aa_tcr_n', 'aa_tcr_m', 'aa_tcr_c', 'kmer_llr'], 'fits': {'em': {'immunogenic': {'mean': [-0.06616284744384984, 0.033257527759931516, 0.7425823626840402, 0.009944481875500932, -0.09744510738436167, -0.10270366444392376, 0.15899611033358801, -0.021620240473703266, 0.0453647892974406, 0.017723454003337618, 0.5022583696379227, 0.029067682576747078, -0.011969435994234565, 0.15024378096183383, 0.23894006374608343, -0.01841692959247472, 0.03140174804271794, 0.03149765470224604, 0.10795577805722431, 0.09403711588635565, 0.285996840399637, 0.25304989654414084, -0.4283587039003903, -0.5300476240404369, 0.08119225093371836, 0.1213505372650956, 0.044784546090369255, -0.0738356488326727, 0.03588980019953098, 0.026272257168881406], 'var': [1.1516692663151664, 1.1084522262605425, 0.19860447709908008, 1.0869453773015285, 1.0310605727504196, 1.1253315125231618, 1.1444420360343173, 1.0927980035842606, 1.1519899538076086, 1.1071525977980514, 0.6586179938029688, 0.8947099943206341, 0.8880246294471223, 1.1369114207966735, 1.0532738268222466, 0.8999572184564513, 1.1610496236862218, 1.3275002839607049, 0.09581536907339404, 0.10353666897503208, 0.2807555549028624, 0.34450507409073905, 4.564916481275358, 4.441787944835681, 0.10721622172558731, 0.1657772376791387, 1.0902456488673937, 1.559667039606054, 1.156856098160668, 1.4032918413480189]}, 'mixing': 0.20847508289499742, 'non_immunogenic': {'mean': [0.017426242443356672, -0.008759504226350764, -0.19558439193947458, -0.0026192184712624652, 0.025665492520388408, 0.027050512871895235, -0.04187704842323269, 0.005694427714860107, -0.011948364485953862, -0.004668076093047893, -0.13228687181118653, -0.007655965597128482, 0.003152559202088637, -0.03957182397370027, -0.06293301514573582, 0.004850726541647888, -0.00827072134405988, -0.008295981633823916, -0.028433836122047553, -0.0247678817254575, -0.07532702220934928, -0.06664932084712008, 0.11282287439673347, 0.1396061197470374, -0.021384748449546936, -0.03196180280491581, -0.011795537647336141, 0.01944713637972825, -0.009452802950278383, -0.006919694974540168], 'var': [0.9585960947283462, 0.9710673715761963, 1.0275840819740554, 0.9770670750613483, 0.9886594572583193, 0.9634797671746005, 0.9535443010666477, 0.9754029519903206, 0.9592834760347848, 0.9716731960596869, 1.0059726022242426, 1.0274505716934175, 1.029444872547767, 0.9564283707210262, 0.9669707503174307, 1.0262368262064476, 0.9572539814923111, 0.9134115157580363, 1.234269793626622, 1.2331716577819047, 1.1622201173805702, 1.1513392221513465, 1.000785366389977e-08, 1e-08, 1.232951486845345, 1.214820890534065, 0.9755633505553665, 0.850778526544332, 0.9582579947100118, 0.8935496732434686]}}, 'supervised': {'immunogenic': {'mean': [0.48458438715948515, 0.06298948453314321, -0.029654583164330322, 0.39128355525065656, 0.1014682115471637, 0.24996067426166946, -0.016188312542218187, -0.3491597579256158, -0.2717712962243248, 0.3675560892923596, 0.18135148896737566, -0.042275395210148434, -0.006615296723916411, 0.20049069236962921, -0.019114653444797746, 0.19690361209248805, 0.6315466359237181, 0.6564536683522727, 0.14271720635612214, 0.12582837414290413, 0.5589047848821895, 0.5188691769230301, 0.26888146312369027, 0.3491639056385894, 0.11709064950644237, 0.18159236648548993, 0.354008713848338, 0.3918615009934163, 0.4387206547144678, 0.7714373541380991], 'var': [1.3930230496478904, 1.1602006424759057, 0.6030177894148014, 1.2263787688544587, 0.9778518154042607, 1.0985013400029648, 1.1708403457939764, 1.1354769271622067, 1.095528758069639, 1.3172639009536289, 0.9862524597940981, 1.054503535191982, 1.0096275333834448, 1.1582998494436165, 0.9218909598402559, 1.0523029114404265, 1.5847444990351371, 2.0242241415857265, 0.6287863446403338, 0.6798051670099481, 1.7587954787066233, 2.182601004693286, 1.5692679535176246, 2.000733627789036, 0.7036790214229164, 1.0860954313051818, 1.986962383550572, 1.9224728597493752, 1.669564234269702, 2.1575359637691927]}, 'mixing': 0.03169589862138353, 'non_immunogenic': {'mean': [-0.01586210115917164, -0.00206186084840703, 0.0009706957352459654, -0.012808046440973456, -0.003321400933777193, -0.008182066129269654, 0.0005298987295771788, 0.011429190761580657, 0.008896002238425973, -0.01203136548461819, -0.005936253291110097, 0.001383817995791796, 0.00021654124361396656, -0.006562744752223986, 0.0006256878566463744, -0.006445327369643187, -0.020672677228585953, -0.021487969422078916, -0.004671621340607238, -0.004118792210880399, -0.018294861475243836, -0.016984359362011706, -0.008801408136358138, -0.011429326530386547, -0.0038327766566152344, -0.005944138035086989, -0.011587913641138271, -0.012826964578101358, -0.014360824636687601, -0.02525177796385514], 'var': [0.9791968914743827, 0.994621969619256, 1.0129648653826153, 0.9874142364597183, 1.0003769476971038, 0.9946635842540551, 0.9943989636734838, 0.9914441481583227, 0.9943762110655474, 0.9850479244412838, 0.9993382267703844, 0.998155507163318, 0.9996833889205654, 0.9934594770280398, 1.0025444345477688, 0.9969773070126338, 0.9673762074590005, 0.9519060687198755, 1.011462555639168, 1.0099458548589053, 0.9646022565242619, 0.9521883135121277, 0.9789219266960645, 0.9631212447947094, 1.0092361349801686, 0.9960670700841824, 0.9634568602941078, 0.9646134000808131, 0.9715762636565185, 0.9419920898218387]}}}, 'kd_threshold': -0.8500000000000001, 'log_odds': {'aa_anchor': [-0.17195912826031323, 1.8451940731045058, -0.3383326314979689, -0.5309006197745685, -0.10548669598002158, 0.14203178413119355, -0.4438454689097804, 0.08188954910316015, -0.12307214034493219, 0.2531804504988795, 0.5553004601198759, 0.1305669975254693, -0.035223985277909264, -0.2042807334750738, -0.3067910455939242, -0.08833926936175907, 0.060135842421909835, 0.15719944510128148, 0.1361199476438575, 0.06362187397922714], 'aa_anchor10': [-0.2022520811905535, 1.6582365394881133, -0.21191815390528346, -0.4273384044275228, -0.19388742686109328, 0.15314775491206678, -0.5778472050268788, 0.0421627706836043, 0.012321574897647736, 0.19777070487777615, 0.5231004866997857, 0.15504047994184544, -0.03311038630277352, -0.07875792042935048, -0.3743890548054498, 0.10555445424948529, 0.1198990535270954, 0.2992389719106292, -0.15750866458101953, -0.04342257958925888], 'aa_anchor11': [-0.16065601260790485, 1.6174140397801606, -0.27465433156603636, -0.14863832184342884, 0.06167147632298109, -0.07367551716753917, -0.1010235917043052, -0.028861400128323833, -0.230528091540513, 0.03736191683968837, 0.36971437324133927, -0.07392299797558932, 0.21726648668495496, -0.4227304106649723, -0.24945682828088422, 0.09779000750694689, 0.15277128278115182, 0.06643388308785925, -0.058436472170595, 0.2639780408845138], 'aa_anchor8': [-0.13629873025390715, 2.6257263950162546, -0.7974534382634779, -0.4950730959299041, -0.08115130790852332, 0.20165648504369704, -0.3868574556412123, -0.028843247661880955, -0.09865958906304018, 0.04648481655892489, 0.4944428841186088, -0.03053823305143455, 0.22466831742077442, -0.16959297760528402, 0.40538983996537015, 0.039470325723848454, 0.11370982206898317, -0.021221858621291112, 0.4424026125901088, 0.16730566685241888], 'aa_anchor9': [-0.17862695719588118, 1.9054169060709576, -0.27658399902294883, -0.5987303966114133, -0.09066545816289695, 0.17642552217525376, -0.43654794822798815, 0.10186616378673508, -0.1596994472827098, 0.2892885908046945, 0.5735795678820441, 0.12129354135612713, -0.06062207561426547, -0.23424807832612737, -0.31999365078425024, -0.18613321931811377, 0.021565292739611497, 0.12556288364885804, 0.3091340996907128, 0.06728014303186924], 'aa_tcr': [0.1269650827576445, 2.0508651169873318, -0.1278672175910449, -0.3965923638827822, 0.312048362043968, 0.05456623245180037, -0.28960607601331123, -0.08651718328776914, -0.3822461012648932, 0.07643556676200847, 0.5557280465342154, 0.13994712385397712, -0.16182712771524965, -0.39178104663270386, 0.13865507523536413, -0.09380087552214711, 0.1356944452714206, -0.06006938642577664, 0.8066684257701158, 0.3563580715841401], 'aa_tcr10': [0.11041695229878457, 1.9695372867631002, -0.27353380791584314, -0.5717224209998402, 0.47369453970111497, 0.004229081709119953, -0.31082693732559763, 0.13940711867842825, -0.33756958013949534, 0.13552090933857208, 0.5602298437823778, -0.003674095757914664, -0.2440546440131821, -0.467221239478429, 0.12067295968179748, 0.030798063215855098, 0.08170025394600078, -0.006905284398639466, 0.9126222570779312, 0.31409292051353965], 'aa_tcr11': [0.034307092582310794, 2.1276411057690137, -0.32117676334623724, -0.22987957712244578, 0.2692742510626802, -0.09214137696067315, -0.1046682419251912, 0.038192708748248094, -0.4901140111953084, -0.00674018569624657, 0.4210891741674745, -0.17233903117264138, 0.08571050978170502, -0.3319401982231489, 0.3854749883302704, -0.06271559344998678, 0.13828374712029357, 0.002301475026118993, 0.6788054088668662, 0.5037665937803593], 'aa_tcr8': [-0.29697431523976325, 2.442364901686199, 0.21768694002602063, -0.06801892407850607, 0.6132789072842479, 0.2080022184566439, -0.6029684606121299, -0.18840420706895777, 0.17071023514109074, -0.18143216168115472, 0.6777731326144099, 0.040877527639617384, -0.004906500768449895, -0.31094769397583377, 0.07652626579107435, -0.25894400383096405, -0.009474325818346951, -0.22749187598759413, 1.0455387632473006, 0.6328266631195407], 'aa_tcr9': [0.17945770071163913, 2.0619637052103297, -0.029981044514179267, -0.3654405343298275, 0.20675501336828495, 0.12656639409957382, -0.2782317593896546, -0.1917114470680965, -0.4077037365276799, 0.07032137075823908, 0.5592548998295923, 0.23170647822326806, -0.17681564687095896, -0.3842853433911029, 0.12292084043406026, -0.12364457611197066, 0.15333852825101246, -0.07875146626659957, 0.7355366135993782, 0.31904688155671534], 'aa_tcr_c': [0.16333550161931853, 1.8707098333760452, -0.1148505690154531, -0.4670323791682849, 0.3314358608017822, 0.0754790845130513, -0.3641240402505592, -0.12193367623157014, -0.43474646569214226, 0.08468515767806473, 0.5125890659921324, -0.016602970922443117, -0.0327462463587751, -0.47481451107669725, -0.20126018609099328, -0.058127301973135204, 0.1656247749868056, -0.05152218711154033, 0.797024826894055, 0.2857599761897811], 'aa_tcr_m': [0.008107412078389054, 2.177840329799901, -0.04579142184387397, -0.3610209167881262, 0.28258237212748005, -0.017691991710642174, -0.2506611603667275, -0.06943184967596983, -0.37064317866220176, 0.14042794973508954, 0.4553026314937103, 0.07890375752309797, -0.4131976989638213, -0.29321347265154074, 0.42957503508856787, -0.156967566477761, 0.1484842222490097, -0.10365523055072456, 0.797835467602722, 0.34723132268913437], 'aa_tcr_n': [0.1954528096969046, 2.2986892206804814, -0.1450903656778295, -0.3338366929044301, 0.26025426122320683, 0.14542760252377196, -0.17975415714020437, -0.0486790082582389, -0.3081890476685425, -0.10720632699825039, 0.7943039763213995, 0.42803131629758884, -0.10253553972469698, -0.3429217934402682, 0.27063702790624644, -0.07733311107678142, 0.0014274289169473597, -0.06758518949137082, 0.8294865346693872, 0.5232115791339824], 'kmer_llr': [-0.017874478710091957, 2.4432388959174407, 0.06144394763864014, -0.6114595687515676, 0.35940076483606376, 0.15473638788561495, -0.24590417844759394, 0.0811799423628008, -0.6759640849584203, 0.07704089966355099, 0.6876447031574555, 0.3596030991356125, 0.2937716689416785, -0.4560810125888013, -0.05588165663171196, 0.18203482103036883, -0.03050596545520179, -0.00279635952314905, 0.9071851400401734, 0.8170942241911412, 2.039704290339335, 3.9193938320740926, 2.1948431125394867, 1.4795343601766895, 2.914371891493655, 2.0803334022280726, 1.4968059466853498, 1.8137364482633638, 1.5286670487543343, 2.1504166725145826, 2.725471363601658, 2.170793857306818, 2.0533775922395456, 1.4943698918874695, 1.7214750513083974, 2.0980755606044923, 2.3600115901071916, 2.1948431125394867, 2.725471363601658, 2.4070176324831225, -0.06015697214592741, 1.5358872967278216, -0.020396323000508865, -0.3722189044643205, 0.0833748258927951, 0.16644368827812972, -0.6500520212923888, -0.1094628095243877, -0.42434483187710015, -0.09896229280251223, -0.02967895563648515, -0.301205713801755, -0.1342283435993199, -0.4390103892007451, 0.12164621760607641, -0.4682697866954566, -0.2576821277454737, -0.23685898356808455, 0.4260161636165396, 0.025040878014332968, -0.20391580887762562, 1.155254164320839, -0.16269973320831443, -0.5903989533073615, -0.01164859093226589, -0.3190510741217656, -0.6669814276717467, -0.5312854263356224, -0.8575558118742936, -0.29560495805036346, -0.21291865568535684, -0.4031722756891618, -0.5110056567435519, -0.5330257352102787, -0.28319065796075105, -0.4342059873718789, -0.2791864755578377, -0.5135232940989543, 0.4797944028794916, -0.11533832734649341, 0.18963524221534822, 2.2791842609732385, 0.12024810700218147, -0.4669025205210362, 0.8707370952122142, 0.27657714973996583, -0.010262266248969532, -0.1378363243030014, 0.06099241079308726, 0.35862932063901276, 0.8689172584952276, 0.19842490667132573, 0.12886468626942893, -0.6399418175702003, 0.09661196042388909, 0.13491852467867016, 0.46800670588808124, 0.26182880456709867, 0.49187914209456274, 0.2553447641253772, 0.2732980410574122, 1.691397596071118, -0.47554477282783836, -0.5523655306917297, 0.3057323636287901, -0.07565643380480847, -0.33350915729699615, 0.41660467136448087, -0.5703655024026721, 0.0015026910617672584, 0.9938760213270443, -0.09863384240278705, -0.2547437887300674, -0.5284531210850592, -0.15679708037862028, -0.12993656045995206, 0.0695155347436236, 0.2597475332304313, 0.9809829427137178, 0.08744478480777396, 0.1801509512494528, 1.9349500190165356, 0.04020119083126694, -0.4600896779808812, -0.1250544857473077, 0.026029537036047934, -0.004052055207287353, -0.25704261886648005, -0.19720133908282822, -0.406360903549948, 0.05132271417512868, -0.4163963601995846, -0.8088785510479832, -0.8793703443232737, -0.060327590374930296, -0.5826355949944864, 0.010560691468139538, -0.5770769369906583, 0.5156995702133491, 0.5719218497681, 0.13442325317276005, 1.4795343601766895, -0.09718222023955647, -0.29812980896112595, 0.03677944408736078, 0.4200697279698362, -0.6745348326903446, 0.01311497127323591, -0.15232161208764694, 0.17450543403149865, 0.0032365275074912603, -0.026570406782087552, 0.020831990687240065, -0.6000141963248016, -0.2664839449216254, -0.29103717070534874, -0.03128501393214478, -0.25070275826406174, 0.6184530541515869, 0.04645821535399275, -0.29901032323831966, 1.9775992833524647, -0.4109574318347695, -0.9157379884941479, -0.14277183178377317, -0.26833598154274885, 0.18466705688795493, -0.9052957635354284, -0.46320955932057384, -0.44721746329507095, 0.20049772121677112, -0.012773128945724466, -0.7254719311318905, -0.7016580441538798, 0.16560676285823117, -0.5086074567861996, -0.5634689233436116, -0.7159228171868346, 0.7474728999948432, -0.16579605042245227, 0.2088597839639048, 2.2608585829907986, -0.08595264814181114, -0.4283145777283499, 0.3384104329689972, 0.23319170397478928, -0.27702095501005264, -0.19597809734008464, -0.4126038907299048, 0.28434529037365763, 0.5224664233499388, 0.19306555387259206, -0.22142578074321229, -0.2131158045146444, -0.1893219674469817, 0.13683051990041495, 0.16909147726837315, 0.14037577558778302, 0.8627115828785152, 0.3932640145762685, 0.5117884541166067, 2.812482740591287, 0.16984918901525958, 0.9683320950202061, 0.3275760908032872, 0.5782072255830606, -0.018758494078054966, 0.4157688028387474, -0.21633256792677802, 0.9722708879772854, 0.9950808407498943, 0.2135519145675886, 0.2977231276536063, -0.2733810372885772, 0.21881228201375258, 0.6338097105870455, 0.4633491342207927, 0.4673915105441626, 2.7705917988821263, 0.9515467620096363, 0.25985995105695014, 2.2926072813053793, 0.21057035794546586, -0.6234391816140699, 0.015704423880392504, 0.2957573725329903, -0.08443321382909108, 0.02530556602359635, -0.2662674394063833, 0.29230182296704044, 0.698698767504009, 0.4002472811006461, -0.20913837658789536, -0.011987105344703153, 0.14851919478157338, -0.042834696846123066, 0.9828618354852141, 0.004139788340288497, 0.06787938376957658, 0.6563926089573675, -0.39889212823543296, 1.689772900344117, -0.2837869060121019, -0.4532860480037648, 0.02121212290199459, -0.6092240560466005, -0.44553778359564067, -0.1557006973960835, -0.8005120159977759, -0.37802716736524555, 0.016846614939833415, -0.63370302828565, -0.58476172709838, -0.23234167509003978, 0.7829774258778581, -0.44081255909242856, -0.06577411175868697, -0.2518356231607708, 0.4401994618092884, 0.10019908895714114, -0.3626984770200412, 1.9231248910767196, -0.42532026771580966, -0.8856544276422138, -0.1523018024581857, -0.2512644092616787, -0.41089203969207944, -0.20638883335792269, -0.9792699948127455, -0.08423249122705201, 0.9711088747135914, -0.17869371642684317, -0.6315087157361994, -0.5934173705977743, -0.8132119992825038, -0.2105109055465597, -0.19073890002552663, -0.4876487987391185, 0.2831243282324536, 0.43398661312012443, 0.3882139879839652, 1.9015539518585172, -0.13935763973408122, -0.2543328423178828, 0.32200297587685434, 0.42514445749120355, 0.05296609991885948, -0.46637578887862396, 0.4897516558162911, 0.2261407326217224, 0.36637406034426956, 0.3375066590058875, -0.4893184549980303, -0.25521777163119364, 1.3221079424272144, 0.22204076525077543, 1.2314716770646736, -0.17816670719198413, 1.2579515625591808, 0.2230423631214551, 0.1315835285550282, 2.4293788126412803, -0.21679488249928092, -0.7731173640686473, 0.41838995642723287, -0.08789326914721851, -0.30670971966036564, -0.29439267003221037, -0.43731499461549994, 0.025641995644305027, 0.49776376954036206, 0.010010671608962518, -0.06112280519738267, -0.6232079837770517, -0.060507698697708, -0.23712292526221024, -0.1335498644104316, -0.27048212442066166, 0.7524592847685527, 0.30650561923395525, 0.09688533385369436, 1.4676193255988172, 0.033254665706303754, -0.1744037055573351, 1.0827025320703392, -0.22264505601066986, -0.31770879204103686, 0.06972534701581701, -0.21007189036285556, 0.4462987112323553, 0.9075421766402174, -0.02262633862982799, -0.02919820239785853, -0.26608553854185324, 0.2453040591775535, 0.07922493041392897, 0.0502389381938988, 0.3183243931040858, 0.9321978895423877, 0.4479342094767844, 0.4037122973838114, 1.835613888795657, -0.11777550944413662, -0.46008228388285133, 0.31984551668843153, -0.26721573757716666, -0.5830871642051401, -0.2016738002887628, -0.5587239273871853, 0.1299446165366298, 0.3357036753999454, -0.04894051319963921, -0.1246659262448313, -0.6427509929034283, -0.09130946412794838, -0.1943142795008006, -0.07826616232160521, -0.009635638888066289, 0.8418035234235655, 0.2337769057593837, 0.6298569626762198, 2.84807368569399, 0.8536691867000652, 0.4567878222832933, 0.9236005847692823, 0.5481569462247675, 0.45273923375729197, 0.9325221262380134, -0.04853680529235049, 1.1551208398756039, 1.1464926586522655, 0.9857682563296555, 1.1278679088145598, 0.5417474734843406, 0.9153627557054067, 0.7614537239189536, 0.5693581745078689, 0.19064140504705662, 0.9337118943736034, 1.802644660747717, 0.2647519394276605, 2.5658412180097745, 0.24630008389840174, -0.0675636803470443, 0.6704483492699671, 0.47248321174698393, -0.07961912154266049, -0.007189324685382914, 0.22771657133580803, 0.3860722974848949, 0.5730046869069314, 0.6245516126134572, 0.15336257357971128, 0.09052081593905204, 0.7479947485785097, 0.20758951699810702, 0.7015694071908234, 0.2812680069271609, 1.2549849112614186, 0.5233853279791925]}, 'log_odds_source': {'aa_anchor': 'anchor', 'aa_anchor10': 'anchor@10', 'aa_anchor11': 'anchor@11', 'aa_anchor8': 'anchor@8', 'aa_anchor9': 'anchor@9', 'aa_tcr': 'tcr', 'aa_tcr10': 'tcr@10', 'aa_tcr11': 'tcr@11', 'aa_tcr8': 'tcr@8', 'aa_tcr9': 'tcr@9', 'aa_tcr_c': 'tcr_c', 'aa_tcr_m': 'tcr_m', 'aa_tcr_n': 'tcr_n', 'kmer_llr': 'pair'}, 'logistic': {'coef': [0.11949579886625111, -0.026388547388491116, 0.03688791951418523, 0.12315135708771073, 0.00793113416429338, 0.032785084028538934, -0.0491770492059215, 0.1283027809168922, 0.0010358114328155197, 0.03544359835457312, -0.09272877328007892, -0.029929037150132728, 0.005234358233192701, 0.054646426854836824, 0.05481812119007732, 0.004802849974804372, 0.031791737568215295, -0.8795515088055536, 0.13246571365233376, 0.11209166951408028, 0.41632421759666804, 0.3710860928169145, 0.21945608459259616, 0.2831765986455691, 0.09310214697766131, 0.12584078767223586, 0.2711009144487395, 0.10677716086875681, 0.2388095572730888, 0.6543935078687705], 'intercept': -3.8119421736439913, 'tau': 4.0}, 'n': 464161, 'n_immunogenic': 14712, 'paratope': {'A': [0.0468, 0.3668], 'C': [0.1994, 0.5024], 'D': [0.1813, 0.5419], 'E': [0.1014, 0.313], 'F': [0.0754, 0.3616], 'G': [-0.0, 0.2369], 'H': [0.2189, 0.5621], 'I': [0.0486, 0.416], 'K': [0.1234, 0.4516], 'L': [-0.0251, 0.1864], 'M': [0.0345, 0.3921], 'N': [0.1193, 0.3291], 'P': [0.0469, 0.2501], 'Q': [0.0384, 0.3411], 'R': [0.0515, 0.421], 'S': [0.0124, 0.2607], 'T': [0.1307, 0.3778], 'V': [0.0143, 0.2991], 'W': [0.0847, 0.4427], 'Y': [0.0172, 0.2474]}, 'prevalence': 0.03169589862138353, 'seed': 20260817, 'standardizer': {'mean': [3.585166758896088, 6.808608481871727, 9.314074211318918, 4.7808659475268716, 1.3706172407505959, -1.1956991886307702, 5.437991241121321, -0.17470653501694436, 0.36020443337546376, 13.744965561518644, 10.890707254593467, 0.06078909482557044, 0.3367987089608995, 1.993329900616381, 1.320692604505764, 0.5490802401177581, -0.1541268070176281, -0.14749849482590252, -0.02042398402213005, -0.014461619053264312, -0.10887835324113007, -0.08397399955209743, -0.02768170726510107, -0.041256419055064476, -0.011095587661907503, -0.023404604320161237, -0.042033855105112226, -0.05205149786415519, -0.07093132420268783, -0.3009136507757868], 'std': [18.338465868430596, 15.272311296957708, 0.7785394170589226, 14.28094041797769, 11.073130188216279, 13.196624288892432, 9.981009433871744, 2.103026551096585, 1.8789707926397332, 2.040641126152638, 2.5470821127038956, 0.029406827409666877, 0.04828902479536415, 1.064927514941948, 0.5583973810425537, 0.23698125064685832, 0.5361092115927888, 0.5265313599296632, 0.20078005110769423, 0.16370954872505514, 0.4592756393360549, 0.40885275689301376, 0.24535545130466027, 0.295520132855347, 0.14842445509414, 0.2117117554672323, 0.27496149723350805, 0.31126309236162925, 0.37329783588644283, 0.9119874447264713]}}, 'mouse': {'alphabet': 'ACDEFGHIKLMNPQRSTVWY', 'anchors': [0, 1, 2, -2, -1], 'arm': 'chowell_rebuilt/mouse', 'features': ['pc1', 'pc2', 'length', 'pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr', 'mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr', 'kd_run_max', 'kd_run_n', 'kd_run_frac', 'aa_anchor', 'aa_tcr', 'aa_anchor8', 'aa_tcr8', 'aa_anchor9', 'aa_tcr9', 'aa_anchor10', 'aa_tcr10', 'aa_anchor11', 'aa_tcr11', 'aa_tcr_n', 'aa_tcr_m', 'aa_tcr_c', 'kmer_llr'], 'fits': {'em': {'immunogenic': {'mean': [0.14312370634486676, 0.16557843346884654, 0.5128065164897488, 0.12339035765818519, 0.07706062311119909, 0.06952353629098441, 0.1574738613288356, -0.16171673327480834, -0.1194230737471546, 0.08672479156214237, 0.48834034342956, -0.043463111475816586, -0.056797325534043576, 0.22450608570720082, 0.21658895480452678, 0.08411004076461961, 0.23662124927264536, 0.2148995074580367, 0.24544731929071378, 0.2476361813953243, 0.3036928644233879, 0.3494929528372943, -0.3806064515522984, -0.36322719184900176, 0.1812460089008381, 0.20797141806254732, 0.15787959100380852, -0.0636343821800119, 0.23362136475552556, 0.18486420078401353], 'var': [1.3366292007778482, 1.1260127995542892, 0.5397194970856732, 1.1464691638520583, 1.036263008956585, 1.2249919507241267, 1.1616885291656642, 1.1251966167126373, 1.2196691562094661, 1.2003553544848022, 0.7856801376368169, 0.920876791408655, 0.9052971398126473, 1.2140570461105102, 1.0359819895084563, 0.9535086599206211, 1.1940143145796225, 1.688603480109305, 0.209819228402907, 0.3471671255512018, 1.0409426960311055, 1.3726619114252188, 3.9374253702654425, 3.954484631726769, 0.15515853974288568, 0.39535124691330403, 1.6292043790268873, 1.5446271786490802, 1.4819835281237557, 1.6470518813210646]}, 'mixing': 0.24221349849914206, 'non_immunogenic': {'mean': [-0.045747045590407244, -0.05292431518252884, -0.16390983498140319, -0.039439618085368305, -0.024631110587640812, -0.02222200965538523, -0.050333827270721665, 0.05168998874322181, 0.03817154361621429, -0.027720096794128773, -0.156089641087793, 0.013892240446834448, 0.018154293980941857, -0.07175953166992306, -0.06922895614473894, -0.026884336408952023, -0.07563193655738568, -0.06868895318676156, -0.07845303892968884, -0.07915267127605743, -0.09707023154348245, -0.11170944671072623, 0.12165434459341934, 0.11609936138149216, -0.0579321877836696, -0.06647450786862545, -0.050463511824108666, 0.020339642237095754, -0.07467307476381334, -0.05908868095494192], 'var': [0.8837619659862849, 0.9481580680597558, 1.0362003405333446, 0.946761697258523, 0.9859043820731401, 0.9260464126732396, 0.9378592831108543, 0.9489520642289404, 0.9237708935277261, 0.9327874264472823, 0.9679148477497845, 1.0244935996266071, 1.028909469481716, 0.9103204564222419, 0.9687121098501856, 1.0118761649627572, 0.9143703416756251, 0.7604202900895842, 1.2271568188380395, 1.182800668684306, 0.9480112139326297, 0.8293643401427054, 1e-08, 1.0000604388067371e-08, 1.256183024584032, 1.1750219909154196, 0.7883718670036672, 0.8242112460149019, 0.8229209257001397, 0.778766088689481]}}, 'supervised': {'immunogenic': {'mean': [0.6210075688060877, -0.01945745406543652, -0.0456845593184066, 0.49096696827720937, 0.009284831870431044, 0.34980986567909433, -0.03776810252455127, -0.5040323448856722, -0.3778243849799156, 0.4491687706693304, 0.22726993717797078, -0.03666951952972699, -0.012015389299309941, 0.2530790343049985, -0.03606705495612654, 0.23105006537598585, 0.6826210046999133, 0.6971941201822071, 0.31680177469563386, 0.33797595438939265, 0.5000992519576556, 0.5171573396090823, 0.3041438640133869, 0.32543117539887206, 0.2156688853186513, 0.25415039199985634, 0.3250055477442498, 0.39990527349316524, 0.520995487946946, 0.7740105907252691], 'var': [1.2813623788649668, 1.0944436080305044, 0.609586325629402, 0.9601960052068509, 1.0038183317945322, 1.2134318049465322, 1.134762221728634, 0.9696211866091288, 1.1479447936088318, 1.1216244835317215, 1.0438402680967365, 1.047034053169049, 0.9965362764094995, 1.293213682273709, 0.8470129928809526, 1.1029711277894307, 1.1194383746445848, 1.8580693977184501, 0.4555433654718886, 0.7542204950732135, 2.205020371283901, 2.4855523448589607, 1.21588981654603, 2.0986188692468333, 0.34107164739227974, 0.8404477191426631, 2.1634449925929142, 1.2683586944065994, 1.7429859405292798, 1.6264339438163766]}, 'mixing': 0.10933389902418328, 'non_immunogenic': {'mean': [-0.0762319108661606, 0.0023885037453731045, 0.005608017404064402, -0.0602687504049052, -0.0011397614314362387, -0.04294098146310941, 0.004636231134459426, 0.06187259337733515, 0.0463799094980813, -0.055137804125906825, -0.02789856752774337, 0.004501374354694335, 0.0014749515660115414, -0.03106676851350549, 0.0044274186929887695, -0.02836259793617407, -0.08379528076556317, -0.0855842065312017, -0.03888906651696485, -0.0414883072672556, -0.061389785752149, -0.06348375478362453, -0.03733524210748737, -0.039948370361688024, -0.026474478038686373, -0.03119828324601642, -0.0398961223520579, -0.049090453474608615, -0.06395490746626635, -0.09501382805216343], 'var': [0.9123094171127833, 0.9883543876943038, 1.0476376718507163, 0.9716638707894784, 0.9995194095793454, 0.9569350395998009, 0.9832606509002214, 0.9687151625307648, 0.9621644611036755, 0.9572636157856593, 0.9874995544093047, 0.9940410108245211, 1.0004053037245377, 0.9551790139305536, 1.0186006765232563, 0.9800021523476821, 0.921116261344951, 0.8276740651836324, 1.0530024117472894, 1.014427394638681, 0.8176077978599909, 0.7807794530560427, 0.9607492076296906, 0.8505425102712357, 1.074476276494231, 1.0106834952367647, 0.842622922007879, 0.9450161810747627, 0.8713841720317269, 0.8405326478905414]}}}, 'kd_threshold': -0.8500000000000001, 'log_odds': {'aa_anchor': [-0.00996606926865784, 2.5951294287960796, -0.7103218600532224, -1.0686311703318228, 0.10898980233447686, -0.0054052232940566824, -0.35002509664809267, 0.1299308601996776, -0.25575430871011307, 0.15914014219243922, 0.8438475469129401, -0.2154052812184819, 0.31073983106715586, -0.42513769705864757, -0.24503836658740097, -0.014759415233200457, -0.024221716914019975, -0.030625644094542714, 0.6050752830524884, 0.2338794791469785], 'aa_anchor10': [-0.17541957143717246, 2.7326276976177146, -1.0432479542571085, -0.9809849679193827, 0.17146525425650072, -0.1838256991781848, -0.3155083646078545, 0.2018746597042953, -0.16829576570932137, 0.23424666492180268, 1.1463362685451473, -0.3809655812624979, 0.12689690853304647, -0.538607351649234, -0.09942709206265876, -0.010023793194412267, -0.27514906857225396, 0.06614841191258813, 0.5991189346676027, 0.3723911710484873], 'aa_anchor11': [-0.3168130445829238, 3.0735274921799496, -1.6538603265323912, -1.8563577734433112, 0.21830068058655527, -0.23502282709736422, -0.5955136197143225, 0.3966240200745763, -0.16122168184454067, 0.30187204112316723, 1.2017253152783587, -0.3825464997080399, -0.053606247740738855, -0.27308169919780223, -0.3357182226650339, 0.04941392608773443, -0.22768212012676958, -0.019387620901657687, -0.37115500142194424, 0.7909691895357787], 'aa_anchor8': [0.02237782061173199, 2.0225972923790074, -1.0656591519964262, -1.4188199223425557, -0.1212825504875088, 0.21822242052087804, -0.2954493358225605, 0.2936764011221187, -0.385223984219631, 0.12734530110734155, 0.714744471202049, -0.04117427452429556, 0.23875121723027437, -0.47595188101171226, -0.20474429774975977, 0.15361262608581772, 0.17148384953597295, -0.06863251022860561, -0.09641546191943107, 0.41637177099455247], 'aa_anchor9': [0.05434809011842612, 2.7801617864888852, -0.35797680422744804, -0.8190052835219146, 0.11857350539615208, 0.0269631609070089, -0.3355168189680158, -0.023599697355983995, -0.20320776641489546, 0.1287692400630729, 0.6980107405386833, -0.21220174658533786, 0.4049254944493561, -0.41469409792973, -0.2473629738604517, -0.09902236803167508, 0.011755231096659546, -0.02067569534130609, 0.8729848450087081, 0.004205353352335628], 'aa_tcr': [0.11158462317236717, 2.773959531831438, -0.0993814377936384, -0.29671810806874266, 0.4950107076279453, -0.00016270415284225237, -0.295647663191398, 0.12410365697561065, -0.6520732243895435, -0.0091045818055, 0.8711861233484659, 0.09271524735876246, -0.2901417838675875, -0.5028664472546183, -0.0702567107737444, -0.05764223762239684, 0.021267736887371047, -0.05961947430965431, 0.8414969403179011, 0.44252305607409204], 'aa_tcr10': [0.14452336323282955, 2.7647428509620076, -0.27303124822977676, -0.30798896302550416, 0.6143544448118448, -0.2331943765628468, -0.04703302854604585, 0.1971390549584613, -0.5847531563232575, 0.18293207436529757, 1.2027446635994674, 0.0816790776425278, -0.5113814737783633, -0.6585624691458074, 0.049693229926436544, -0.005135993202697087, 0.028414528536732764, 0.015028342421871344, 0.7667575271402747, 0.30589199181565396], 'aa_tcr11': [0.042926325241647856, 2.9003697945156524, -0.7895690061266629, -0.7915433129800777, 0.6393630303776776, -0.3833691741191667, -0.2225349870647575, 0.24237121287574004, -0.47042485007288226, 0.38893639039422423, 1.4626645383965529, 0.15788930858279437, -0.5669489563014309, -0.5750469202268622, 0.10258790549005825, -0.10261002927518481, 0.1098189843463917, -0.006438385693241955, 1.1720230083245715, 0.7555479833927148], 'aa_tcr8': [-0.3626440402975324, 2.5690093765719366, 0.025152763677318113, -0.16113505155131636, 0.5716458918736516, -0.2940218366346805, -0.3270023489716065, 0.15256964736345102, -0.7693940827296135, -0.4177094648848714, 0.14329159817426618, 0.4257479627254508, -0.33763705409399103, -0.5355484143287876, -0.11204150782514999, 0.016121861455391517, 0.10389869076731628, -0.34296188767169866, 0.44423758739279595, 0.727599998700061], 'aa_tcr9': [0.25611875334301626, 2.7565035137077114, 0.0315792749221524, -0.30293376920372683, 0.4833436678602476, 0.2763942471895562, -0.39608284166471996, 0.06641716904970618, -0.6421198031734301, -0.08097270631773501, 0.7710438091180425, -0.1006770393040215, -0.053726795463244326, -0.4737200041343188, -0.10586162477797378, -0.09686090704837946, -0.05122900072975867, -0.011995701167633932, 0.9028414672177298, 0.3915866895934963], 'aa_tcr_c': [0.1567402196952874, 2.7264368854118315, -0.18612278745969535, -0.26665195917354545, 0.7469810008938964, 0.05105625357002408, -0.30566025168897903, 0.12538129624249628, -0.9114980709696425, 0.0758398085986891, 0.9547759452427305, -0.05394852113794668, -0.05793638638900056, -0.5162769946096821, -0.2905931903804695, -0.04758468295787388, 0.04809361857382566, 0.004081676160733405, 0.929273848867064, 0.4292549334284228], 'aa_tcr_m': [0.06040864281957914, 2.944928961060338, -0.39027352332045373, -0.5272985419894227, 0.3926480736280371, -0.08845857572090399, -0.3938329496610251, 0.07786352302480992, -0.9161797421571483, 0.02228140375683152, 0.8838121918738211, 0.14101175178008019, -0.7282021360854594, -0.5074887196245412, 0.12285779355736359, -0.16303490498628337, 0.02092417847246919, -0.13251087115944982, 0.5824957217012088, 0.48492599355688126], 'aa_tcr_n': [0.08560626373448521, 2.6587189318351836, 0.16515666836662835, -0.22547288708232927, 0.3320831434654905, 0.04564772848685372, -0.19907053579036305, 0.15704050294825533, -0.1325443306694063, -0.23930737878525132, 0.7037971757516512, 0.25022020515325893, -0.24378592296794555, -0.49363813072452123, 0.10065796644980018, 0.02027761719567467, -0.08179738019077876, -0.11926920006029063, 0.970448592894039, 0.41046372086635596], 'kmer_llr': [0.15700106006952907, 2.6020064931053426, -0.5148701638077364, -0.18437555126677374, 0.45976405018647615, 0.698371827198268, -0.8381063048064856, 0.14787150198409638, -0.8799659287420241, 0.19629552225488744, 0.87478554501486, 0.1514119883061218, -0.3684079724643574, -0.8460551975330297, -0.1334426821849819, -0.011634023525791903, 0.4300882820403604, -0.14396294285916067, 1.119320286310387, 0.2153382829007553, 4.106083889881617, 4.656130226800889, 2.8643707575728348, 2.091180869339352, 2.091180869339352, 3.035642478180204, 2.6789675342414725, 3.7006187817734526, 1.1748901374651979, 2.000209091133626, 2.091180869339352, 2.1601737408263038, 3.007471601213507, 2.6507966572747748, 2.378862941791134, 2.2988202341175974, 4.211444405539444, 2.879638229703623, 1.8034987968875704, 4.083611034029559, 0.20225765165764997, 2.245331549166611, -0.05952369236197441, -0.521208213193745, 0.39490341598443557, -0.24821819677740908, -0.361149610014035, -0.07923245054620942, -0.8194520736966346, -0.09136454533381766, 0.924745984332481, 0.7098005230218911, -0.9473714013975663, -0.6442683059509715, 0.03279273685734907, -0.20513461064109784, -0.2216959096912401, -0.41179078784628675, -0.05725354382743486, -0.20782640230680904, 0.06447562374045912, 2.879638229703623, -1.2299738046492665, -0.6597463561192498, -0.1472259979632824, -0.15531475700364972, -0.738103196576974, -0.362693671200919, -1.1978992142020966, -0.7102434025403159, 0.252901384476405, -0.04549890419540059, -0.43614649633259983, -0.5934078450838598, -0.5963708102145171, -0.5453221804708095, -0.11434251081156166, -0.31961780829492614, -0.1945971053383122, -0.2615540073401643, 0.7316646273008551, 2.6890178700949727, 0.045640879335198115, 0.2620338680396772, 1.2058736829851648, -0.1710161627494058, -0.12802261471564158, 0.7307277394033491, -0.4058471409809927, 0.5904761570417181, 1.01582544283551, 0.38410448734133595, 0.06701852338429148, -0.5103858513604074, 0.0004397724055831276, 0.5003548040998513, 0.2838204596688172, 0.3440535541320662, 1.2203525115419556, 0.9825182448177419, 0.21553991613059598, 2.0111381616658166, -0.22103511671565457, -0.29100727425366024, 1.1293668824411132, 0.1245348218048612, -0.912850207029333, -0.15907840465295298, -1.3992476460507453, -0.12698688184211004, 0.8316382138586063, 0.25216851526075157, -0.31892456736594976, -0.12679134019111515, -0.25715024544289733, -0.5733842198231187, -0.011075350106655613, 0.5441548892957826, 1.161644910715177, 0.38845502248169517, 0.007120381803221498, 2.378862941791134, -1.3298191396189818, -0.5829677800871762, 0.17425825715729193, -0.9651760260310729, 0.07116274813031787, -0.7504007243873803, -0.6902968003176753, -0.6408225727853507, 0.7048865082194613, 0.04597249930175895, -0.9915627812042684, -0.7652893368811311, -0.3067144034590177, -0.11342381529448886, -0.26638635807204647, -0.38908540089214494, -0.029082666860739792, 0.2193786924377612, 0.0015869561954984235, 2.6665450142429137, -0.3067144034590177, -0.1979812033225521, 0.5574547786409481, -0.3039557810199378, 0.10704950746384156, 0.5762084231580111, -1.291088502437284, 0.13511834882002116, 0.7972598283504766, 0.09663877144600352, 0.1752729705878382, -0.3603465583748129, 0.06422951028091362, 0.003957187217461744, -0.04352335101553262, 0.045315623193121546, 1.3492435246099754, 0.7474461226382578, -0.6633893473977501, 2.9666496066932533, -0.754212893273734, -0.7898884958944992, 0.026777205024057338, -0.968379445748611, -0.7155408592698871, -0.6423484170638423, -1.0020237470234585, -0.8258591356389733, 0.5713551155949395, -0.41255327358662175, -1.1006662831409288, -0.4860010565578188, -0.10176104620486548, -0.4345477749689026, -0.3049302816655164, -0.7972047905851829, 0.08584729981323846, 0.041431117466919076, -0.06742887186908675, 2.3424952976202587, -0.20922794423209812, -0.20482525255665074, 0.46090887000241043, -0.33021237472632237, -0.6053085093617812, 0.3292743632609785, -0.8339338370006848, 0.3597360501326827, 0.6830674608248222, -0.3215633342004178, -0.8004527183772341, -0.5564113627257434, -0.07270728772927537, -0.1216082245379786, -0.032235585920354026, 0.21617868970708987, 1.778806184297201, 0.6464981899813074, 0.814887403433791, 3.5952582661156267, 1.481415297718458, 1.197362993317256, 1.5605526182771827, 0.822669543875846, 0.2193786924377621, 0.9040151833297978, -0.133442682184981, 1.3343178743932995, 1.3439664675091318, 0.7141889008817754, 0.580588791541885, 0.14527072028403953, 0.3301930587780513, 0.9639952082271863, 0.9242493366937659, 0.992568580671243, 3.7959289615777774, 0.8102470238772881, 0.12260192047132623, 2.7843280498992975, 0.18750348390731197, -0.6013651452268922, 0.9381333741162265, 0.23379335643994636, -0.5580288317399242, -0.003764858876448507, -0.3937257804486478, 0.16297029738344015, 1.2087916891408783, 0.13778004855488124, -0.3098442964679453, -0.46936319465426735, 0.42767573563497674, -0.08389090335573357, 0.06303262204706694, -0.3188521114734417, 1.120401952181128, 0.7235786412316143, -0.5605509011726335, 1.8680373180251433, -0.5506199462216568, -0.7504007243873803, 0.6027812852823082, -0.4666204948568975, -0.17750267197901248, -0.17825201633779297, -1.5516546462731764, -0.3271110313666634, 0.7048865082194622, -0.677088191755912, -0.3435759967631764, -1.3648931225486374, -0.6024805943107578, -0.1833837113225183, -0.582083215677442, -0.703762498456233, -0.5976379698349703, -0.09314030316095856, -0.532987847782155, 3.631625910286502, -0.19385991400378355, -1.2193621440546716, -0.47856466038567724, -0.07977578117625672, -1.2184493187973127, -1.0370405872607176, -1.5490334127933032, -0.5973824440893614, 1.1046858787919485, -0.3645154023488466, -0.9480365161543771, -1.0720578752234529, -0.8995388623910943, -0.6556150983386999, -0.6789047226823106, -0.24968641022570992, 0.3565798139512468, 0.418658981478675, 0.318469595081992, 2.5611844985850887, -0.21763477340532944, -0.22690841019065822, 0.3757374879841793, 0.03567209387114456, -0.07469317811279019, -0.6744391835843349, -0.5780294984465932, -0.1907849364519576, 1.1025694758855717, 0.21019026638335525, -0.3004250745514536, -0.6628544450046814, 1.6493481170603133, -0.12179206496500594, 0.2834210587648567, -0.15868744148228853, 1.323925716625685, 0.2193786924377612, 0.2575189084572882, 2.208963904995736, -0.36669710806072864, -0.7051619389085353, 0.3723538696023292, 0.03557180302761509, -0.37973953847390796, 0.16180002223524959, -0.9596707034357177, 0.1874710747744155, 1.0413587448406751, -0.2258160023159963, -0.22718809135595475, -0.8381063048064856, -0.7495939474174405, -0.27187604540971044, 0.14095105913952288, -0.05864146910228296, 0.5797233652654565, 0.3388802368518453, 0.2486490747378749, 3.14100299383803, -0.2952857076353954, -0.8223133059776666, 0.26752803633534317, -0.1854899348892225, -0.19882544144783232, 0.18312094441513693, -0.5215591519585328, 0.33552640475341367, 1.0957528169064732, -0.22690841019065822, -0.3579066826436952, -0.5898406593749383, -0.1127551425848532, -0.10809920717896304, 0.23728961900429102, -0.03452239015655678, 1.0795799576608731, 0.461264425561712, 0.008273119683029684, 3.295153673665289, -0.33656736660869857, -0.637347863106398, 0.3757374879841793, -0.22310026341788447, -0.275942744792264, -0.0028594717616359233, -1.1993509444352837, -0.07400435425225016, 1.4190870979772399, -0.232431590846768, -0.44559761115548824, -0.33923759516457785, -0.6837199926588085, -0.35037533957503353, -0.14705892056428826, 0.13778004855488124, 0.5042158127573115, 0.6635012920566075, 0.3162285184276783, 2.4966459774475167, 0.14527072028404042, 1.5521843686066656, 1.0473768171662377, 0.3181135331234506, 0.40478191576912437, 0.7694250293570333, -0.023351992151753542, 1.4991298056507762, 1.6857157612311875, 1.2847050034724043, 0.27880211290856227, 1.1596226653344095, 0.3235189516903585, 0.11401817677993531, 0.9254292778336151, 0.3864327771009277, 2.091180869339353, 1.302723508975082, 0.5322268623372439, 2.9384787297265564, 0.0702355343411254, 0.3864327771009277, 0.8577941798071231, -0.09487040739874164, 0.600107556986285, 1.0278145197329174, -0.14952881993660494, 0.2680557408330779, 1.146719260498501, 0.8464759621737468, 0.06830967914791142, 0.5318673440878205, 0.1728586768028686, 0.3522638814856771, -0.026975184521174533, 0.12506801296652004, 0.9672507726869535, 1.0351281950900386]}, 'log_odds_source': {'aa_anchor': 'anchor', 'aa_anchor10': 'anchor@10', 'aa_anchor11': 'anchor@11', 'aa_anchor8': 'anchor@8', 'aa_anchor9': 'anchor@9', 'aa_tcr': 'tcr', 'aa_tcr10': 'tcr@10', 'aa_tcr11': 'tcr@11', 'aa_tcr8': 'tcr@8', 'aa_tcr9': 'tcr@9', 'aa_tcr_c': 'tcr_c', 'aa_tcr_m': 'tcr_m', 'aa_tcr_n': 'tcr_n', 'kmer_llr': 'pair'}, 'logistic': {'coef': [0.12551420796218662, 0.011746965690129688, -0.22796120873496709, 0.07389043313873397, -0.10648502804956149, 0.0981711722426051, 0.12875607543264733, -0.07270088127999416, 0.22781527923445063, -0.01464490734396946, 0.19703550119972588, 0.05942800972836916, -0.04333446425231443, 0.032323813691251826, 0.06151924942123461, -0.10680728245257694, 0.08415955976711734, -1.0316874083306022, 0.3513886251185817, 0.38937808996351464, 0.37747313651651576, 0.44008083790377955, 0.31284697323611715, 0.3089903633919363, 0.2950373758403601, 0.2913008156208962, 0.22854742491454486, 0.0651838617708407, 0.2907564651418195, 0.893787772998956], 'intercept': -2.6803856922142644, 'tau': 4.0}, 'n': 47140, 'n_immunogenic': 5154, 'paratope': {'A': [0.0468, 0.3668], 'C': [0.1994, 0.5024], 'D': [0.1813, 0.5419], 'E': [0.1014, 0.313], 'F': [0.0754, 0.3616], 'G': [-0.0, 0.2369], 'H': [0.2189, 0.5621], 'I': [0.0486, 0.416], 'K': [0.1234, 0.4516], 'L': [-0.0251, 0.1864], 'M': [0.0345, 0.3921], 'N': [0.1193, 0.3291], 'P': [0.0469, 0.2501], 'Q': [0.0384, 0.3411], 'R': [0.0515, 0.421], 'S': [0.0124, 0.2607], 'T': [0.1307, 0.3778], 'V': [0.0143, 0.2991], 'W': [0.0847, 0.4427], 'Y': [0.0172, 0.2474]}, 'prevalence': 0.10933389902418328, 'seed': 20260817, 'standardizer': {'mean': [3.649061829380585, 8.026884001018251, 9.072910479422996, 5.918703917394879, 4.0730983529274685, -2.2696420880144097, 3.9537856480908196, -0.3019153585065778, 0.5219495120916435, 13.831227195587637, 10.148506788290465, 0.06310424664121059, 0.33805213484655444, 1.8381417055579126, 1.277895630038184, 0.5352962805826635, -0.252232403137967, -0.1718184419660689, -0.12433160035260404, -0.0722116039466836, -0.06933721421291336, -0.0646821815674958, -0.049932359407905566, -0.04023770735202519, -0.08874299682438072, -0.0755211620108445, -0.03413589022035551, -0.07822360968723212, -0.0987321163119205, -0.3785270097660761], 'std': [18.32119223144693, 14.18739320828701, 0.9419106801415313, 13.983141047815161, 10.34894494135274, 12.899403706872747, 9.853255541060484, 1.9410970099623657, 1.8385842630913234, 2.0065068003403557, 2.7392906053643205, 0.02962738503853292, 0.0469905931559011, 1.0282499024909733, 0.5717053374996469, 0.24464656713775987, 0.8864423156372201, 0.6707030040588131, 0.6658127889389893, 0.416676136943885, 0.494762791203886, 0.4544240965858623, 0.41044452275654136, 0.34657991976206737, 0.5802170482251391, 0.45662314333473913, 0.28677994099002807, 0.46936947578760807, 0.5137388926023074, 1.2583479310885775]}}}#

standardizer, linear head, the fitted log-odds tables, and the EM / supervised Gaussian parameters kept for comparison. Class-I artifacts are loaded eagerly because every caller that predates the class split expects them present at import.

Type:

The frozen class-I models, one per species

mhcmatch.complement.SPECIES = ('human', 'mouse')#

Species with a fitted table. Both are the chowell_rebuilt arm of their own host – human 464,161 rows / 14,712 immunogenic, mouse 47,140 / 5,154 – never pooled, because the two hosts have different MHC and different thymic repertoires and a fit across them is fitting a mixture.

mhcmatch.complement.CLASSES = ('mhc1', 'mhc2')#

The two MHC classes, which are two different constructions and not one with a parameter.

mhcmatch.complement.BLOCKS = {'aa': ['aa_anchor', 'aa_tcr', 'aa_anchor8', 'aa_tcr8', 'aa_anchor9', 'aa_tcr9', 'aa_anchor10', 'aa_tcr10', 'aa_anchor11', 'aa_tcr11', 'aa_tcr_n', 'aa_tcr_m', 'aa_tcr_c'], 'kmer': ['kmer_llr'], 'motif': ['kd_run_max', 'kd_run_n', 'kd_run_frac'], 'phys': ['pc1', 'pc2', 'length'], 'pot': ['mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr'], 'role': ['pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr']}#

it is the class-I layout and every vendored class-I artifact is written against it.

Type:

Class-I feature blocks. BLOCKS keeps its name and its contents

mhcmatch.complement.BLOCKS_MHC2 = {'aa': ['aa_anchor', 'aa_tcr', 'aa_nflank', 'aa_core', 'aa_cflank', 'aa_anchorL0', 'aa_tcrL0', 'aa_anchorL1', 'aa_tcrL1', 'aa_anchorL2', 'aa_tcrL2', 'aa_anchorL3', 'aa_tcrL3'], 'kmer': ['kmer_llr'], 'motif': ['kd_run_max', 'kd_run_n', 'kd_run_frac'], 'phys': ['pc1', 'pc2', 'length'], 'pot': ['mj_anchor', 'mj_tcr', 'para_tcr', 'para_sd_tcr'], 'role': ['pc1_anchor', 'pc2_anchor', 'pc1_tcr', 'pc2_tcr', 'kf4_anchor', 'kf4_tcr']}#

Class-II feature blocks. The aa block has the same shape as class I’s – the pooled role pair, a position key, and a length key – and differs only in what the position key is: relative thirds of the TCR face at class I, register zones at class II. See encode().

mhcmatch.complement.FITTED = {'aa_anchor': 'anchor', 'aa_anchor10': 'anchor@10', 'aa_anchor11': 'anchor@11', 'aa_anchor8': 'anchor@8', 'aa_anchor9': 'anchor@9', 'aa_tcr': 'tcr', 'aa_tcr10': 'tcr@10', 'aa_tcr11': 'tcr@11', 'aa_tcr8': 'tcr@8', 'aa_tcr9': 'tcr@9', 'aa_tcr_c': 'tcr_c', 'aa_tcr_m': 'tcr_m', 'aa_tcr_n': 'tcr_n', 'kmer_llr': 'pair'}#

Columns computed from a fitted log-odds table rather than from the peptide alone: feature -> which count matrix it weights. <matrix>@<bin> is that matrix restricted to the rows of one length bin, so a per-length table costs a mask rather than another count matrix.

mhcmatch.complement.FITTED_MHC2 = {'aa_anchor': 'anchor', 'aa_anchorL0': 'anchor@0', 'aa_anchorL1': 'anchor@1', 'aa_anchorL2': 'anchor@2', 'aa_anchorL3': 'anchor@3', 'aa_cflank': 'cflank', 'aa_core': 'core', 'aa_nflank': 'nflank', 'aa_tcr': 'tcr', 'aa_tcrL0': 'tcr@0', 'aa_tcrL1': 'tcr@1', 'aa_tcrL2': 'tcr@2', 'aa_tcrL3': 'tcr@3', 'kmer_llr': 'pair'}#

The class-II equivalent. <matrix>@<bin> here indexes mhc2_length_bin(), not length_bin() – encode() writes the right one into counts["bin"] per class.

mhcmatch.complement.ZONES = {'mhc1': ('tcr_n', 'tcr_m', 'tcr_c'), 'mhc2': ('nflank', 'core', 'cflank')}#

Which zone matrices a class partitions its TCR-facing face into. The pooled tcr matrix is their sum in both cases, so the split costs no extra pass and posbayes stays recoverable.

mhcmatch.complement.MHC2_ZONES = ('nflank', 'core', 'cflank')#

the residues before the 9-mer core, the non-anchor residues inside it, and the residues after it.

This was built as the class-II replacement for LENGTH_BINS, on the reasoning that a floating core makes total length uninformative. The measurement says join, not replace (bench/results/complementarity_mhc2.md): the zones earn +0.0029 human / +0.0034 mouse AUROC over the pooled pair, total length earns +0.0070 / +0.0159, and carrying both beats either on both hosts and both metrics. The reasoning was right about the core and wrong about the ligand – an 18-mer and a 13-mer with the same core do present the same residues, but a class-II ligand’s length is the length of its flanks, which is its own covariate and not a register question.

Type:

Register-relative zones for class II

mhcmatch.complement.TCR_THIRDS = ('tcr_n', 'tcr_m', 'tcr_c')#

Relative thirds of the class-I TCR-facing face. The same cell means the same fraction along the peptide at every length, which is what the contact profile already does for its per-position weights.

mhcmatch.complement.LENGTH_BINS = (8, 9, 10, 11)#

8, 9, 10 and 11+. Closed at both ends, which is a shippability property and not only a variance one – a model with one table per observed length cannot score a 12-mer at all, and real inputs contain those.

Type:

Length bins for the class-I aa block

mhcmatch.complement.MHC2_LEN_EDGES = (14, 16, 19)#

Class-II total-length quartile edges, from the corpus’s own distribution (11-25, median 15). LENGTH_BINS cannot be reused: it clamps to 11 and would put every class-II ligand in one bin.

mhcmatch.complement.mhc2_length_bin(L)[source]#

Which class-II length quartile (0-3) a ligand of length L falls in. Closed both ends.

Return type:

int

mhcmatch.complement.length_bin(L)[source]#

Which of LENGTH_BINS a peptide of length L falls in.

Return type:

int

mhcmatch.complement.KD_THRESHOLD = -0.8500000000000001#

the median of the Kyte-Doolittle scale itself over the 20 standard residues (-0.85), so it is a property of the scale rather than a constant tuned on any corpus. It admits ACFGILMSTV and excludes the other ten. Same rule as mhcmatch.immuno._aggregate().

Taken from the plain dict with the stdlib rather than from BASIS with numpy, because sphinx mocks numpy at doc-build time and a module-level float(np.median(...)) then raises on a Mock – the whole module fails to import and its page renders empty.

Type:

The “hydrophobic” cut for the motif run features

mhcmatch.complement.PHYS_SCALE = 'Rose'#

The chemistry half of the recognition axis, as a scale name. "Rose" is the shipped choice – the average fractional area a residue buries on folding – selected out of 576 candidates by the BIC change it produced inside the general model, not by standalone AUROC. Any other name here is exploratory: it re-parameterises a published result and should be reported as a comparison, never substituted silently. The shipped aggregate records the same name as phys_scale; this constant is what burial() actually reads.

mhcmatch.complement.blocks(cls='mhc1')[source]#

The feature blocks of one class: BLOCKS or BLOCKS_MHC2.

Parameters:

cls (str)

Return type:

dict

mhcmatch.complement.fitted(cls='mhc1')[source]#

The fitted-column map of one class: FITTED or FITTED_MHC2.

Parameters:

cls (str)

Return type:

dict

mhcmatch.complement.mhc2_anchors(peptide, register=None)[source]#

0-based P1/P4/P6/P9 of a class-II ligand’s 9-mer core, memoised.

Delegates to mhcmatch.store.anchor_indices(), so the register is the same one decompose, the logos and the signatures use. register pins the frame – pass mhcmatch.diffusion.AnchorModel.best_register() to annotate with the frame the model actually scored with, rather than the allele-agnostic heuristic.

The heuristic register is an argmax over up to seventeen offsets per peptide and is the only non-vectorised step in encode(), so it is cached: a corpus repeats its peptides across alleles and folds, and the register does not depend on either.

Parameters:
  • peptide (str)

  • register (int | None)

Return type:

tuple

mhcmatch.complement.table(species='human', cls='mhc1')[source]#

The fitted parameters for one (species, cls).

Parameters:
  • species (str)

  • cls (str)

Return type:

dict

mhcmatch.complement.phys_scale(scale='Rose')[source]#

Resolve a residue scale by name to {residue: value}.

Accepts a key of mhcmatch.data.aa_tables.HYDROPHOBICITY (45 scales, including the shipped "Rose"), a "FAMILY:COMPONENT" key into DESCRIPTORS (e.g. "KIDERA:KF4"), or a ready dict over AA. Raises rather than guessing on an unknown name.

Return type:

dict

mhcmatch.complement.burial(peptides, cls='mhc1', scale='Rose', registers=None, per_residue=True)[source]#

C_phys: a residue scale averaged over the TCR-facing positions.

With the default scale="Rose" this is the shipped chemistry term – the Rose burial propensity, i.e. the average fraction of a residue’s surface buried on folding, over the face a receptor reads. It has no fitted residue parameters: the basis is imported, which is why it cannot memorise the corpus’s cysteine gradient (correlation with per-peptide cysteine count +0.108, against +0.688 for the full fitted score()).

``per_residue=True``, and the old sum was a length detector. The TCR face is L - 5 residues wide and the Rose scale is strictly positive (0.52 to 0.91), so summing it gives roughly 0.75 (L - 5): on 60,000 fit-corpus peptides the summed column correlates with peptide length at Pearson +0.954, against +0.052 for a centred scale like Kidera KF4. A chemistry term that is 91 % length variance is not measuring chemistry, and it made the two scales incomparable – one carrying length, one not. Dividing by the face width is the same correction mhcmatch.mimicry.corpus_R() makes with its per-window divisor, for the same reason. Pass per_residue=False to restore the summed form.

scale= is for exploration, not for scoring. Passing another basis re-parameterises a result that was selected by BIC inside the general model over 576 candidates, so a number produced with a different scale is a comparison and must be reported as one. The obvious comparisons – Kidera factors, against which the Chowell-family literature is usually written – are reachable as "KIDERA:KF4", "KIDERA:KF2" and so on; on the neoantigen corpus they lose to "Rose".

>>> burial(["GILGFVFTL"])[0] > 0
True
>>> abs(burial(["GILGFVFTL"], scale="KIDERA:KF4")[0]) >= 0
True
Parameters:
  • cls (str)

  • per_residue (bool)

Return type:

list[float]

mhcmatch.complement.encode(peptides, cls='mhc1', registers=None, positions='mask', paratope='loop')[source]#

(features, counts) for an iterable of peptides – the vectorised feature builder.

Everything additive is a matrix product against a residue vector, so the peptide-side feature set is two (n, 20) count matrices (one per role) times seven property vectors. Run statistics need position order and use a loop over positions (at most 25 iterations), never over peptides. Peptides are grouped by residue layout so each group is one batch of array operations.

counts carries anchor and tcr (n, 20) – the matrices the aa log-odds tables weight – plus that class’s ZONES matrices and the adjacent TCR-facing residue pairs as a sparse pair list, pair_code (which of 400 pairs) and pair_row (which peptide). A 9-mer has at most 3 such pairs, so a dense (n, 400) matrix would be 99% zeros and would cost 1.5 GB of temporaries per pass on a 500k-peptide corpus; apply_log_odds() sums the sparse form instead.

The two classes differ in one place and it is deliberate. Class I is anchored at fixed peptide positions (ANCHORS), so its geometry is a function of the peptide’s length: the aa block carries one table per LENGTH_BINS bin, and the TCR face is split into relative thirds. A class-II ligand is anchored by a 9-mer core that floats inside an 11–25-mer, so its total length is the length of its flanking regions and carries no register information – an 18-mer and a 13-mer with the same core present the same residues to the receptor. Binning a class-II table on total length would therefore split it on a variable that says nothing about the object being modelled. The class-II construction bins on the register instead (MHC2_ZONES): before the core, inside it, after it. That both classes reach the same aa_anchor / aa_tcr pooled pair on top of it is what keeps mhcmatch.posbayes.llr() a strict special case of either.

registers (class II only) pins each peptide’s core to an explicit frame, element for element with peptides; None uses the allele-agnostic heuristic register.

positions selects which peptide positions the TCR-facing chemistry is read over – the binary anchor mask ("mask", the shipped default) or the positional contact profile ("profile"). paratope selects which marginalisation of the TCRen potential the pot block projects onto – PARATOPE ("loop", the shipped default) or PARATOPE_CONTACT ("contact"). The two compose, and both are explained in full on score(), which is where the measured consequences of changing either are recorded.

The hydropathy stretch, exactly. A residue enters the motif block when all three hold: it is TCR-facing (an anchor is buried in the groove and is not part of any stretch the receptor reads), it is one of the 20 standard residues, and its Kyte-Doolittle value exceeds KD_THRESHOLD. kd_run_max is then the longest run of consecutive such positions, kd_run_n the number of runs (counted as rising edges, so IIDI is two runs and IIDD is one), and kd_run_frac the above-threshold count divided by the number of TCR-facing positions. An anchor breaks a run rather than bridging it: two stretches either side of a buried residue are two stretches.

Non-standard residues (X masks, B/J/O/U/Z) contribute to no count matrix and to no property sum, but they still count toward length – and in the motif block they behave exactly like a below-threshold residue rather than like a gap. A mask breaks a run (AAAIIXIAA gives kd_run_max = 2, the same as AAAIIDIAA, not the 3 it would give if the mask were transparent) and it sits in kd_run_frac’s denominator while never entering its numerator. That is the conservative reading of an unknown residue – it is not evidence of a continued hydrophobic stretch – and it is stated here because the opposite was once claimed.

Parameters:
  • cls (str)

  • positions (str)

  • paratope (str)

Return type:

tuple[dict[str, numpy.ndarray], dict[str, numpy.ndarray]]

mhcmatch.complement.apply_log_odds(counts, source, weights)[source]#

A fitted log-odds table applied to one count structure – one value per peptide.

anchor / tcr / the three TCR_THIRDS are dense (n, 20) and this is a matrix product. pair is the sparse pair list, and summing weights[code] per row with numpy.bincount() costs O(pairs) rather than O(n x 400) – the difference between seconds and a gigabyte of temporaries on a corpus.

"<matrix>@<bin>" is that matrix restricted to one LENGTH_BINS bin. The restriction is applied to the result, not by materialising a masked copy: a per-length table is then a length-n mask rather than another (n, 20) matrix, which is what makes eight of them affordable.

Parameters:
  • counts (dict)

  • source (str)

Return type:

numpy.ndarray

mhcmatch.complement.design(peptides, species='human', cls='mhc1', registers=None, mask_cys=False, positions='mask', paratope='loop')[source]#

The (n, k) design matrix in feature_names() order, before standardization.

mask_cys, positions and paratope are forwarded unchanged; see score() for what each one does and what it costs.

Parameters:
  • species (str)

  • cls (str)

  • mask_cys (bool)

  • positions (str)

  • paratope (str)

Return type:

numpy.ndarray

mhcmatch.complement.feature_names(species='human', cls='mhc1')[source]#

Column order of the design matrix, matching the vendored coefficients.

Parameters:
  • species (str)

  • cls (str)

Return type:

list[str]

mhcmatch.complement.features(peptide, species='human', cls='mhc1')[source]#

The feature vector for one peptide, as a name -> value dict. For inspection; use score() (which takes a list) for anything at scale.

Parameters:
  • peptide (str)

  • species (str)

  • cls (str)

Return type:

dict[str, float]

mhcmatch.complement.score(peptides, species='human', cls='mhc1', registers=None, blocks=None, mask_cys=False, positions='mask', paratope='loop')[source]#

Log-odds of immunogenic vs not, one per peptide. Carries no prior.

Larger is more immunogenic. The training corpus’s own base rate is divided out, so this is directly comparable across settings and composes with any prevalence via posterior().

Accepts a single string as well, and still returns an array of length 1 – so a caller never has to branch on whether they passed one peptide or a million.

blocks restricts the score to a subset of BLOCKS. The head is linear over a fixed block partition, so this is an exact partial sum of the shipped score, not a refit: the parts add back to the whole, up to the constant offset that belongs to no block. With off = logistic.intercept - log(prev / (1 - prev)) from table(),

score(p, blocks=(“phys”, “role”, “pot”, “motif”)) + score(p, blocks=(“aa”, “kmer”)) + off

== score(p)

exactly. The offset is dropped from a block score deliberately – an intercept is not attributable to any block, and a constant cannot change a ranking. Whole blocks only: aa_tcr is fitted at -0.8796 while every per-length aa_tcr{8,9,10,11} and every aa_tcr_{n,m,c} is positive, so the pooled column is collinear with its own decompositions and the fit pushed it negative to compensate. Dropping columns within a block without refitting inverts the dominant term; dropping a whole block cannot.

mask_cys scores with cysteine zeroed out of the fitted log-odds tables, as mhcmatch.posbayes does by construction. See _mask_cys() for the size of the artifact and why the default is off.

positions="profile" reads the TCR-facing chemistry over the positional contact profile (mhcmatch.immuno.contact_profile(): per-position TCR-peptide contact frequency from 8,062 contacts over 370 crystals, keyed by class and length) instead of the binary anchor mask. It touches the role, pot and motif blocks only – aa and kmer are count matrices keyed on a role and a weighted count is a different object, so they come back bit-identical. The anchor face is untouched for the same reason the TCR face is re-weighted: the profile measures contact with a receptor and carries no information about burial in the groove.

paratope="contact" reads the pot block’s TCRen projection off PARATOPE_CONTACT – the potential marginalised over the receptor residues that actually contact peptide – instead of PARATOPE’s flat whole-loop composition. Independent of ``positions``: one changes which peptide positions are read, the other which receptor residues the potential was averaged over, and they compose.

The coefficients are unchanged under either, so both are the shipped head reading a different encoding, not a refit. Whether either should be refit is the question a benchmark answers; the defaults stay "mask" and "loop".

Parameters:
  • species (str)

  • cls (str)

  • mask_cys (bool)

  • positions (str)

  • paratope (str)

Return type:

numpy.ndarray

mhcmatch.complement.kidera_design(peptides, anchors=None, roles=None)[source]#

All ten Kidera factors, each summed over the anchors, the TCR face and the whole peptide.

The fitted model uses one of the ten – kf4_anchor and kf4_tcr, the hydropathy axis – because that is the factor the role split was established on, not because the other nine were compared and rejected. This builds the full set so they can be. Label-free: it reads only the vendored Kidera table and ANCHORS, so it needs no fitted artifact and no per-fold refit.

>>> kidera_design(["GILGFVFTL"]).shape
(1, 30)
Return type:

numpy.ndarray

mhcmatch.complement.kidera_names()[source]#

Column names of kidera_design(), in order.

Return type:

list[str]

mhcmatch.complement.posterior(peptides, prior, species='human', cls='mhc1')[source]#

P(immunogenic | peptide) at an explicit prior. Exact, because score() has none.

prior has no default on purpose. The training corpus runs at ~3.2% positives, a viral proteome scan nearer 3.0e-3 and the NCI screen 4.2e-4 – a default would silently pick one setting’s base rate for every caller and overstate the rest by up to 75x.

Parameters:
  • prior (float)

  • species (str)

  • cls (str)

Return type:

numpy.ndarray

mhcmatch.complement.parameters(species='human', cls='mhc1')[source]#

The fitted model as a plain dict (a copy of the vendored file).

Parameters:
  • species (str)

  • cls (str)

Return type:

dict

mhcmatch.recognition module#

The head dispatcher over the recognition axis: complement (the six-block model above, and the default), plus posbayes, physchem_glm and esm64_glm, each fitted alone so their BIC is comparable. See Complementarity — the I and C of EPIC.

Recognition score for MHC epitopes: how immunogenic a peptide looks, given how it is presented.

Four heads. default_head() – what score() uses when none is named – is complement, the six-block, 30-feature model of mhcmatch.complement, which is served by its own module and has no vendored table here. The other three are each fitted alone so that their fit criteria are comparable and each score is readable on its own terms; lowest_bic_head() reports the parsimony winner among those three, currently posbayes for both species. The two answer different questions – BIC asks which head buys its parameters on one training arm, the default asks which recognition term to score with – so they do not have to agree, and they do not.

complement

Six feature blocks over roles, contact potentials, hydropathy runs, residue identity and adjacent residue pairs; 30 features, linear head. mhcmatch.complement, and the source of the shipped C_phys term. The default.

posbayes

Naive Bayes over amino-acid identity conditioned on face – MHC-facing or TCR-facing – scored as a summed log-likelihood ratio. Two 20-cell tables, three parameters. Conditioning on the face rather than on absolute position is what keeps it well defined when peptide length varies, and it is why the two tables can disagree in sign without anything being told to flip. Pure numpy, no optional dependency, and the whole model prints in forty numbers.

physchem_glm

Raw sums of the Kidera factors over each face. Length is not a separate feature: summing a constant-1 factor over a face gives that face’s size, so it enters as \(KF_0\) and the two face sizes add to the peptide length. 22 features, all interpretable.

esm64_glm

64 principal components of a whole-peptide ESM2 pool. The most accurate head on mouse and the least explainable; needs pip install 'mhcmatch[esm]'.

All three take what the caller knows, at whatever resolution they know it: peptide with both masks given, peptide with an MHC (masks from the allele’s layout), or peptide alone (the class-I default).

Coefficients come from chowell_iedb_full_matched – the rebuilt Chowell corpus with negatives resampled so the allele group carries no signal about the label, so no coefficient can be paid for recognising which allele happened to be typed. Selection, transfer and the full metric matrix are in bench/results/shipped_models.md.

mhcmatch.recognition.HEADS = ('complement', 'posbayes', 'physchem_glm', 'esm64_glm')#

complement is the six-block, 30-feature model of mhcmatch.complement and is the default. The other three are fitted alone on one arm so their BIC is comparable to each other; posbayes wins that comparison at three parameters. That is a statement about parsimony on one corpus, not about which term to score with: in the integrated neoantigen fit it is the six-block form that carries the recognition signal, and a 3-parameter head is a different claim that must not be substituted for it silently. BIC across the three is still reported by table().

mhcmatch.recognition.FITTED_HEADS = ('posbayes', 'physchem_glm', 'esm64_glm')#

The three heads with their own vendored artifacts. complement is served by its own module.

mhcmatch.recognition.table(head=None, species='human')[source]#

Fitted parameters for one head.

head=None resolves to lowest_bic_head(), not to default_head(): this returns a vendored table, and the default head is complement, which has none – it is mhcmatch.complement, whose parameters come from complement.parameters(species).

Parameters:
  • head (str)

  • species (str)

Return type:

dict

mhcmatch.recognition.default_head(species='human')[source]#

What score() uses when no head is named: the six-block complement model.

lowest_bic_head() still reports the parsimony winner among the three separately fitted heads. They answer different questions – BIC asks which head buys its parameters on one training arm, and the default asks which recognition term to score with.

Parameters:

species (str)

Return type:

str

mhcmatch.recognition.lowest_bic_head(species='human')[source]#

The lowest-BIC head among the three with their own fitted tables.

Parameters:

species (str)

Return type:

str

mhcmatch.recognition.feature_names(head=None, species='human')[source]#

Column names of design(), in order. head=None follows table().

Parameters:
  • head (str)

  • species (str)

Return type:

list[str]

mhcmatch.recognition.roles_for(peptides, mhc=None, anchors=None, store=None, cls='mhc1')[source]#

Per-peptide boolean masks, True where the residue faces the MHC.

Resolution order is explicit anchors, then the store’s layout via mhcmatch.store.Store.decompose(), then the class-I default. Everything downstream reads this, so it is the one thing worth getting right.

The two non-explicit routes report different anchor sets, deliberately. The default is the five-pocket set the shipped posbayes tables were fitted on (P1–P3, PΩ-1, PΩ). The store route is the store’s own layout: class-I P2/PΩ, or – the reason to take this route – class-II’s P1/P4/P6/P9 within the floating 9-mer core, which no length-only default can give. Pass cls="mhc2" with a store for a class-II ligand.

mhc is forwarded to decompose as its allele, which accepts it for forward-compat; the layout it returns is the class default and is not allele-conditioned today.

Parameters:

cls (str)

mhcmatch.recognition.design(peptides, species='human', head=None, *, mhc=None, anchors=None, store=None, cls='mhc1', roles=None)[source]#

The design matrix for one head, in feature_names() order, before standardization.

head=None resolves the same way table() does – to lowest_bic_head() – because only the three separately fitted heads have a design matrix here. The complement default is encoded by mhcmatch.complement.design().

Parameters:
  • species (str)

  • head (str)

  • cls (str)

Return type:

numpy.ndarray

mhcmatch.recognition.score(peptides, species='human', head=None, **kw)[source]#

Recognition log-odds. Higher is more likely to be immunogenic.

Not a probability – pass it through posterior() with a prior for the set being scored.

Parameters:
  • species (str)

  • head (str)

Return type:

numpy.ndarray

mhcmatch.recognition.posterior(peptides, prior, species='human', head=None, **kw)[source]#

sigmoid(score + logit(prior)). prior is the immunogenic rate expected of the set.

Parameters:
  • prior (float)

  • species (str)

  • head (str)

mhcmatch.recognition.embed(peptides, batch=256)[source]#

Whole-peptide mean-pooled ESM2 embeddings, (n, 1280).

Batched within length groups: slicing batches out of a length-sorted list and keeping only what matches the batch’s first length silently leaves the rest unembedded.

Parameters:

batch (int)

Return type:

numpy.ndarray

mhcmatch.recognition.mhc2_core(peptides, register_start=None)[source]#

The register-anchored 9-mer core of each class-II peptide, and where it starts.

mhcmatch.recognition.score_mhc2(peptides, species='human', head=None, *, register_start=None, warn=True)[source]#

Class-II score on an MHC-I-fitted model. Read this before using the number.

There is no fitted class-II recognition model here, and this is not one: no class-II immunogenicity corpus with a usable negative set exists to fit against. This applies the class-I coefficients to the class-II binding core, with the groove-facing positions redefined as P1/P4/P6/P9 of the register-anchored 9-mer.

Scoring the core rather than the whole peptide keeps every feature in the range the model was fitted on – nine residues, face sizes of four and five. Scoring a 15-mer directly would put the face sizes far outside it. But the coefficients were fitted where the groove-facing residues are the termini and the TCR-facing ones a contiguous middle, and in class II the two faces interleave, so a coefficient learned on one geometry is being read on another. The register is also a heuristic unless supplied, and an error in the frame moves every residue between faces.

Rank class-II peptides against each other with it. Do not compare the values with class-I scores, do not read them as calibrated, and say which model produced any number you report. nan where no 9-mer core can be assigned.

Parameters:
  • species (str)

  • head (str)

  • warn (bool)

mhcmatch.recognition.MHC2_ANCHORS = (0, 3, 5, 8)#

0-based P1/P4/P6/P9 within the register-anchored 9-mer core.

mhcmatch.recognition.log_odds_table(species='human')[source]#

posbayes’s two 20-cell tables as {face: {residue: log-odds}}. The whole model.

Parameters:

species (str)

Return type:

dict

mhcmatch.immuno module#

Physicochemical featurization of epitopes for immunogenicity prediction.

The signal this targets is the Chowell/Calis one: immunogenic epitopes differ from presented-but- non-immunogenic ones in the physicochemistry of their TCR-facing residues, not their anchors. Two things follow, and both are parametrized here rather than hardcoded:

Which positions are TCR-facing is not agreed even within this toolchain – mhcmatch.store.anchor_indices() masks class-I P2+PΩ, while seqtree.layout.presentation_features() and mhcmatch.diffusion.MHC1_ANCHORS mask P1-P3+PΩ-1,PΩ. ANCHOR_SCHEMES keeps all of them selectable so the choice is an ablation axis with a reported number, not a silent constant. "contact" is the third option: a continuous per-position weight from observed TCR-peptide contact frequency, which needs no anchor call at all.

How residues are aggregated matters as much as which scale is used. Summed/averaged descriptors are the established positive result (Chowell 2015 on hydrophobicity; Pogorelyy 2018 associates epitope length and summed Kidera factors 6 and 10 with precursor frequency), so sum/mean are always emitted. But a contiguous hydrophobic stretch is a different object from the same residues scattered across the peptide, and no sum can express it – hence run_max/run_n/run_frac.

length is emitted as a feature, deliberately. Ligand length distribution is allele-specific and is part of what distinguishes a real ligand set; it is signal here, not a nuisance to regress out.

All scales are plain dict[str, float] (see mhcmatch.data.aa_tables), so any AAindex-style table can be passed through the same call.

mhcmatch.immuno.ANCHOR_SCHEMES: dict[str, tuple] = {'full': (), 'p2_pomega': (2, -1), 'pockets': (1, 2, 3, -2, -1)}#

Class-I anchor position sets, 1-based (negatives count from the C-terminus), matching the three definitions that coexist in the toolchain. full masks nothing (whole-peptide baseline); contact is continuous and ignores this table entirely.

mhcmatch.immuno.DEFAULT_SCALES = ('KIDERA', 'VHSE', 'MJ', 'KyteDoolittle')#

VHSE, MJ potential, Kidera – plus Kyte-Doolittle, the axis Chowell used, so the reproduction gate runs on the same scale the paper did.

Type:

The scales named in the original brief

mhcmatch.immuno.scales(names=('KIDERA', 'VHSE', 'MJ', 'KyteDoolittle'))[source]#

Resolve scale names to residue->value tables.

A name may be a descriptor family ("KIDERA" -> its 10 components, each its own scale), a hydrophobicity scale ("KyteDoolittle"), or "MJ" for the Miyazawa-Jernigan partition energy. Unknown names raise rather than being silently skipped.

Return type:

dict[str, dict[str, float]]

mhcmatch.immuno.position_weights(peptide, cls='mhc1', scheme='p2_pomega', register_start=None, contact_profile=None)[source]#

Per-position weights: 0 at anchors, 1 at TCR-facing positions.

scheme is a key of ANCHOR_SCHEMES, or "contact" – which requires contact_profile, a callable (length) -> list[float] of continuous per-position weights (observed TCR-peptide contact frequency, normalized per length).

Class II ignores scheme and always masks the register-anchored core P1/P4/P6/P9, because that definition is agreed across the toolchain. Pass register_start from mhcmatch.diffusion.AnchorModel.best_register() so the annotated frame matches the scored one; None uses the allele-agnostic heuristic register.

Parameters:
  • peptide (str)

  • cls (str)

  • scheme (str)

  • register_start (int | None)

Return type:

list[float]

mhcmatch.immuno.features(peptide, cls='mhc1', scheme='p2_pomega', scale_names=('KIDERA', 'VHSE', 'MJ', 'KyteDoolittle'), register_start=None, contact_profile=None)[source]#

Physicochemical feature vector for one peptide.

Returns length plus, per scale, the seven statistics in _STATS keyed "{scale}_{stat}". Non-standard residues (X masks, B/J/O/U/Z) are dropped from the aggregation rather than scored as zero – a zero is a real value on a centred scale like Kidera and would be a silent bias.

Parameters:
  • peptide (str)

  • cls (str)

  • scheme (str)

  • register_start (int | None)

Return type:

dict[str, float]

mhcmatch.immuno.feature_names(scale_names=('KIDERA', 'VHSE', 'MJ', 'KyteDoolittle'))[source]#

Column order matching features(), for building a matrix without a dict round-trip.

Return type:

list[str]

mhcmatch.immuno.contact_profile(cls='mhc1')[source]#

Structure-derived per-position weights: (length) -> list[float], for scheme="contact".

Backed by mhcmatch.data.contact_profile – 8,062 TCR<->peptide residue contacts over 370 crystal structures. Two derived steps turn a contact frequency into a weight:

Zeroing. A position contacted less than half as often as a uniform footprint would predict (frac < 1/(2L)) is treated as not TCR-facing. The threshold is derived from the profile’s own uniform expectation, not tuned. On class-I 9-mers this zeroes exactly P1, P2, P3 and PΩ – the anchor set the contact data supports, recovered without being told anchors exist.

Rescaling. Surviving positions are scaled to mean 1, so the weighted statistics in features() are on the same scale as the binary schemes (where a kept position weighs 1) and the two are directly comparable.

Lengths with fewer than MIN_STRUCTURES observed structures fall back to the class’s pooled relative-position profile, interpolated onto the requested length – the TCR footprint scales with the peptide, so relative position transfers where absolute position does not.

Parameters:

cls (str)

mhcmatch.posbayes module#

Position-role naive Bayes over amino-acid identity: anchor and TCR-facing residues get separate conditional distributions, because for several amino acids their contributions carry opposite signs and pooling averages that away. Emits a log-likelihood ratio, so the caller supplies the prior.

Grouped 5-fold CV 0.712 human / 0.758 mouse; size-matched transfer 0.731 human→mouse, 0.692 mouse→human. Cysteine is masked — see the module warning.

Position-role naive Bayes: amino-acid evidence for immunogenicity, split by anchor / TCR-facing.

A whole-peptide physicochemical predictor scored a peptide from pooled physicochemical descriptors and did not distinguish where a residue sits. But anchor and TCR-facing positions are different channels – an anchor residue is buried in the MHC groove and a TCR-facing one is contacted by the receptor – and the benchmark finds their contributions carry opposite signs for several amino acids. Pooling averages that away.

This model keeps them apart in the simplest form that can: for each role separately, the conditional amino-acid distribution given the class, Laplace-smoothed, scored as a summed log-likelihood ratio.

score(p) = sum_i log [ P(p_i | role(i), immunogenic) / P(p_i | role(i), non-immunogenic) ]

Each peptide reduces to two 20-vectors of counts, which is what lets 8-11mers pool in one model.

It emits a log-likelihood ratio, not a probability. That is deliberate: the LLR carries no prior, so a caller can supply whatever base rate their setting actually has –

logit P(immunogenic) = llr(peptide) + log(prior / (1 - prior))

– and posterior() does exactly that. The distinction matters because the training corpus runs at ~3.2% positives while a viral proteome scan runs at ~3.0e-3 (counted from distinct 9-mers against known epitopes) and the NCI screen at 4.8e-4. Reading a corpus-prevalence probability as an operational one overstates it by 11-66x.

Measured performance, peptide-grouped 5-fold cross-validation (no peptide in both train and test), on the IEDB positive-T-cell-assay vs self-eluted-ligand corpus:

metric

human

mouse

rows

464,310

47,203

immunogenic

14,712

5,154

AUROC (this model)

0.712

0.758

Size-matched cross-species transfer, mean over 10 matched subsamples:

  • human -> mouse: 0.731 (sd 0.003)

  • mouse -> human: 0.692 (sd 0.000)

is not like-for-like. It is quoted because an in-sample baseline that still loses is the conservative direction, not because it is a fair contest. The module is gone; the measurement is not, and neither is this row.

Warning

Cysteine is masked, and the reason is an assay artefact worth knowing about. The negatives in this corpus are mass-spectrometry-eluted ligands, and cysteine is systematically under-detected in immunopeptidomics unless alkylated. Measured on the training corpus: Cys-containing peptides are 11.59% of the T-cell-assayed positives, 1.73% of the IEDB eluted negatives and 0.17% of the thymus MS negatives – a 6.5x depletion driven by platform, not biology. Fitted freely, Cys took the single largest coefficient in the model (+1.84 anchor / +2.05 TCR-facing). It is therefore zeroed here. The cost is small and measured: grouped-CV AUROC 0.712 -> 0.690.

Any model trained on MS-eluted negatives against assayed positives inherits this, including any model whose training corpus was built the same way.

mhcmatch.posbayes.ANCHORS = (0, 1, 2, -2, -1)#

Anchor positions, signed, matching mhcmatch.immuno.ANCHOR_SCHEMES "pockets".

mhcmatch.posbayes.HUMAN = {'anchor': (-0.17202, 0.0, -0.338211, -0.530797, -0.105505, 0.141961, -0.443731, 0.08185, -0.122986, 0.253121, 0.554954, 0.130571, -0.035569, -0.204233, -0.306762, -0.088272, 0.060235, 0.157231, 0.136212, 0.063724), 'n': 464310, 'n_immunogenic': 14712, 'prevalence': 0.031685727208115265, 'tcrface': (0.126897, 0.0, -0.127745, -0.396513, 0.312013, 0.054715, -0.289776, -0.086445, -0.38218, 0.076293, 0.555381, 0.140134, -0.161836, -0.391744, 0.13859, -0.093884, 0.135715, -0.059949, 0.806245, 0.356368)}#

Fitted on 464,310 rows / 14,712 immunogenic (human host, MHC-I, 8-11mers).

mhcmatch.posbayes.MOUSE = {'anchor': (-0.009639, 0.0, -0.711281, -1.067674, 0.107452, -0.005274, -0.350006, 0.130257, -0.255666, 0.159464, 0.843716, -0.215143, 0.305866, -0.425195, -0.243769, -0.014672, -0.024961, -0.03003, 0.604519, 0.234526), 'n': 47203, 'n_immunogenic': 5154, 'prevalence': 0.10918797534055039, 'tcrface': (0.11129, 0.0, -0.099216, -0.295732, 0.495405, 6.9e-05, -0.295119, 0.1245, -0.652293, -0.009727, 0.870243, 0.093684, -0.29161, -0.502684, -0.07, -0.05781, 0.021781, -0.060342, 0.841119, 0.443017)}#

Fitted on 47,203 rows / 5,154 immunogenic (mouse host, MHC-I, 8-11mers).

mhcmatch.posbayes.llr(peptide, species='human')[source]#

Log-likelihood ratio of immunogenic vs non-immunogenic. Carries no prior.

Larger is more immunogenic. Non-standard residues are skipped; cysteine contributes 0 by construction (see the module warning).

Parameters:
  • peptide (str)

  • species (str)

Return type:

float

mhcmatch.posbayes.posterior(peptide, prior, species='human')[source]#

P(immunogenic | peptide) at an explicit prior. Exact, because llr() has none.

prior is not optional and has no default on purpose. The training corpus runs at ~3.2% positives; a viral proteome scan is nearer 3.0e-3 and the NCI screen 4.8e-4, so a default would silently pick one setting’s base rate for every caller.

Parameters:
  • peptide (str)

  • prior (float)

  • species (str)

Return type:

float

mhcmatch.posbayes.roles(length)[source]#

1 at anchor positions, 0 at TCR-facing ones, for a peptide of this length.

Parameters:

length (int)

Return type:

list[int]

mhcmatch.posbayes.table(species='human')[source]#

The fitted tables for "human" or "mouse".

Parameters:

species (str)

Return type:

dict

mhcmatch.luksza module#

The Łuksza recognition term \(R = Z/(1+Z)\) – a soft partition function over near-matches, replacing a hard distance cut, so how many near-matches a candidate has and how near they are both enter. viral_R is not a term of the shipped aggregate — the sum lived only in the benchmark repo, which made mhcmatch.rank.aggregate_score() callable with a feature no installed user could supply. EPIC does not score with it; the quantity is still the published recognition term and is still computed here. k and a0 are read from the shipped artifact when one carries them, so a refit needs no code change.

The Łuksza recognition term \(R = Z/(1+Z)\), computable without the benchmark repo.

The shipped EPIC aggregate does not score with this term (see SHAPE). It ships because the quantity is the published recognition term, and a refit or a head-to-head against the Łuksza model needs it computable from an installed library rather than only inside the benchmark repository.

The model. A soft partition function over near-matches, replacing a hard distance cut (Balachandran, Łuksza et al., Nature 551:512–516, 2017, doi:10.1038/nature24462; Łuksza et al., Nature 606:389–395, 2022, doi:10.1038/s41586-022-04735-9):

\[Z = \sum_e e^{-k(a_0 - a_e)}, \qquad R = \frac{Z}{1 + Z}\]

a_e is the alignment score against reference epitope e, k an inverse temperature and a0 a horizontal offset. R saturates at 1 when some reference is close and decays smoothly otherwise, so how many near-matches there are and how near they are both enter — which a distance cut discards.

The alignment score is identities, not Smith–Waterman, matching the fit: for a peptide of length L at d substitutions a_e = L - d, so Z collapses to a sum over distances weighted by how many reference windows sit at each. That is cruder than a gapped alignment and monotone in the same direction, and the masked-Hamming search makes it exact rather than approximate.

``k`` and ``a0`` are read from the shipped artifact, not hardcoded. They were fitted by profile likelihood on our own corpus rather than carried over from the papers, which estimated them on a different reference set, a different alignment and a different cohort.

Warning

``R`` is only comparable against the reference set it was fitted with. The shipped standardizer has sigma = 3.8e-8 because the sum saturates near zero for almost every peptide, so an R computed against a different viral ligandome is not on the same axis as the fitted coefficient. viral_r() therefore defaults to the same reference and radius the fit used. Supply a different one only if you are refitting.

Where the time goes, measured rather than assumed. End to end against the 57,331-peptide viral ligandome at radius 4, on 20,000 peptides: 57,000 peptides/s, of which the seqtree neighbour search is 98.6 %, counts_by_distance() 1.3 % and r_term() 0.1 %. The two functions here are already an order of magnitude faster than the search that feeds them, so vectorising them further optimises 1.4 % of the run – do not bother. If this path ever needs to be faster, the search is the thing to attack.

mhcmatch.luksza.r_term(counts, lengths, k=None, a0=None)[source]#

R = Z/(1+Z) from per-distance reference counts.

counts[i, d] is how many reference windows sit at exactly d substitutions from peptide i; lengths[i] is its length. k/a0 default to the shipped fitted shape.

The exponent is clipped to [-60, 60]: exp(-k(a0 - a)) overflows for small a0 and large k on a peptide with thousands of near-matches, and at 60 the term is already far past saturating R.

>>> import numpy as np
>>> r = r_term(np.array([[0, 0], [5, 0]]), np.array([9, 9]), k=1.0, a0=9.0)
>>> bool(r[1] > r[0])
True
Parameters:
  • k (float | None)

  • a0 (float | None)

Return type:

numpy.ndarray

mhcmatch.luksza.counts_by_distance(peptides, hits, category, max_subs=4)[source]#

(counts, lengths) from a mhcmatch.mimics.neighbours() result.

hits is {peptide: {category: [(n_subs, ref_peptide), ...]}}. Distances above max_subs are dropped rather than folded into the last bin, because Z weights them by exp and a fold-in would silently inflate the tail.

Parameters:
  • hits (dict)

  • category (str)

  • max_subs (int)

mhcmatch.luksza.viral_r(peptides, ref_sets=None, *, max_subs=4, k=None, a0=None, threads=1, category='viral')[source]#

viral_R end to end – the published recognition term, computed standalone.

Searches peptides against the viral ligandome with mhcmatch.mimics.neighbours() and turns the per-distance counts into R. With ref_sets=None it loads the same default reference the coefficient was fitted against. The shipped aggregate does not consume this column – mhcmatch.rank.aggregate_score() scores EPIC, which has no viral_R feature.

>>> from mhcmatch import luksza
>>> luksza.viral_r(["NLVPMVATV", "GILGFVFTL"])
array([...])
Parameters:
  • max_subs (int)

  • k (float | None)

  • a0 (float | None)

  • threads (int)

  • category (str)

Return type:

numpy.ndarray

mhcmatch.luksza.shape(artifact=None)[source]#

(k, a0) — the values the viral_R coefficient was fitted with.

Reads an artifact’s luksza block when one carries it, so a vendored refit still overrides; otherwise returns SHAPE.

Parameters:

artifact (dict | None)

Return type:

tuple

mhcmatch.mimicry module#

Mimicry as a signed, per-component immune-response risk. Three references — viral (priming), self (tolerance, and the autoimmunity read-out) and thymus (negative selection) — each split into an anchor and a TCR-facing channel, because a whole-peptide distance averages two different measurements. Scores are log-odds; mhcmatch.mimicry.probability() is a separate step that requires a named corpus, because the screens behind any calibration run from 0.048 % to 46.8 % positive and an unqualified probability is mostly a statement about which intercept was used.

The tested-neoantigen database is exposed as mhcmatch.mimicry.annotate() — prior evidence, never a fitted term. Every labelled screen we hold sits inside that database, so a coefficient on it would be memorisation; held out of the fit, fuzzy matching at two substitutions still recovers 0.08–0.34 of a fresh screen’s positives against 0.00–0.26 for exact lookup, which is what makes it worth reporting. Command line: mhcmatch neoag.

Mimicry as a signed, per-component immune-response risk.

Three reference sets, each answering a different question, and never summed into one “similarity” number (mhcmatch.mimics.KINDS makes the same point about the raw scan):

viral a foreign presented ligandome. A hit says a pre-existing anti-pathogen repertoire may

cross-react, which raises expected immunogenicity.

thymus the thymic self-immunopeptidome. A hit says reactive precursors met the peptide during

negative selection.

self the host proteome. The same tolerance argument without the presentation guarantee, and

simultaneously the autoimmunity read-out: whatever cross-reactive clones survived selection are the ones that would attack the tissue displaying the mimic.

Every component is split into two channels, because a whole-peptide distance averages two different measurements. The anchor channel counts substitutions only over mhcmatch.complement.ANCHORS; the tcr channel counts them over the complement. The two partition the peptide, so no position is weighted twice.

There are two conditionings with two different sign patterns, and they must not be confused. The shipped coefficients are the first one.

Standalone – this artifact, bench/results/mimicry_model.md: the six columns plus screen indicators and nothing else – the sign follows the reference, the way the design predicts. viral is positive on both channels (+0.60 anchor, +0.44 tcr), self is negative on both (-0.30, -0.46): priming and tolerance respectively. thymus is positive on the anchor channel (+0.37) and unresolved on the TCR channel (+0.08, |z| = 1.1).

Residual to a model that already contains a whole-peptide physicochemical term and a foreignness term – bench/results/mimicry_residual.md – a different pattern appears: across all four references tried, anchor-restricted similarity is positive and TCR-face-restricted similarity is negative, with whole-peptide similarity between them and near zero. That is a statement about what mimicry adds to those terms, not about mimicry on its own, and quoting the second pattern as though it were the first is a mistake this paragraph exists to prevent.

Mechanistically the channels are different questions either way, which is why they are kept apart. Anchor similarity to a presented reference is largely presentation – the peptide carries an anchor motif that reference’s alleles present – and it correlates with the binder score (r = +0.25 to +0.33). TCR-face similarity correlates with nothing in the binding stack (|r| < 0.11 against presentation and affinity) but strongly with the whole-peptide physicochemical log-odds (r = +0.73 to +0.82; the row count behind that range was not recorded alongside it), which is precisely why its sign moves once that term enters the model.

That earlier arrangement keeps its recorded coefficients: the letter V names the generation, not the module.

Scores are log-odds, calibration is separate and explicit. score() returns signed contributions and their sum on the log-odds scale, which is corpus-free. probability() maps that sum to a risk of immune response against a named fitted corpus, because an absolute probability is a property of the corpus’s prevalence and candidate generation, not of the peptide. Callers who want a number in [0, 1] should say which corpus they mean.

The tested-neoantigen database is an annotation, never a fitted term. annotate() reports the nearest validated-immunogenic neoantigen and its distance, and that is all it does. Every labelled screen we hold is inside that database – retrieval recall at exact match is 1.000 on all seven – so a fitted coefficient on it would be memorisation. Held out properly it still earns its place as prior evidence: rebuilt without the test screen, fuzzy matching at two substitutions recovers 0.08-0.34 of a screen’s positives where exact lookup recovers 0.00-0.26.

from mhcmatch import mimicry s = mimicry.score([“GILGFVFTL”], refs) # per-component log-odds + aggregate p = mimicry.probability(s, corpus=”screens”) # optional, named a = mimicry.annotate([“GILGFVFTL”], refs) # prior evidence, outside the model

mhcmatch.mimicry.COMPONENTS = ('viral', 'self', 'thymus')#

Reference categories entering the fitted aggregate, in feature order. neoag is deliberately absent – see annotate().

mhcmatch.mimicry.CHANNELS = ('anchor', 'tcr')#

The two signed channels every component is split into.

mhcmatch.mimicry.params(cls='mhc1')[source]#

The frozen model: standardizer, coefficients with posterior sds, the radius per channel, the reference window totals behind the density normalisation, and the fit’s provenance.

Loaded on first use rather than at import, because annotate() and masks() are useful without a fitted aggregate and must not be blocked by its absence.

Parameters:

cls (str)

Return type:

dict

class mhcmatch.mimicry.MimicryScore(peptide, components, logodds, autoimmune, density=<factory>, nearest=<factory>)[source]#

Bases: object

One peptide’s mimicry read-out.

components is {component: {channel: signed log-odds contribution}}; logodds is their sum. autoimmune is the self component’s total, reported separately because a self mimic is simultaneously a tolerance argument and a cross-reactivity liability for a vaccine, and those two license different decisions.

Parameters:
  • peptide (str)

  • components (dict[str, dict[str, float]])

  • logodds (float)

  • autoimmune (float)

  • density (dict[str, float])

  • nearest (dict)

peptide: str#
components: dict[str, dict[str, float]]#
logodds: float#
autoimmune: float#
density: dict[str, float]#
nearest: dict#

{component: {channel: {"subs", "peptide", "source", "n"}}} – which reference peptide was hit and what protein it came from. Without this the self/thymus channels are a bare number and safety() cannot be reached, which is the question a vaccine needs answered.

as_dict()[source]#

Flat {column: value} – one {component}_{channel} key per score, plus the nearest reference peptide/source/substitution-count for each, ready for a TSV row.

Return type:

dict

mhcmatch.mimicry.masks(length, cls='mhc1', peptide=None, register=None)[source]#

Positions each channel counts substitutions over, for one peptide.

anchor is the MHC-facing set and tcr is its complement, so the two channels partition the peptide and no position is counted twice.

The two classes take their anchors from different places, and the class is never inferred. Class I is anchored at fixed peptide positions (mhcmatch.complement.ANCHORS, the same five the shipped role model calls MHC-facing), so length alone determines the split. A class-II ligand is anchored by a 9-mer core that floats inside an 11–25-mer, so its split is a function of the register, not the length: pass cls="mhc2" with the peptide, and optionally a register to pin the frame instead of using the allele-agnostic heuristic (mhcmatch.complement.mhc2_anchors()). Reading a class-II ligand on the class-I layout labels the wrong residues as anchors and returns a confident, wrong face.

>>> masks(9)["anchor"]
[0, 1, 2, 7, 8]
>>> masks(15, "mhc2", "AAAKFVAAWTLKAAA")["anchor"]      # P1/P4/P6/P9 of the floating core
[4, 7, 9, 12]
Parameters:
  • length (int)

  • cls (str)

  • peptide (str | None)

  • register (int | None)

Return type:

dict[str, list[int]]

mhcmatch.mimicry.features(peptides, refs, cls='mhc1', *, threads=1)[source]#

Per-(component, channel) mimic density: hits per million same-length reference windows.

Density and not a raw count because the window totals span three orders of magnitude across components, channels and lengths, so a count standardized across that mix is largely measuring which length the peptide is. log1p because the counts are heavy-tailed by construction.

The radius per channel is the fitted one (params()): the anchor channel is searched wider than the TCR channel because it projects onto more positions and so saturates later.

Batched per (component, channel, length), not per peptide. The indices are already built per that key by load_references(), so a per-peptide search() asked the same index six questions per peptide – one C++ entry per question, GIL held between them – where search_batch asks it once for every peptide of that length with the GIL released. Same index, same radius, same hits; the loop that was there is the one rule 2 forbids.

Parameters:
  • refs (dict)

  • cls (str)

  • threads (int)

Return type:

list[dict]

mhcmatch.mimicry.corpus_R(peptides, spectrum, cls='mhc1', registers=None)[source]#

The Luksza corpus density per component, exactly – one table lookup per query window.

spectrum comes from corpus_spectrum(). Returns {component: rho} per peptide, where

\[\rho_k(q) \;=\; \frac{1}{m_k(q)\,N_k}\sum_{i=0}^{m_k(q)-1}\; \sum_{x} T_k[x]\,\beta^{\,d_H(f(q)[i:i+k],\,x)}\]

with \(f(q)\) the TCR face, \(m_k(q)\) its sliding k-mer count, \(T_k\) the reference k-mer frequency table, \(N_k=\sum_x T_k[x]\) its total mass and \(\beta=e^{-\kappa}\). The inner sum runs over the whole reference set – there is no radius and no k-nearest cutoff, because \(\beta^{d}\) is the threshold. It is evaluated as a table lookup because the weight factorises over positions (corpus_spectrum()).

What the two divisors are for. \(N_k\) makes the value a density, so thymus (26,513 peptides) and self (12 M proteome windows) are on one scale and the standing claim that thymus makes the other two redundant is a comparison rather than a size effect. \(m_k(q)\) makes it per query window, which is what removes the length artefact: the old fixed-face column varied 17x in mean across lengths 8-11 and correlated with length at Spearman -0.502, against +0.036 here (3,600 real epitopes, 900 per length, none in the reference).

So \(\rho\in[0,1]\) is the expected mismatch weight between a uniformly chosen query window and a uniformly chosen reference window. The Luksza \(Z/(1+Z)\) saturation is dropped as redundant: it existed to bound an unbounded count, \(\rho\) is already bounded, and the old column never left its linear regime anyway (Z stayed below 1.32e-3, so R was Z to three figures). \(a_0\) is gone with it – it was a scale the standardizer absorbed, and the length compensation exp(kappa*(L - a0)) it carried is now the explicit \(m_k\) divisor.

A peptide whose face cannot supply a k-mer yields no key, exactly as an unindexed length did before. With the shipped ``k = 3`` this cannot happen for a canonical class-I ligand: the face is L - 5 wide and the shortest ligand is an 8-mer. It fires only for a non-standard residue.

>>> spec = corpus_spectrum(components=("thymus",))
>>> corpus_R(["GILGFVFTL"], spec)[0]["thymus"]
Parameters:
  • spectrum (dict)

  • cls (str)

Return type:

list[dict]

mhcmatch.mimicry.corpus_counts(pmhc_dir=None, cls='mhc1', comp='thymus', k=3, self_species='human', weights=None, mask='slice')[source]#

(T, N): the sliding-k-mer count table over one reference component’s TCR faces.

mask must match the one the queries will use (CORPUS_MASKS); it selects the same face construction on the reference side, so the two are comparable by construction.

T is a flat A**k array of window counts with multiplicity – one increment per reference peptide per window, the published Luksza form, not per distinct face – and N = T.sum() the total reference window mass. Memoised per (cls, comp, k, species, weights) for the process; see _COUNTS.

weights="locus" divides each reference peptide’s contribution by the size of its locus_weights() component, so a protein region that appears many times over – a recurrent hotspot, a family of overlapping registers – counts once instead of many times. The density stays in [0, 1] because N is the same weighted mass. Off by default: it changes what the column means, so it is an arm to measure rather than a correction to assume.

It applies to the assayed deposits (``thymus``, ``viral``) and is a no-op for ``self``, and that is the semantics rather than an optimisation. Locus weighting corrects a sampling bias: a peptide deposit over-represents whatever was detected or tested most, so one region can enter it many times. A proteome is not a sample – every window appears exactly once by construction, there is no “assayed more often” to undo, and sliding windows overlap their neighbours by definition, so a sequence-overlap grouping over 122 M of them would collapse each protein (and every homolog of it) into one degenerate component at ruinous cost.

Every length the class admits contributes, which is the point of sliding: a 9-mer reference informs an 11-mer query because the table is keyed on the k-mer, not on the length. A reference whose face is narrower than k contributes nothing and is skipped rather than padded.

Parameters:
  • cls (str)

  • comp (str)

  • k (int)

  • self_species (str)

  • weights (str | None)

  • mask (str)

mhcmatch.mimicry.contract(T, kappa, k=3, kernel=None)[source]#

Apply the mismatch kernel along every axis of an A**k count table. One-time, ~1 ms.

A is inferred from len(T) (alphabet()), so a wildcard-masked 21**k table works unchanged. kernel defaults to the Hamming form K = (1-beta)I + beta*J with beta = exp(-kappa), kept only so older results stay reproducible; pass blosum62_kernel() for the graded score. Any position-additive, ungapped score factorises this way; a gapped alignment does not, which is the one real limit.

Parameters:
  • kappa (float)

  • k (int)

mhcmatch.mimicry.corpus_spectrum(pmhc_dir=None, cls='mhc1', components=None, k=3, shapes=None, self_species='human', weights=None, mask='slice', kernel=None)[source]#

Contracted sliding-k-mer tables over the TCR face, one per component. No search.

Returns {component: (table, n_kmers, k, mask)} where table is a flat A**k array and n_kmers the total reference window count. Feed it to corpus_R(), which reads the mask back off the tuple so a query cannot be posed against a table built the other way.

kernel is passed to contract(). None keeps the original Hamming form; pass a callable kappa -> (A, A) array – blosum62_kernel() partially applied, or functools.partial(blosum62_kernel, mask=mask) – to grade the substitutions, since each component carries its own kappa.

Why there is no index here. The Luksza sum Z = sum_r exp(-kappa*(a0 - s(q,r))) weights every reference by its similarity, and with an ungapped position-additive score the weight factorises:

beta**hamming(u, x) = prod_p K[u_p, x_p],    K = (1-beta)*I + beta*J,   beta = exp(-kappa)

so the sum over the whole reference set is one 20x20 matrix applied along each axis of the k-mer frequency table – a tensor contraction, computed once. Every query is then a table lookup. That is exact where a radius-capped neighbour search is a truncation: measured against a brute-force all-vs-all over every reference k-mer, the contraction agrees to 5.6e-16, while the radius-2 search it replaced captured a median 0.4999 of the true sum (IQR 0.4115-0.5556, min 0.1539 over 600 real 9-mers). Cost went from ~46 s to 2.3 ms for 340,876 queries, and the ~7.5 GB proteome index the self channel needed became a 64 KB table.

The two halves are split on purpose: corpus_counts() builds the table (expensive, and memoised) and contract() applies kappa (~1 ms), so profiling kappa costs one build, not one per grid point.

shapes supplies kappa per component (corpus_shapes()); a0 is not used and does not need to be, because the length compensation it stood in for is now done explicitly by normalizing per query window (see corpus_R()).

Species. self_species keys every component, not just self, and this function honours it literally – pass "mouse" and you get the mouse tables, which is what makes the substitution measurable. But the SCORER does not call it that way: it resolves each component through reference_species() first, and that map sends every mouse component to "human", in both classes. So a mouse run is scored against the human tables throughout, not the mouse ones, and this docstring once said the opposite.

The mouse tables remain reachable and are the arm that measures the difference: they stand on 25,264 and 40,244 reference windows against 140,482 and 136,618 for class I. See reference_species() for the per-component transfer measurements and for why thinness is not the explanation.

Parameters:
  • cls (str)

  • k (int)

  • shapes (dict | None)

  • self_species (str)

  • weights (str | None)

  • mask (str)

Return type:

dict

mhcmatch.mimicry.face_kmers(peptide, cls='mhc1', k=3, register=None, mask='slice')[source]#

Packed sliding k-mers of the peptide’s TCR face; empty when it cannot carry one.

The class-I anchor set is {P1, P2, P3, POmega-1, POmega}, so under mask="slice" the TCR face is contiguous – peptide[3:L-2], width L - 5 – at every length; class II gathers its face from around the floating core instead, and the k-mers slide over that projection. Sliding rather than taking the whole face is what lets a query of one length be compared against references of another: the table is keyed on the k-mer, not on the length.

Under mask="wildcard" the anchors are replaced by WILDCARD in place rather than deleted, so the window count is L - k + 1 and a k-mer may span a pocket. See CORPUS_MASKS for why that changes which k are admissible.

The packing base follows the mask (alphabet()), so a table built under one mask cannot be indexed with the other – the sizes differ and corpus_R() checks it.

>>> face_kmers("SIINFEKL").size                          # W = 3, exactly one 3-mer
1
>>> face_kmers("SIINFEKL", mask="wildcard").size         # 8 - 3 + 1
6
Parameters:
  • peptide (str)

  • cls (str)

  • k (int)

  • register (int | None)

  • mask (str)

mhcmatch.mimicry.SHAPES: dict = {'self': 5.0, 'thymus': 3.0, 'viral': 8.0}#

Fitted decay kappa per component, by profile deviance on the neoantigen corpus (bench/immuno/corpus_exact.py). Passed as shapes= only to re-measure it. The shipped aggregate records the same values as corpus_shapes; corpus_shapes() reads that copy so a re-vendored refit moves the scored column with it, and this one is the fallback.

``a0`` is gone, and the value is one number rather than a pair. It was a scale the standardizer absorbed, and the length compensation exp(kappa*(L - a0)) it carried is now the explicit per-window divisor in corpus_R().

The three differ, and they differ for a reason. Writing gamma = (1-beta)/(1+19*beta) with beta = exp(-kappa) for the fraction of order-1 sequence structure the contraction keeps (contract()), thymus sits at an interior optimum with gamma = 0.49 – graded tolerance, near-misses count – while viral runs to gamma = 0.99, near-exact 3-mer matching. self carries no shape information at all: its profile moves 0.12 deviance units across a 24-fold range of kappa, because the human proteome occupies essentially every cell of the 3-mer table, so smoothing cannot change the ordering. That is the mechanism behind self reading as how many rather than whether.

mhcmatch.mimicry.CORPUS_K: int = 3#

Width of the sliding k-mer over the TCR face. 3 is not a tuning choice, it is the only width every class-I length can supply: the face is L - 5 residues wide (the anchors are P1-P3 and POmega-1/POmega) and the shortest ligand is an 8-mer, so W_min = 3. At k = 4 an 8-mer has no window at all and at k = 5 neither an 8- nor a 9-mer does – which reads as a low score and is really a structural zero, the exact failure mode this module removed. Measured on 3,600 real epitopes balanced 900 per length, none of them in the thymic reference, the normalized density correlates with peptide length at Spearman +0.036 for k=3 against +0.587 (k=4) and +0.830 (k=5) – and against -0.502 for the fixed-face column this replaced.

mhcmatch.mimicry.LOCUS_W: int = 7#

Substring width at which two peptides are called the same locus by locus_weights(). Seven residues links every register of one mutation – VVVGAVGVGK, VVGAVGVGK and GADGVGKSAL are one KRAS G12 locus at 7 and three separate peptides at 11 – while being long enough that an incidental match between unrelated proteins is rare: there are 20**7 = 1.28e9 7-mers against ~1.2e7 windows in a whole proteome.

mhcmatch.mimicry.locus_weights(peptides, w=7)[source]#

Per-peptide weights that make each locus count once, not each peptide.

Two peptides are the same locus when they share a substring of w residues; the relation is closed transitively, and every peptide in a component of size m gets weight 1/m, so the component contributes unit mass however many peptides it supplied.

Why this exists. A peptide set is not a set of independent observations. One recurrent hotspot – KRAS G12, which is both the most-tested and the most-validated public neoantigen – contributes many overlapping registers of the same eleven residues, and an unweighted k-mer table reads that as evidence about immunogenicity when it is evidence about what got assayed. Measured on the immunogenic arm of the neoantigen fit corpus, the effect is large enough to dominate the set’s most distinctive 3-mers (VGA, GAV, AVG – all fragments of VVVGA[VDC]GVGKSA). See bench/results/kmer_spectrum.md.

Grouping is by sequence rather than by gene symbol because the corpora that need it do not all carry one, and because the register-level overlap is what actually repeats. Where a real source annotation exists, pass its own weights to corpus_counts() instead.

>>> [round(x, 3) for x in locus_weights(["VVVGAVGVGK", "VVGAVGVGK", "SIINFEKLAA"])]
[0.5, 0.5, 1.0]
Parameters:

w (int)

Return type:

list

mhcmatch.mimicry.corpus_shapes(artifact=None)[source]#

{component: kappa} – the decay C_corpus was fitted with.

Reads the shipped aggregate’s corpus_shapes when it carries one, so a refit that is re-vendored moves the scored column rather than leaving it on a stale module constant; otherwise returns SHAPES. Same convention as mhcmatch.luksza.shape().

Accepts the older (kappa, a0) pair form for reading an old artifact and keeps only kappa; nothing writes that form any more.

Parameters:

artifact (dict | None)

Return type:

dict

mhcmatch.mimicry.score(peptides, refs, cls='mhc1', allow_missing=False, *, threads=1)[source]#

Signed per-component log-odds contributions and their sum, one per peptide.

Raises if refs is missing a component. A missing feature standardizes to zero, so the aggregate would silently be a different, smaller model rather than an error – and the usual way to get here is load_references(with_self=False), which drops the component carrying the largest coefficients. Pass allow_missing=True to accept that deliberately — its channels then come back NaN, and so do MimicryScore.logodds and MimicryScore.autoimmune, because the fitted coefficients describe the full set. A zero on autoimmune reads as “no self-similarity found”; the truth under with_self=False is “never looked”, and those license different decisions.

Parameters:
  • refs (dict)

  • cls (str)

  • allow_missing (bool)

  • threads (int)

Return type:

list[MimicryScore]

mhcmatch.mimicry.probability(scores, corpus='screens', cls='mhc1')[source]#

Map the aggregate log-odds to a risk of immune response against a named corpus.

The intercept is the corpus’s own base rate, so this number is not transferable: the seven neoantigen screens behind "screens" run from 0.048 % positive to 46.8 %, and a probability quoted without naming the corpus is quoting one of those prevalences by accident. Use MimicryScore.logodds to rank; use this only to report.

Parameters:
  • corpus (str)

  • cls (str)

Return type:

list[float]

mhcmatch.mimicry.annotate(peptides, pmhc_dir=None, cls='mhc1', max_subs=2, *, threads=1)[source]#

Nearest validated-immunogenic neoantigen and its distance. Prior evidence, not a score.

This is kept out of score() on purpose. Every labelled screen we hold is contained in the tested-neoantigen database, so retrieval recall at distance 0 is 1.000 on all seven and a fitted coefficient would be measuring the answer key. Held out honestly the channel is still useful – with the test screen removed from the database, matching at two substitutions recovers 0.08-0.34 of its positives against 0.00-0.26 for exact lookup – which is why it is reported at all.

Parameters:
  • cls (str)

  • max_subs (int)

  • threads (int)

Return type:

list[dict]

mhcmatch.mimicry.NEOAG_COLUMNS: tuple = ('neoag_distance', 'neoag_nearest', 'neoag_n_within', 'known')#

The columns annotate() appends, in order, after peptide. It lives here rather than inline in the CLI or in a pipeline stub for the same reason mhcmatch.rank.BASE_COLUMNS does: a consumer has to be able to name the schema without running the command, and a schema typed out a second time is a schema that drifts.

mhcmatch.mimicry.load_references(pmhc_dir=None, cls='mhc1', with_self=True, self_species='human', lengths=None)[source]#

Reference window sets per (component, channel, length), ready for features().

with_self=False skips the host proteome, which dominates the cost. The aggregate is not defined without self – it carries the largest coefficients in the fit – so score() raises unless the caller passes allow_missing, and mhcmatch rank --score aggregate refuses the combination outright.

``lengths`` is the difference between usable and not at class II. The index is per-length and the cost is per-length: one proteome pass is ~11 s for window_array plus ~1.0 min to resolve a class-II register for each of its 12,685,964 windows. The class admits fifteen lengths (11-25), so building all of them is ~19 min; class I admits four, at ~65 s. A run queries the lengths its own peptides have – usually one or two – and building the other thirteen is work nobody asked for. Pass the queried lengths and pay for those: ~75 s for one class-II length, ~17 s for one class-I length.

None (the default) keeps every length the class admits, so an existing caller is unchanged. A length outside mimics._LEN[cls] was never indexed and still is not; features() skips a query it has no index for.

The build itself is vectorized: mhcmatch.proteome.Proteome.window_array() replaces a per-window Python loop (2.7x), and the per-channel projection is one np.unique over a fixed-width byte view rather than a setdefault over 12 M strings.

Parameters:
  • cls (str)

  • with_self (bool)

  • self_species (str)

Return type:

dict

mhcmatch.mimicry.safety(scores, top=5, symbols=None)[source]#

Where the self/thymus mimics are expressed – the autoimmunity read-out, made actionable.

There is no tumour argument, and there used to be one that did nothing. tumor sat at positional #2 and was never read – mhcmatch.expression.safety_profile() conditions on no context at all – so safety(scores, "SKCM") returned the pooled profile while reading as if it had been conditioned. On a read-out whose job is to say which tissue you cannot afford to damage, a caller believing they narrowed the question is the dangerous direction. Removed rather than accepted-and-ignored; add it back only when safety_profile can actually take one.

A self or thymus hit says the candidate resembles a peptide the body presents; the decision it feeds is whether the tissue presenting it is one you can afford to damage. That needs the mimic’s source protein, which is why load_references() carries it and features() keeps it rather than collapsing everything to a density.

Returns, per peptide, the tolerance-side hits with mhcmatch.expression.safety_profile() resolved for each source that maps to a gene.

The ``thymus`` deposit names its sources as UniProt accessions (Q8WZ42) while mhcmatch.expression.safety_profile() is keyed on HGNC symbols (TTN), so the join needs a map: pass symbols= from mhcmatch.proteome.gene_symbols(path, key="accession")(). Without it an accession simply fails to resolve and profile comes back empty – which reads as no risk found and is the dangerous direction to be wrong in, so mhcmatch.vector.self_origin_risk() refuses to run without the map rather than defaulting it.

One gap remains and it returns an empty ``profile`` rather than a guess. The self component is built from proteome windows, which carry no source column at all, so it is not resolvable to a gene however good the map is. The mimic peptide and its raw source are always returned.

gene is the resolved symbol and source stays the deposit’s own identifier, because a withdrawal decision has to be traceable back to the record it came from.

Parameters:

top (int)

Return type:

list[dict]

mhcmatch.mimicry.CORPUS_REFERENCE: dict = {('mouse', 'self'): 'human', ('mouse', 'thymus'): 'human', ('mouse', 'viral'): 'human'}#

(query species, component) -> reference species, for the components where the two differ. Read by reference_species(), which is the only thing that should consult it, and applied by cli._aggregate_channels when it builds the scoring channels.

One entry: a mouse ``thymus`` query is scored against the HUMAN thymic table. The mouse thymic deposit is one haplotype. Of its 6,661 class-I peptides, every one of the 2,663 that carries an allele annotation is H-2Db (1,574) or H-2Kb (1,089) – no other haplotype appears at all – so its k-mer table is the H-2b groove’s motif rather than a measure of what the thymus presents. That column is then applied to a fit spanning six H-2 allotypes, and it is collinear with binder, which already scores the groove.

The mouse class-I corpus block reads human tables throughout – see reference_species() for the per-component measurement behind one uniform rule.

mhcmatch.mimicry.NATIVE_CORPUS_COMPONENTS: tuple = ('self', 'thymus')#

The host compartments, and the two components native=True will route back to the query species’ own table. viral is deliberately not one of them: it is not a host compartment, so “the mouse’s own” makes no sense for it in the way it does for a proteome or a thymus. What a mouse viral table samples is 9 H-2 allotypes against the human table’s 129 – a thinner sample of the same pathogen ligandome, not a different compartment – so there is nothing to recover by switching, and CORPUS_REFERENCE keeps it human even under native.

mhcmatch.mimicry.reference_species(species, comp, native=False)[source]#

Which species’ reference table a species query of component comp is scored against.

native=True overrides the redirect below for the host components named in NATIVE_CORPUS_COMPONENTS and returns species itself, so a mouse query is matched against the mouse tables. It is not the default, and the reason is measured, not assumed – see the per-component transfer figures below, and in particular thymus at r = 0.3245 with the matched-mass arm that rules out thinness as the explanation. mhcmatch rank –native-corpus warns on every run that uses it. All twelve tables ({mhc1,mhc2}|{self,thymus,viral}|{human,mouse}|3) ship, so this is a routing switch and nothing is fetched or rebuilt.

A coefficient fitted under one routing does not transfer to the other. Every shipped mouse artifact was fitted with the human tables, so --native-corpus scores those coefficients against a different feature. Use it to measure the mouse tables, not to rank against a fit that never saw them.

Usually species itself. The exception is mouse class I, whose whole corpus block – ``thymus``, ``self`` and ``viral`` alike – is matched against the human tables, because the mouse reference deposits are too small and too groove-skewed to be a reference.

Nothing is trained on human data by this. A corpus channel is a k-mer density lookup: the query peptide’s TCR-facing windows are matched against a counted reference table, and the table is the only thing that is human. Every coefficient in aggregate_mhc1_mouse.json is fitted on mouse neoantigens, and presentation, expression and physicochemistry all read mouse sources. What crosses the species line is the reference corpus being matched to, nothing else.

The corpus channel transfers to the extent that its reference deposit is not one groove’s motif, and the three components sit at three points on that axis. Measured on the 921-row mouse class-I fit population (bench/epic/corpus_transfer.py in the benchmark repo), r being the Pearson correlation between the same peptide’s density under the two species’ tables:

  • self – no groove at all (an unpresented proteome window is not presented by anything), 112,565,681 mouse against 121,968,158 human reference windows, r = 0.9990. The same table twice over: the substitution is free in either direction and taking human keeps one reference source.

  • viral – 9 mouse allotypes (H-2Kb 50.2 %) against 129 human, r = 0.8382. A 9-allotype sample of a 129-allotype space.

  • thymus – 2 mouse allotypes (H-2Db 1,574, H-2Kb 1,089, nothing else) against a pooled human donor panel, 25,264 against 140,482 windows, r = 0.3245. The H-2b motif, and nothing else.

It is not a sample-size effect, and that was measured rather than assumed. The mouse thymic table stands on 25,264 reference windows against human’s 140,482, so thinness is the obvious explanation and it is the wrong one: thinning the human deposit at the peptide level to the mouse table’s window count, 40 draws, still reproduces the full human column at r = 0.8933 (range 0.8728-0.9109) and still disagrees with the mouse table at 0.2903 (0.2467-0.3310). A human table cut to mouse’s size does not become the mouse table. What differs is which grooves each deposit sampled – so depositing more mouse thymic peptides from the same two allotypes would not close it, and the substitution is the fix rather than a stopgap.

This is the same failure background="ligand" had: a pooled statistic over a pool dominated by one allotype measures that allotype. Mouse class II is the extreme case – its thymic deposit is 1,490 peptides, all I-Ab.

The redirect is not keyed on class, and that matters. This function takes (species, comp) only, so a mouse class-II query is routed to the human class-II tables by the same rule – cls is carried separately and selects which of the twelve tables is read. The point was once moot: both class-II artifacts were the six-term neoantigen fits, which declare no corpus block, so nothing read a class-II corpus table. The two class-II pathogen fits do declare one, and they were fitted under this redirect – C_corpus_self resolves in both (p = 1.1e-03 human, 3.5e-10 mouse) and C_corpus_thymus in neither mouse cell (p = 0.24), which is what the I-Ab-only thymic deposit above predicts.

Expression is NOT covered by this and must not be. Human and mouse organs and tumours are different tissues, so a human expression level is not a stand-in for a mouse one at any sample size; mhcmatch.expression stays species-keyed throughout. Mapping a gene to its mouse orthologue is a different operation – it fixes gene identity and still reads a mouse transcriptome.

Parameters:
  • species (str)

  • comp (str)

  • native (bool)

Return type:

str

mhcmatch.mimics module#

Cross-reactivity by category, never summed: a hit in the thymic immunopeptidome argues tolerance and autoimmune risk, a hit in the host proteome argues peripheral cross-reactivity without implying presentation, and a viral or bacterial hit argues a pre-existing repertoire may cross-react. KINDS states which is which; neighbours() runs the whole query set through one threaded C++ index per (category, length).

Molecular-mimicry annotation for strong binders.

For each strong-binding neoantigen, search reference peptide sets for mimics — near-identical presented peptides — and report the presentation-aware E-value (mhcmatch.search.find_mimics(), lower = more significant mimicry) per category:

  • thymus — the thymic self-immunopeptidome (HLA Ligand Atlas). A significant thymic mimic means the neoantigen resembles a self-peptide presented during negative selection: reactive T cells were likely deleted (reduced immunogenicity) and it flags cross-reactivity / autoimmune risk for a cancer vaccine.

  • self — a window of the host proteome with no evidence of thymic presentation. Encoded is not presented, so the tolerance argument is weaker than thymus; the peripheral cross-reactivity risk is the same. Kept as its own category rather than merged into thymus because those two hits license different conclusions.

  • viral / bacterial — foreign presented peptides / pathogen and commensal proteomes. A foreign mimic can raise immunogenicity (a pre-existing anti-pathogen repertoire cross-reacts) — molecular mimicry.

  • neoag — the tested-neoantigen database: has this (or a near-identical) neoantigen been reported.

KINDS states what each category argues; PROTEOME_REFS names the proteomes behind self and bacterial, which are opt-in (load_reference_sets(..., proteomes=("bacterial",))) because building their window sets is the expensive part.

This scores cross-reactivity, not presentation or immunogenicity directly; compose it with the presentation / affinity scores from mhcmatch.predict. Reference data: the isalgo/pmhc_data compendium (thymus/, ligandome/, immunogenicity/, proteome/).

mhcmatch.mimics.KINDS = {'bacterial': 'foreign', 'neoag': 'database', 'self': 'self', 'self_mouse': 'self', 'thymus': 'self', 'viral': 'foreign'}#

What a hit in each category argues. Not the same question per category, which is why they are never summed into one “mimicry score”:

category

a hit means

thymus

presented in the thymus, so reactive clones met it during negative selection. Lowers expected immunogenicity; raises autoimmune risk for a vaccine.

self

a window of the host proteome not known to be thymically presented. Weaker tolerance argument – being encoded does not imply being presented – but the same peripheral cross-reactivity risk. Kept distinct from thymus on purpose.

viral

a foreign presented peptide. Raises expected immunogenicity: a pre-existing anti-pathogen repertoire may cross-react.

bacterial

the same, from pathogen and commensal proteomes.

neoag

already tested as a neoantigen somewhere. Prior evidence, not biology.

mhcmatch.mimics.DEFAULT_REFS = {'neoag': ('neoantigens/neoag_tested.tsv.gz', 'database'), 'thymus': ('thymus/thymus_immunopeptidome.tsv.gz', 'self'), 'viral': ('ligandome/viral_foreign_iedb.tsv.gz', 'foreign')}#

(folder/file under pmhc_data, kind). self is the tolerance reference passed as find_mimics’ self_set; the rest are foreign/database sets.

Type:

Default reference categories

mhcmatch.mimics.SPECIES_REFS = {('thymus', 'mouse'): 'thymus/thymus_immunopeptidome_mmu.tsv.gz'}#

Per-species overrides of a DEFAULT_REFS path.

The two deposits carry species differently, and the difference is in the data, not a style choice. viral is one file whose own mhc_species column holds both – 2,773 distinct mouse peptides among 44,993 – so it needs no entry here and is selected with species= at load time. thymus is one file per species, because the human deposit is a single multi-donor atlas and the mouse one is assembled from three studies with different antibodies and search pipelines; pooling them into one table would bury that in a column nobody filters on.

Resolve through ref_path() rather than indexing this directly, so a category with no override falls back to DEFAULT_REFS instead of raising.

mhcmatch.mimics.CLS_REFS = {('self', 'mhc2', 'human'): 'ligandome/self_ligandome_mhc2_hsa.tsv.gz', ('self', 'mhc2', 'mouse'): 'ligandome/self_ligandome_mhc2_mmu.tsv.gz', ('thymus', 'mhc2', 'human'): 'ligandome/hla_ligand_atlas_mhc2_presented.tsv.gz', ('viral', 'mhc2', 'human'): 'ligandome/viral_ligandome_mhc2_pan.tsv.gz', ('viral', 'mhc2', 'mouse'): 'ligandome/viral_ligandome_mhc2_pan.tsv.gz'}#

Class-II overrides, keyed ``(category, cls, species)``, consulted before SPECIES_REFS.

Two of the three class-II references are a different object from their class-I namesakes, and both differences are forced by what the reference has to argue:

  • self at class I is the host proteome and stays one – C_corpus_self is a fitted term of two shipped class-I artifacts, so redefining it would be a model change. At class II the proteome saturates the table (8,000 of 8,000 3-mer cells, ~111 M window mass), and a near-uniform density cannot discriminate. The class-II channel therefore reads a self ligandome – IEDB class-II ligands whose source protein is the host’s own, 421,174 human and 25,314 mouse peptides – which is what “presented” means and what tolerance is about.

  • viral is pan-species at class II. The deposit carries no mhc_species column, so load_peptides() keeps every row whichever species asks. Splitting it left the mouse cell on 464 peptides against human’s 14,468; a viral k-mer density is not species-specific the way a proteome is, and pooling deepens the mouse reference 31x while moving human 14,468 -> 14,645.

  • thymus at human class II is the whole displayed proteome, not the thymus – every benign tissue the HLA Ligand Atlas sampled on a class-II molecule, 132,818 peptides against the thymic fraction’s 27,987. The channel was never thymic: measured across the atlas’s 29 tissues on 7,948 human class-II T-cell rows of which 5,150 are immunogenic, lung leads at contrast AUROC 0.5496 (95% CI 0.5365-0.5627) and thymus is second at 0.5475 (0.5345-0.5606) – the same result class I gives, where thymus is 6th of 29. The slot keeps the name thymus because C_corpus_thymus is a released column name and renaming it would misname the model a reader downloads; every reader-facing label says displayed self. Mouse class II is untouched by this and still reads its own thymic deposit.

mhcmatch.mimics.ref_path(category, species='human', refs=None, cls=None)[source]#

The compendium path for category in species, for cls when it matters.

Resolution order is most specific first: CLS_REFS (category, cls, species), then SPECIES_REFS (category, species), then the species-agnostic DEFAULT_REFS. Only a deposit that is physically split needs an entry above the last.

cls=None skips the class layer, which is what a caller with no class in hand wants; it is not the same as cls="mhc1", because a class-I-only override would then be reachable from a call that never named a class.

Parameters:
  • category (str)

  • species (str)

  • refs (dict | None)

  • cls (str | None)

Return type:

str

mhcmatch.mimics.PROTEOME_REFS = {'bacterial': ('ecoli_K12_UP000000625', 'saureus_UP000008816', 'lreuteri_UP000001991', 'mgnavus_UP000018690', 'egallinarum_UP000254807'), 'self': ('human',), 'self_mouse': ('mouse',)}#

Reference proteomes per category, as mhcmatch.store.fetch_proteome() stems. Gut commensals (L. reuteri, M. gnavus, E. gallinarum), a gut/lab strain (E. coli K12) and a skin/nasal pathogen (S. aureus) – the exposures a human repertoire has plausibly seen.

mhcmatch.mimics.CANONICAL_LEN = {'mhc1': range(8, 12), 'mhc2': range(11, 26)}#

Canonical presented lengths per class, and what every shipped number rests on. The class-I groove is closed at both ends, so 8-11 is the range the corpus tables, the anchor models and every fitted artifact were built over.

mhcmatch.mimics.EXTENDED_LEN = {'mhc1': range(12, 15)}#

Extended class-I lengths, 12-14 – real ligands, and deliberately NOT implemented.

MHC I does present past 11 by bulging the peptide out of the groove, and the deposits carry those rows: filtered to mhc_class = MHCI, viral_foreign_iedb.tsv.gz holds 157 peptides of length 12-13 (0.3 % of its 55,084) and thymus_immunopeptidome.tsv.gz 195 (0.8 % of 25,891); the human neoantigen deposit holds 84,622 rows carrying 276 immunogenic peptides outside 8-11. Mouse holds 1.

They are excluded on purpose and the exclusion is held, not fixed: admitting them would rebuild every corpus table, move the human fit population that artifact version 11 was fitted on, and make no published number reproduce. length_range() therefore refuses the extended range rather than returning it – a range no downstream path was fitted for is worse than an error, because it would score silently and wrongly.

mhcmatch.mimics.length_range(cls, extended=False)[source]#

The presented lengths cls admits. extended=True is not implemented and raises.

Parameters:
  • cls (str) – "mhc1" or "mhc2".

  • extended (bool) – ask for the extended class-I range (12-14) instead of the canonical one.

Returns:

a range of peptide lengths.

Raises:
  • NotImplementedError – when extended is true, naming what would have to change.

  • ValueError – when cls is not a known class.

The two ranges are CANONICAL_LEN and EXTENDED_LEN; see the latter for how many real ligands the canonical cut leaves out and why the cut is held anyway.

class mhcmatch.mimics.MimicResult(binder, allele, category, n_exact, n_near, top_mimic, top_subs, e_value, n_hits, significant)[source]#

Bases: object

Per-(binder, category) mimicry summary.

A mimic is a reference peptide of the same length within near_subs substitutions of the binder (T cells cross-react across a few substitutions). n_exact / n_near count identical and near-identical mimics; top_mimic / top_subs are the closest one. e_value / n_hits are the raw presentation-aware search stats, kept for reference.

Parameters:
  • binder (str)

  • allele (str)

  • category (str)

  • n_exact (int)

  • n_near (int)

  • top_mimic (str)

  • top_subs (int)

  • e_value (float)

  • n_hits (int)

  • significant (bool)

binder: str#
allele: str#
category: str#
n_exact: int#
n_near: int#
top_mimic: str#
top_subs: int#
e_value: float#
n_hits: int#
significant: bool#
mhcmatch.mimics.load_peptides(pmhc_dir, rel_path, cls, species='human')[source]#

The peptide column of a compendium TSV, filtered to cls / species and plausible presented lengths. Rows without a class/species field are kept (some sets are unlabelled).

pmhc_dir=None fetches the file from the public HF dataset instead (cached, and overridable with $MHCMATCH_PMHC_DIR), so a fresh install needs no pre-staged mirror.

Parameters:
  • rel_path (str)

  • cls (str)

  • species (str)

Return type:

list

mhcmatch.mimics.proteome_peptides(category, lengths)[source]#

Every distinct lengths-mer window of the reference proteomes of one PROTEOME_REFS category, standard residues only.

A proteome is a sequence reference, not a ligandome: these windows are what the source organism encodes, with no claim that any of them is presented. That is the whole point of keeping self separate from thymus – see KINDS.

lengths is required rather than defaulted to the full class range because the cost is linear in it and dominated by the host: the human proteome has ~11.4 M distinct 9-mers, so asking for all of 8-11 is several GB. For the human self category prefer mhcmatch.proteome.Proteome, which indexes on demand instead of materialising a set.

Parameters:

category (str)

Return type:

list

mhcmatch.mimics.proteome_window_array(category, L)[source]#

Sorted |S{L} array of every distinct standard-AA window of a PROTEOME_REFS category – the vectorized counterpart of proteome_peptides() for one length.

A consumer that projects or indexes these windows wants the array; materialising a 12 M-element Python set to hand it one is most of what proteome_peptides() costs.

Parameters:
  • category (str)

  • L (int)

mhcmatch.mimics.load_reference_sets(pmhc_dir=None, cls='mhc1', species='human', refs=None, proteomes=())[source]#

(self_set, foreign_sets) for scan(). self_set is the single tolerance reference (the self-kind entry, thymus by default); foreign_sets is {name: [peptides]} for the rest. refs overrides DEFAULT_REFS.

proteomes adds PROTEOME_REFS categories ("bacterial", "self") built from FASTA windows over the class’s plausible presented lengths. They land in foreign_sets whatever their KINDS entry says, because find_mimics() computes its E-value against exactly one background and that slot is already the thymic set – KINDS is how a caller reads a category, not how it is passed.

Parameters:
  • cls (str)

  • species (str)

Return type:

tuple

mhcmatch.mimics.neighbours(peptides, ref_sets, max_subs=2, threads=1)[source]#

{peptide: {category: [(n_subs, ref_peptide), ...]}} – same-length mimics, in batch.

The Hamming half of scan(), and 4,300x faster than the path :func:`scan` used to take. One seqtree.Index per (category, length), queried with search_batch in parallel C++ with the GIL released: measured at 237,000 queries/s against 55 for one find_mimics() call per peptide, returning identical counts and distances (benchmark repo, bench/results/neighbour_search_speed.md).

What it does not give is the per-allele presentation-aware E-value – that genuinely needs the k-mer/allele index find_mimics() builds, which is why this is a second entry point and not a replacement. Use it when the question is how close is the nearest reference peptide, and scan() with evalue=True when the significance of the match is the question.

Hits are nearest-first and exclude the query itself (distance 0); same length only, which is what Hamming distance means.

Reference peptides are deduplicated, and that is a fix. The compendia repeat a peptide once per allele/source-organism it was reported under – the viral IEDB set is 57,331 rows over 26,640 distinct peptides – so counting rows makes n_near a function of how often a sequence was deposited rather than of the sequence neighbourhood. On 400 Chowell peptides this changes exactly one result (AAAAATMAL had n_near = 3, all three the same peptide EAAAATCAL); top_mimic and top_subs are unaffected either way.

>>> neighbours(["GILGFVFTL"], {"viral": ["GILGFVFTA", "DDDDDDDDD"]})
{'GILGFVFTL': {'viral': [(1, 'GILGFVFTA')]}}
Parameters:
  • max_subs (int)

  • threads (int)

Return type:

dict

mhcmatch.mimics.scan(binders, self_set, foreign_sets, cls='mhc1', max_subs=2, near_subs=2, self_name='thymus', exclude_query=False, evalue=False, threads=1)[source]#

Mimic-scan an iterable of (peptide, allele) binders. Returns list[MimicResult] (one per binder × category with >=1 same-length reference peptide within near_subs substitutions).

self_set is the tolerance reference (category self_name); foreign_sets is {name: [peptides]}. max_subs is the fuzzy-search radius. find_mimics() excludes the exact query (a neoantigen’s identical peptide is its source, not a mimic), so n_exact is a direct set-membership check and n_near counts same-length reference peptides 1..``near_subs`` substitutions away (from the fuzzy hits, by exact Hamming distance). One find_mimics() call per binder scores every category at once.

``exclude_query=True`` is required whenever the output becomes a model feature. By default n_exact answers is this peptide in the reference set, which is the right question for a lookup (“is this a known viral ligand?”) and the wrong one for a feature: a reference assembled from the same deposit as the labels will contain the positives, so a known epitope scores n_exact = 1 for free and the model reads self-identity as foreignness.

Not hypothetical. 45% of the Gfeller cohort’s immunogenic peptides are exact matches to the viral IEDB ligand set, and a foreignness term built that way scored 0.714 AUROC there against 0.554 once self-matches were excluded (bench/results/gfeller_contamination.md). With exclude_query=True a peptide is never its own mimic and only genuine neighbours at 1..``near_subs`` substitutions contribute.

``evalue`` defaults to ``False``, the fast path, from 1.20.0 — it used to default to ``True``. On the fast path the search runs through neighbours() (one batched, threaded index query per category and length) instead of one find_mimics() call per binder, which is 4,300x the throughput on measured identical answers: every field except e_value and n_hits is identical, and tests/test_mimics.py pins that equality.

The default moved because the expensive path was the one you got without asking, while the docstring recommended the other – and find_mimics builds its index per call, so the cost is quadratic in the obvious usage. What you give up is real and not silent: e_value comes back nan and n_hits becomes the near-hit count, so a caller reading significance sees nan rather than a plausible wrong number. Pass evalue=True when the per-allele presentation-aware significance is the question; it now honours threads, which it previously accepted and dropped.

mhcmatch.mimics.patient_summary(results, binders)[source]#

Aggregate scan() output into patient-level counts for a dashboard row.

binders is the full strong-binder list (so “0 mimics” binders are counted too).

Return type:

dict

mhcmatch.mimics.write_table(results, path)[source]#

Write per-(binder, category) mimic results as a TSV (one row per category with a near mimic).

Parameters:

path (str)

Return type:

None