Ranking neoantigens#
The fitted EPIC aggregate over a candidate list.
mhcmatch.rank module#
Neoantigen candidate ranking, from a mutation-spanning window FASTA or an already-scored table.
The default score is a fitted aggregate (--score aggregate). Which one is a lookup on
(cls, species, mode) — see mhcmatch.rank.AGGREGATE_ARTIFACTS and
mhcmatch.rank.models() — and a combination that was never fitted raises rather than
being scored with another fit’s coefficients. A second fit of one cell is reachable through
mhcmatch.rank.AGGREGATE_VARIANTS and the variant= argument, which from 1.20.0 is how
the nine-term class-II corpus fit is read; it is not the default, because at class II the corpus
block was measured and adds nothing. Which cells ship is not restated here — that
list is generated from the artifacts themselves at build time, so it cannot drift; see
The shipped models, or run mhcmatch models --all against an install. This paragraph carried a
hand-typed copy of it once and was four facts stale by the time it was replaced.
Two other scores exist: the noisy-AND gate, a product of sigmoids, as --score gate /
mhcmatch.rank.GATE; and --score features /
mhcmatch.rank.FEATURES_ONLY, which computes every fitted column and scores nothing — what
a refit needs before its own artifact exists.
Rank neoantigen candidates: presentation, agretopicity, physicochemistry, expression, databases.
Two entry points, matching the two shapes real pipeline output arrives in:
rank_fasta()– a mutation-spanning window FASTA plus the donor’s HLA types. Windows are tiled, presented k-mers called, the wild-type counterpart recovered, and every component computed here. This is the full path.rank_table()– a table already scored by another tool (the 57-column pipeline.scored.csv). Its presentation columns are kept for comparison but affinity, agretopicity and the recognition terms are recomputed, because a score is only interpretable next to the model that produced it.
The combination is a gate, not a sum. Presentation and recognition are close to orthogonal (measured), and an additive predictor has to average away the fact that a recognition term is worth almost nothing on a peptide that is not presented and a great deal on one that is. So the aggregate is a noisy-AND – a product of two sigmoids, one per axis:
P(immunogenic) = sigmoid(a * presentation + b) * sigmoid(c * recognition + d)
It is monotone in both axes, it collapses to presentation alone when the recognition sigmoid
saturates, and it returns a probability rather than an uncalibrated sum. Coefficients are fitted in
the benchmark repo and vendored here as GATE; they are not tuned at call time.
A database hit is a flag, never a score contribution. If a candidate is an exact match to a
known immunogenic neoantigen, that is far stronger evidence than any model output, and burying it in
a weighted sum would let a mediocre model score dilute it. rank_fasta() reports it in
known_epitope and sorts those candidates into a top tier of their own, with the model score still
shown so the two can be compared.
The reference sets are built in by default (mhcmatch.known): confirmed tumour neoantigens
from NCI/Gartner, the epitope-resolution screens and the aggregated cohorts; peptides those screens
tested and found negative; IEDB-immunogenic epitopes; the thymic self-immunopeptidome; and the
viral ligandome. A neoantigen_neg hit is as informative as a neoantigen one and is reported
the same way – it is the only label that says this exact peptide was tried and did not work.
- mhcmatch.rank.GATE = {'a': 0.4546, 'b': -3.2283, 'c': 1.509, 'd': -0.1743, 'pres_mu': -0.3965, 'pres_sd': 0.8481, 'recog_mu': 0.5116, 'recog_sd': 1.2631}#
Noisy-AND coefficients and the standardizer they were fitted with, from
bench/neoag/gate_fit.pyon the presentation-matched IEDB-ligandome corpus (44,904 rows, 811 positive; seebench/results/neoag_gate.md). The recognition axis ismhcmatch.complement.score().The four
mu/sdentries are not decoration.a-dare coefficients on z-scores – the fit standardizes both axes first – so feeding a raw-log10(%rank)and a raw log-odds through them is applying coefficients to the wrong scale. That is what the previous values here did (they carriedmu = 0, sd = 1placeholders because the fitting script never wrote the standardizer out), and it moved the ranking, not merely the calibration: a product of two sigmoids is not rank-preserving under a monotone rescaling of one axis. Measured cost of the bug, same rows, corrected vs previous: TESLA 0.597 vs 0.473, Neopep 0.802 vs 0.662, Gfeller 0.782 vs 0.702 AUROC – every cohort improved.Refitted whenever the recognition axis changes, because
recog_mu/recog_sddescribe that axis. The length-awareaablock moved them (0.4642/1.1712 -> 0.5116/1.2631); the holdouts are unchanged within noise (TESLA 0.592, Neopep 0.804, Gfeller 0.784), which is the expected result – the length arm improves the recognition axis on the corpora it is fitted and cross-validated on, and the gate’s holdout performance is dominated by presentation.
- class mhcmatch.rank.Ranked(peptide, allele, allele_scored='', gene='', source='', keep=0, keep_reason='', presentation=nan, binder=nan, agretopicity=nan, occupancy=nan, d_occupancy=nan, wt_absent=0.0, physchem=nan, expr_pct=0.5, expression=nan, expression_imputed=False, wt_peptide='', n_alleles_presenting=0, alleles_presenting='', imputed='', known_epitope='', variant_type='', core='', core_offset=-1, core_source='', score=nan, p_response=nan, rank=0, components=<factory>, row=<factory>)[source]#
Bases:
objectOne ranked candidate, with every component kept separate so a rank can be explained.
- Parameters:
peptide (str)
allele (str)
allele_scored (str)
gene (str)
source (str)
keep (int)
keep_reason (str)
presentation (float)
binder (float)
agretopicity (float)
occupancy (float)
d_occupancy (float)
wt_absent (float)
physchem (float)
expr_pct (float)
expression (float)
expression_imputed (bool)
wt_peptide (str)
n_alleles_presenting (int)
alleles_presenting (str)
imputed (str)
known_epitope (str)
variant_type (str)
core (str)
core_offset (int)
core_source (str)
score (float)
p_response (float)
rank (int)
components (dict)
row (dict)
- peptide: str#
- allele: str#
- allele_scored: str = ''#
The allele the row’s numbers are actually against. Equal to
alleleexcept where the input named several (split_alleles), in which casealleleis the cell as supplied and this is the best presenter of the set. Keeping both means a caller can still join on what it sent.
- gene: str = ''#
- source: str = ''#
- keep: int = 0#
1when a whitelist rule named this row. Such a row survives anyrank_thresholdand says so in its own column: “kept because whitelisted” and “kept because it scored well” are different facts, and a reader cannot separate them from survival alone.
- keep_reason: str = ''#
epitope(exact hit against a validated immunogenic neoantigen),epitope~1(one substitution from one),gene(whitelisted symbol), or"". A gene hit says the gene is of interest and nothing about this peptide; an epitope hit is direct evidence about this peptide. Collapsing both intokeepwould file one as the other.- Type:
Which rule kept it
- presentation: float = nan#
-log10(presentation %rank); larger = better presented.
- binder: float = nan#
-log10(calibrated combined binder %rank) – the aggregate’s
B. Distinct frompresentation, which is the presentation head alone.
- agretopicity: float = nan#
The differential agretopicity index,
log10(Kd_WT / Kd_MT)against the recovered wild type; larger = more differential. This is the same quantity and the same orientation asmhcmatch.predict.Prediction.dai.Warning
mhcmatch.predict.Prediction.agretopicityis a different quantity under the same name: the raw ratioKd_MT / Kd_WT, which is the pipeline convention and runs in the opposite direction (there,< 1means the mutant binds better). A figure sourced from one path and labelled like the other has its sign flipped. Preferdai, which names one quantity on both paths, and readagretopicityonly where an existing consumer requires the name.
- property dai: float#
agretopicityunder the name that means the same thing on both code paths.
- occupancy: float = nan#
Fraction of MHC this peptide occupies at equilibrium,
a/(1+a)witha = [P]/Kd(PEPTIDE_NM). Unlike a %rank this is an absolute quantity, and unlike agretopicity it needs no wild type – so it is defined for a frameshift or fusion product that has none.Two properties of the range are worth knowing before it is read as a physical occupancy. The Kd is a predicted competition IC50 used as a dissociation constant in a Langmuir expression, which is standard practice in this literature and is an approximation rather than an identity. And the low tail is one tied mass point, not biology:
mhcmatch.affinity.y_to_ic50()clamps the predicted Kd to[1, 50000]nM first, so at the shipped[P] = 10nM occupancy is confined to[1.9996e-4, 0.909091]and cannot reach either bound. Measured over 669,974 scored rows, 23.6 % sit at exactly Kd = 50,000 nM and therefore share occupancy 1.9996e-4 exactly. The term is a compressed high-affinity detector; ranking within its low tail is ranking within a tie.
- d_occupancy: float = nan#
occupancy(Kd_MT) - occupancy(Kd_WT)– agretopicity in Michaelis-Menten form, bounded in[-1, +1]and defined with or without a wild type (d_occupancy()). Emitted and measured, and not fitted – it has no axis of its own onceoccupancyis in the model, and entered alongside it the coefficient flips sign.
- wt_absent: float = 0.0#
1.0 when no wild type was recoverable, so
d_occupancyfell back to the mutant’s own occupancy andagretopicityis undefined. The same object as the corpus’snoncanonicalflag: a frameshift, fusion or other product with no germline counterpart.
- physchem: float = nan#
calibrated physicochemical log-probability of immunogenicity.
- expr_pct: float = 0.5#
expression’s percentile within the scored batch, in (0, 1);0.5where there is no value – seeexpr_percentile(). Emitted, not fitted. It was the fitted expression term once and this line went on saying so after the fit moved: the shipped artifact’sfeatureslist has carriedexpr_lvlandexpr_normsince v9, andexpr_pctappears in neither. Being batch-relative it is also not comparable across runs, which is the reason it is not fitted.
- expression: float = nan#
log1p(TPM), observed if the input carried one, else the tissue/tumour reference median.
- expression_imputed: bool = False#
- wt_peptide: str = ''#
- n_alleles_presenting: int = 0#
How many of the queried allotypes present this peptide at or below the breadth band, and which. A peptide presented by three of a donor’s six class-I allotypes is a different bet from one presented by one: in the block model of
mhcmatch.portfolioit spans three blocks by itself. Derived from the per-allelemhcmatch.predict.binder_score()the ranker already runs, so it costs nothing extra. Column only – not a term of any fitted model until a benchmark says it earns one.
- alleles_presenting: str = ''#
- imputed: str = ''#
Which of the model’s features had to take their training mean for this row, joined by “;”. Empty when every feature was observed. A candidate with no IC50 has no occupancy and a frameshift has no wild type; those are candidates with incomplete data, not a different model, so they are scored and the substitution is declared here instead of being silent.
- known_epitope: str = ''#
name of the reference set an exact match was found in, “” if none.
- variant_type: str = ''#
What kind of variant produced the candidate, from the FASTA header’s
typefield on thefastapath and thetypecolumn on thetablepath. Carried so a cassette can hold a quota of non-conventional epitopes – a frameshift or fusion product is foreign over a stretch rather than at one position, so it fails differently from a missense and is worth a budget of its own (mhcmatch.portfolio.compose()). Reported, never scored.
- core: str = ''#
The NetMHCpan
core/Ofpair and the register behind it – seemhcmatch.store.binding_core(). Filled on therank fastapath, where the class-II model register is already in hand;rank tablehas no register and reports the heuristic, which is whatcore_sourcesays. Emitted only under--core, and never scored.
- core_offset: int = -1#
- core_source: str = ''#
- score: float = nan#
- p_response: float = nan#
Calibrated
P(this candidate elicits a detectable response)at the pool prevalence the run was given –probability(). Only as good as that prevalence, which is a prior the caller owns and which the model cannot supply: the fit gave every screen its own intercept precisely so base rate stayed out of the slopes.
- rank: int = 0#
1-based dense rank by
scoredescending; ties share a rank. Distinct from the row’s position in the file, because known epitopes are floated to the top of the listing and that is a display choice, not a ranking.
- components: dict#
- row: dict#
The input row this candidate came from, verbatim, on the paths that read a table (
rank_pairs()). Empty elsewhere. Whatrank --passthroughemits ahead of its own columns, so a caller’s table comes back annotated and re-ordered rather than replaced – and so the join that would otherwise reconstruct it is not needed. It cannot be done by the caller safely: a cell naming several alleles is split and the best presenter stands for the row, so the output has neither the same length nor the same allele column as the input.
- mhcmatch.rank.rank_fasta(store, fasta_path, alleles, cls='mhc1', *, tissue=None, tumor=None, refs=None, rank_threshold=None, top=None, gate=None, score='aggregate', channels=None, prevalence=0.06016260162601626, keep=None, species='human', expr_floor=None, expr_prefilter=0.0, mode='neoantigen', **kw)[source]#
Rank every presented k-mer in a mutation-spanning window FASTA.
storeis amhcmatch.Store;allelesthe donor’s HLA types in pipeline form.tissue(GTEx) and/ortumor(TCGAcancer_type,SKCMfor melanoma) supply reference expression where the FASTA header carries none.refsmaps a reference-set name to a set of peptides for the exact-match flag; ``None`` uses the built-in sets frommhcmatch.known, and{}disables the lookup.The wild type comes from
mhcmatch.predict.predict_fasta(), which recovers the position-aligned WT k-mer from the window’s own wild-type sequence; where the header has none, the nearest self peptide is the WT by construction – it is the first hit of a self-similarity search, which is the same operation.channelssupplies the aggregate’s three corpus features (CHANNEL_COLUMNS) as a callablelist[peptide] -> {name: sequence}. Withscore="aggregate"it is required: the model scores on the features it declares or not at all.score="gate"does not use them.rank_thresholddrops nothing by default (None->mhcmatch.predict. RANK_DEFAULT_TIER); pass"wb"/"sb"for the published class-aware cut or a number for an arbitrary percentile.keepis amhcmatch.predict.Keep– separate gene-symbol and epitope whitelists – whose rows survive any cut and are flagged inkeep/keep_reason.n_alleles_presentingno longer follows this threshold: an allele counts as presenting at its class’s weak cut, which is the published convention and does not move when a caller changes their own filter.- Parameters:
fasta_path (str)
cls (str)
tissue (str | None)
tumor (str | None)
refs (dict | None)
top (int | None)
gate (dict | None)
score (str)
prevalence (float)
species (str)
expr_floor (float | None)
expr_prefilter (float)
mode (str)
- Return type:
list[Ranked]
- mhcmatch.rank.rank_table(path, *, channels=None, keep=None, tissue=None, tumor=None, refs=None, store=None, cls='mhc1', gate=None, score='aggregate', prevalence=0.06016260162601626, species='human', expr_floor=None, expr_prefilter=0.0, mode='neoantigen', alleles=None, best_by='presentation', allele_policy=None)[source]#
Rank a table already scored by another tool, recomputing what this package can compute.
Reads a wide candidate CSV, resolving each column by name from
PEPTIDE_COLUMNSand its siblings (soepitope/MT Epitope Seq,best_allele/HLA Allele,tpm/Gene Expression,gene_name/Gene Name), plusref_seq/seqand any built-inscore. Presentation is recomputed withstorewhen one is given, so the ranking is this package’s own rather than a re-sort of someone else’s; the incomingscoreis preserved incomponents['score_builtin']so the two can be compared.- Parameters:
path (str)
tissue (str | None)
tumor (str | None)
refs (dict | None)
cls (str)
gate (dict | None)
score (str)
prevalence (float)
species (str)
expr_floor (float | None)
expr_prefilter (float)
mode (str)
alleles (list[str] | None)
best_by (str)
allele_policy (dict | None)
- Return type:
list[Ranked]
- mhcmatch.rank.gate_probability(presentation, recognition, gate=None)[source]#
The noisy-AND aggregate. Inputs are the raw axes; standardization happens here.
presentationis-log10(%rank)andrecognitionismhcmatch.complement.score(), both larger-is-better. They are z-scored with the fit corpus’s constants fromGATEbefore the coefficients apply – pass raw values, not pre-standardized ones.- Parameters:
presentation (float)
recognition (float)
gate (dict | None)
- Return type:
float
- mhcmatch.rank.BASE_COLUMNS: tuple = ('rank', 'peptide', 'allele', 'allele_scored', 'gene', 'score', 'p_response', 'presentation', 'binder', 'occupancy', 'd_occupancy', 'wt_absent', 'agretopicity', 'physchem', 'expression', 'expr_pct', 'expr_imputed', 'n_alleles_presenting', 'alleles_presenting', 'imputed', 'wt_peptide', 'known_epitope', 'variant_type', 'keep', 'keep_reason')#
The mhcmatch rank output schema, one source of truth. It lives here rather than inline in the CLI because a consumer – a pipeline module’s stub, a downstream join – has to be able to name the columns without running the command, and a schema typed out a second time is a schema that drifts. That is not hypothetical: the nextflow module’s stub carried an 18-column header against a 57-column table until 2026-09-20.
- mhcmatch.rank.MIMICRY_PAIRS: tuple = (('viral', 'anchor'), ('viral', 'tcr'), ('self', 'anchor'), ('self', 'tcr'), ('thymus', 'anchor'), ('thymus', 'tcr'))#
(reference, channel) in the order the mimicry columns are emitted.
- mhcmatch.rank.EXTENDED_COLUMNS: tuple = ('mimicry_logodds', 'autoimmune', 'viral_anchor', 'viral_tcr', 'self_anchor', 'self_tcr', 'thymus_anchor', 'thymus_tcr')#
the fitted aggregate and its six signed channels.
- Type:
Appended by
--extended
- mhcmatch.rank.ANNOTATE_COLUMNS: tuple = ('nearest_viral_anchor', 'source_viral_anchor', 'subs_viral_anchor', 'nearest_viral_tcr', 'source_viral_tcr', 'subs_viral_tcr', 'nearest_self_anchor', 'source_self_anchor', 'subs_self_anchor', 'nearest_self_tcr', 'source_self_tcr', 'subs_self_tcr', 'nearest_thymus_anchor', 'source_thymus_anchor', 'subs_thymus_anchor', 'nearest_thymus_tcr', 'source_thymus_tcr', 'subs_thymus_tcr', 'neoag_distance', 'neoag_nearest', 'neoag_n_within')#
what each channel’s nearest reference peptide actually was, then the tested-neoantigen lookup. Prior evidence, reported and never fitted (
MODELS.md).- Type:
Appended by
--annotate
- mhcmatch.rank.columns(extended=False, annotate=False, score='aggregate', core=False, cls='mhc1', species='human', mode='neoantigen')[source]#
The exact mhcmatch rank header for a given flag combination.
scorematters: the fitted aggregate emits its own four recognition channels, because a row should carry the features that produced it and nothing else.score="gate"does not use them and does not emit them.``(cls, species, mode)`` matter too, and this is the one implementation. A run emits the columns it COMPUTED and no others, which is a property of the artifact being scored rather than of the flags:
mode="pathogen"drops the wild-type columns and the whole expression block, and any fit whosefeatureslist names noC_corpus_thymus– which is both class-II artifacts – drops the corpus channels rather than emitting three NaN columns under an unchanged header.cli._rank_columnscalls this rather than restating it: two implementations of one header is how the nextflow module stub reached 18 columns against 57, which is the failure this function was written to prevent.- Parameters:
extended (bool)
annotate (bool)
score (str)
core (bool)
cls (str)
species (str)
mode (str)
- Return type:
list
- mhcmatch.rank.aggregate(cls='mhc1', species='human', mode='neoantigen', variant='')[source]#
The fitted
EPICartifact for(cls, species): features, coefficients, standardizer.EPIC – Expression, Presentation, Immunogenic Complementarity – names the four blocks the nine columns are fitted in, not the order they enter in; the pipeline order is presentation, expression, physchem, corpus, and the two recognition blocks are the two halves of Complementarity.
Fitted by
bench/epic/fit.pyover the neoantigen screens as a ridge logistic regression with an unpenalised per-screen intercept, which also writes this file;bench/run_epic.shis the whole chain that leads to it. The corpus it was fitted on is in the artifact, underfit–rows,positives,screensandbic– rather than in this docstring, which quoted a superseded corpus for two refits running.EPIC is hierarchical and Complementarity is kept whole. The nine columns enter in four blocks, in pipeline order, each on top of the last – see
AGGREGATE_BLOCKS. A recognition coefficient is therefore what the term is worth after presentation and expression, not in competition with them.presentation–binder(the calibrated Fisher combination of the presentation%rankwith the Potts affinity%rank) andlog10a(occupancy’s log-odds, see_logit10()). A%rankis a within-allele quantity where occupancy is absolute, so the two are not one axis entered twice: measured, they share Spearman +0.7431, whilebinderand the bare presentation rankpresshare +0.8797.presandoccupancyare both emitted; neither is fitted.expression–expr_lvl, this candidate’s own abundance, andexpr_norm, the same gene in the tumour’s matched normal tissue, bothlog2(1 + TPM/c)on the floor the tumour type’s own transcriptome sets. Two free terms rather than an imposed ratio, and the fit says the ratio would have been the wrong constraint: a difference of logs requires equal and opposite coefficients, and every fit since v9 has returned these two both positive — so abundance and normal-tissue level carry separate information rather than a ratio’s worth between them. Their sizes move at every refit and are not restated here;mhcmatch rank --coefficientsprints what the installed artifact actually uses.expr_pct, the within-batch percentile, is still emitted and is not fitted.physchem–C_phys_buriedandC_phys_charge,mhcmatch.complement.burial()over the TCR face on the Rose burial propensity and on Atchley AF5 electrostatic charge. Imported scales, so zero fitted residue parameters; burial carries a cysteine loading of +0.108 against the retired 30-columncomplementblock’s +0.693, and charge’s is +0.0056, the lowest of 141 scales swept. The two are orthogonal by measurement, r = +0.008 per peptide, which is what a chemistry block needs and what a burial/hydropathy pair does not have (r = -0.837, not identified).corpus–C_corpus_thymus,C_corpus_self,C_corpus_viral,mhcmatch.mimicry.corpus_R()against three reference corpora, split by when a T cell meets them. The thymic immunopeptidome is a biased sample of self – mTECs express tissue-restricted antigens under Aire and Fezf2 precisely to purge the clones worth purging – so similarity to it reads as danger and its coefficient is positive, whileself(the periphery) reads as tolerance and is negative.
The corpus channels are the exact Luksza sum under a graded kernel. Evaluated as a k-mer table contraction rather than the radius-2 trie walk that captured a median 0.4999 of it, and computed for every row rather than read from a peptide-keyed cache whose query set never contained three of the nine screens. Since v4 the kernel is identity-normalised BLOSUM62,
K[u,x] = exp(kappa (S[u,x] - S[u,u])), at k = 3 over the sliced face, which beats Hamming on held-out mean and median under an identical kappa-refit protocol.C_corpus_missingwent with the cache – there is no gap left to flag, so the column would be identically zero.bench/results/corpus_exact.mdandepic_corpus_kernel.md.An aggregate score is cheap. The
selfchannel costs a 64 KB table rather than the host-proteome reference index a neighbour search would force (6 min 15 s, ~7.5 GB), because the contraction needs counts rather than a trie.--no-selfand--score aggregateare not in conflict.There is no intercept and that is deliberate. Each screen was given its own, unpenalised, precisely so prevalence and candidate generation stayed out of the slopes; no single intercept transfers, and a new cohort has its own base rate. What ships is a ranking. A probability needs a named corpus, and the nine screens behind this fit span four orders of magnitude of base rate – from 0.0060 % positive (Neopep, 19 of 318,197 candidates) to 59.7 % (ITSNdb, 89 of 149). A probability quoted without naming the corpus is quoting one of those prevalences by accident.
- Parameters:
cls (str)
species (str)
mode (str)
variant (str)
- Return type:
dict
- mhcmatch.rank.aggregate_features(cls='mhc1', species='human', mode='neoantigen')[source]#
The feature names
(cls, species, mode)’s artifact declares, in order.AGGREGATE_FEATURESis the human class-I tuple as a module constant, because it is public API and is pinned against the artifact by a test. This is the general form, and it is what a caller building a feature dict for a mouse run should read.- Parameters:
cls (str)
species (str)
mode (str)
- Return type:
tuple
- mhcmatch.rank.AGGREGATE_ARTIFACTS: dict = {('mhc1', 'human', 'neoantigen'): 'aggregate_mhc1.json', ('mhc1', 'human', 'pathogen'): 'aggregate_mhc1_pathogen.json', ('mhc1', 'mouse', 'neoantigen'): 'aggregate_mhc1_mouse.json', ('mhc1', 'mouse', 'pathogen'): 'aggregate_mhc1_mouse_pathogen.json', ('mhc2', 'human', 'neoantigen'): 'aggregate_mhc2_human.json', ('mhc2', 'human', 'pathogen'): 'aggregate_mhc2_human_pathogen.json', ('mhc2', 'mouse', 'neoantigen'): 'aggregate_mhc2_mouse.json', ('mhc2', 'mouse', 'pathogen'): 'aggregate_mhc2_mouse_pathogen.json'}#
Which artifact scores a ``(cls, species, mode)``. One library, several fits, resolved by lookup rather than by branch – the same shape as
mhcmatch.diffusion._VENDORED_MODELS.models()reads it out with each artifact’s own identity attached.
- mhcmatch.rank.AGGREGATE_MODES: tuple = ('neoantigen', 'pathogen')#
The two immunological modes, and why the key carries one. A tumour neoantigen and a pathogen epitope are answered by different mechanisms – autoimmunity is not inflammation – so they are two models, never one model with an extra covariate. They also do not admit the same terms:
pathogendrops the expression block, which is undefined for a peptide the host does not transcribe. What it does with the corpus block is not a property of the mode but of the deposit a given fit was trained on – seeTERMS_PATHOGEN_EXPECTED– so twopathogenartifacts may legitimately carry different corpus channels, and every consumer here reads the artifact’s ownfeatureslist rather than a table keyed on the mode.Every
(cls, species)cell not listed inAGGREGATE_ARTIFACTSrefuses by name rather than silently serving a neighbour’s coefficients.
- mhcmatch.rank.AGGREGATE_VARIANTS: dict = {('mhc2', 'human', 'neoantigen', 'corpus'): 'aggregate_mhc2_human_corpus.json'}#
A second fit of a cell that already has one, reachable only by asking for it by name. A variant is not a mode and not a cell:
(cls, species, mode)still resolves to exactly one default throughAGGREGATE_ARTIFACTS, and nothing here changes whatmhcmatch rankscores. Kept in its own table for that reason – a key added to the registry above would move a default, and _modeldoc.ORDER and four tests pin that registry precisely so a default cannot move by accident.("mhc2", "human", "neoantigen", "corpus")is the nine-term fit of the class-II neoantigen cell, carrying the corpus block the six-term default drops. It ships because the block became computable in 1.20.0 – the HLA Ligand Atlas gives 132,818 class-II peptides over 29 benign tissues, where the class-II thymic fraction alone gave 27,987 – and it is not the default because computing it answers the question in the negative. On the 1,112 CD4+ rows (656 immunogenic, 157 references) every one of the three coefficients spans zero:C_corpus_thymus+0.0170 (95 % CI -0.4944 to +0.3995, sign held in 53 % of resamples),C_corpus_self+0.5227 (-0.3044 to +0.9815),C_corpus_viral-0.3612 (-0.8419 to +0.2928). At class II the corpus block adds nothing, and the six-term default is the honest release. Use this one to reproduce that measurement, not to score with.
- mhcmatch.rank.TERMS_MOUSE_EXPECTED: tuple = ('binder', 'log10a', 'expr_lvl', 'expr_norm', 'C_phys_buried', 'C_phys_charge', 'C_corpus_thymus', 'C_corpus_self', 'C_corpus_viral')#
The registry’s two rules, stated once here because the values below cannot carry them.
Human class I keeps the bare legacy name. It is the artifact every recorded number in the manuscript was produced under, it is named in
PROVENANCE.md, in_build.TARGETSand in the digesttests/test_aggregate_terms.pypins; renaming it to_humanfor symmetry would move all of that and buy nothing.A missing key is a refusal, not a fallback. All four
(cls, species)neoantigen cells are fitted, and so are all four pathogen cells – eight of eight. An unregistered cell raises rather than being served a neighbour’s coefficients; scoring class-II candidates with class-I ones is the mistake the lookup exists to make impossible. The branch that raises is kept even with the table full, because a cell can leave the registry – a refit withdrawn, an artifact not vendored – and the honest refusal is what must survive that, not a table that happens to be complete today.The term set the mouse class-I artifact declares. It is
AGGREGATE_FEATURES– that fit was run on the human specification deliberately, so the two are comparable coefficient by coefficient – and it is named separately because a mouse refit is free to move it and the human one is not. Both class-II artifacts declareTERMS_MHC2_EXPECTEDinstead.Nine names and, from mouse model version 5, nine free parameters. Versions 2-4 spent seven: the last three were one fitted scalar on the human artifact’s corpus direction rather than three independent coefficients, and a projection with one free scalar is three coefficients constrained to be proportional. v5 retired that axis – C_corpus_thymus carries +0.2919 at sign stability 0.94 on its own. The file lists all nine so one library scores both artifacts with no branch.
- mhcmatch.rank.TERMS_MHC2_EXPECTED: tuple = ('binder', 'log10a', 'expr_lvl', 'expr_norm', 'C_phys_buried', 'C_phys_charge')#
The class-II NEOANTIGEN artifacts carry six terms, not nine, and that is a specification. A corpus channel is a density over a reference set of peptides – displayed self, encoded self, viral – and the two cells reach that tuple for different reasons, which is worth keeping apart:
human, since 1.20.0, because the block was computed and answers in the negative. The HLA Ligand Atlas supplies class-II references, the nine-term fit ships as a variant (
AGGREGATE_VARIANTS), and on its 1,112 CD4+ rows all three coefficients span zero. The six-term default is a measurement, not a gap.mouse, because the reference tables are still class-I geometry, so the block would be fitted against a mismatched register and is dropped outright.
Either way the design has no C_corpus_* column and aggregate_score reads only the names the artifact lists, so one library still scores all eight cells with no branch. The nine names stay registered in
AGGREGATE_FEATURESbecause a recorded result cites the model version that produced it, and a registry that drops a name cannot say what those numbers were.It is not a property of the class, and the human variant is the proof: the block is computable on class-II references and was computed. Only the artifacts that declare this tuple drop it. Do not read the class-II pathogen cells as a second proof – they carry no corpus block either, for the mouse reason above, and PROVENANCE records that the four-cell refit superseded every pathogen term set that did.
- mhcmatch.rank.TERMS_PATHOGEN_EXPECTED: dict = {'mhc1': ('binder', 'C_corpus_self'), 'mhc2': ('binder', 'C_phys_buried', 'C_phys_charge')}#
The shipped pathogen artifacts carry two terms on class I and three on class II, against the nine of the neoantigen design. Each omission has its own reason and only some are weak coefficients; the paragraphs below give them in the order they were established.
log10ais well-defined here and still dropped, and the two facts are unrelated. It islog10([P]/Kd)of the candidate itself (_logit10()), so it needs no wild type at all –Ranked.occupancyis “defined for a frameshift or fusion product that has none”. It goes because it does not earn its parameter: corr(log10a, binder) = +0.8123 on the 38,106-row fitted design, at coefficient -0.0519, p = 0.137 and sign stability 0.887. Dropping it cost 0.0006 AUROC and raised PPV in the top 100 from 0.1000 to 0.1300.The genuinely wild-type-dependent quantities –
agretopicity,d_occupancy,wt_absent– are degenerate rather than absent in this mode (always NaN, alwaysoccupancy, always 1.0) and none is fitted anywhere; seeWT_COLUMNS.expressionis undefined, not missing. A pathogen epitope comes from an organism the host does not transcribe, so there is no source-gene abundance to measure and no matched normal to compare it against;expr_lvlwould be a number for a quantity that does not exist. Nothing is imputed, because there is nothing to impute from.C_corpus_viralis dropped by this fit’s deposit, not by the mode, and not because the reference is impure. A corpus channel is defined by the nature of its compartment –viralis “foreign in origin, observed presented on MHC” – and every peptide in that deposit satisfies it, including the 12.3 % that also carry a positive T-cell assay. What decides it is what the channel would be fitted against. This fit’s non-self negatives come fromligandome/viral_foreign_iedb.tsv.gzand theviralk-mer table is counted from that same file: on the fitted foreign arm, 35,472 of 35,472 negatives and 2,634 of 2,634 positives – 100 % of both classes – are exact members of the table. A channel every row on both sides belongs to is reading class membership, not similarity, and its coefficient would transfer to nothing.A
pathogenfit whose negatives do not overlap the table keeps the channel and carries one term more. That is why nothing in this module selects channels by mode: the admissible set is read off the artifact’s ownfeatureslist, perstand_in()andaggregate_features().C_corpus_selfstays on class I, and it is the theoretically interesting one: it is a host compartment, so a foreign epitope resembling host self should be seen by a tolerised repertoire and be less immunogenic. There is no circularity – the negatives are foreign, the table is host.C_corpus_thymuswent with the allele-balanced refit: the two correlate at +0.683 to +0.792 on these corpora and took near-equal-and-opposite coefficients, so what was fitted wasthymus - selfrather than two channels, and only one of the pair survives alone.The complement of a T-cell flag is not a negative set, which is why no arm here filters on one. IEDB’s T-cell export is positives-only, so “no positive record” mostly means nobody measured it.
Keyed by class, because the two classes do not have the same admissible terms. Class I keeps one corpus channel and drops the physchem pair; class II keeps the physchem pair and drops the corpus block entirely. That is not symmetry for its own sake – it is what the fits measured on the balanced corpora (2026-09-20):
class I:
C_corpus_selfis the term carrying the non-presentation signal (mouse -0.1487, p = 1.6e-08, sign stability 1.000 on 10,398 peptides), and adding Rose burial and AF5 on top moves ROC-AUC by +0.0008 (human) and +0.0003 (mouse). They did not earn a parameter.class II: there is no comparable corpus channel to keep – the host tables are class-I geometry – so the physchem pair is what remains beside presentation.
A consumer must still read the artifact’s own
featureslist rather than this table (stand_in(),aggregate_features()); this is the specification a shipped fit is checked against, not the thing that selects channels at run time.
- mhcmatch.rank.FEATURES_ONLY: dict = {'features': ['expr_lvl', 'expr_norm', 'C_phys_buried', 'C_phys_charge', 'C_corpus_thymus', 'C_corpus_self', 'C_corpus_viral'], 'model': '', 'version': None}#
Every fitted column, and no coefficients. The artifact-shaped stand-in that
score="features"supplies in a real artifact’s place.Computing the design matrix and scoring it are two things, and earlier only the second could ask for the first –
_finishdrives every column offa["features"], so a species or class with no artifact yet could not be measured, which is exactly what fitting one requires. That is a bootstrap the library owed a refit: bench/epic/fit_mouse.py needs these nine columns for H-2 rows before aggregate_mhc1_mouse.json exists.It declares no
expressionblock, so there is no recorded pooled floor to fall back on and afeaturesrun must resolve a real one from--tissue/--tumoror be givenexpr_floor. That is the correct failure: a fit whose floor silently became a cohort constant is the defect expr_norm already had once.
- mhcmatch.rank.models()[source]#
One record per shipped aggregate: what it is, what version it is, and what it was fitted on.
A model version is not the library version, and the manuscript pins the model. The paper quotes numbers produced by one specific fit; the library keeps moving underneath it while mouse and class II are worked on. So each artifact carries its own identity –
model_id(mhc1.human.neoantigen), an integerversionthat moves on a specification change, andrelease, the dotted package version the fit was accepted under. Citingmhc1.human.neoantigen v12 (release 1.20.0)names a fit that no later library version can move, which is exactly what quoting a package version cannot do. Every shipped cell was accepted at 1.20.0 and the package has moved past it since, which is the ordinary state: thereleasefield is what keeps the citation attached to the fit rather than to the wheel that happened to carry it.Returns records sorted by
model_id, each withcls,species,mode,model_id,version,release,file,featuresand the fit’s own row/positive counts where the artifact records them.- Return type:
list
- mhcmatch.rank.aggregate_score(features, imputed_out=None, cls='mhc1', species='human', mode='neoantigen')[source]#
Rank-score candidates with the fitted aggregate.
featuresis{name: sequence}.Every column in
AGGREGATE_FEATURESis standardized with the mu and sigma it was fitted with, then weighted. A missing column, or a non-finite value inside one, becomes the training mean – the same convention the fit used, so a candidate that lacks a wild type or a gene expression value is scored on the terms it does have rather than dropped. Higher is better.imputed_out, when given, is filled with one list per candidate naming the features that fell back to the training mean –["expr_lvl", "expr_norm"]for a peptide whose gene symbol the expression reference does not carry. Pass[]; it is grown to length. A row that names a feature here was scored without it, which is the difference between “this gene is silent” and “we have never heard of this gene”, and only this list tells them apart.Two things the caller owns, because getting them wrong is silent:
Compute each feature the way the fit did.
binderis-log10of the calibrated combined %rank – the Fisher statistic over the presentation rank and the Potts affinity rank, not the presentation rank alone, which is the separate columnpres–occupancyisa/(1+a)fora = 10 nM / Kd, the expression pair isexpr_level()andexpr_norm_level()on the floormhcmatch.expression.context_floor()returns for the tumour type, theC_physpair ismhcmatch.complement.burial()on the two scales ofPHYS_COLUMNS, and the threeC_corpuschannels aremhcmatch.mimicry.corpus_R().The ``C_corpus`` channels are densities, not counts. Each is a per-window mean over the whole reference set divided by that set’s total window mass, so it lands in [0, 1] and the three corpora are on one scale despite spanning 140,482 to 122 M reference windows. The standardizer is specific to the reference deposits, the mask and the shape parameters it was fitted with (
tcr5,mhcmatch.mimicry.SHAPES,k = 3); a density computed against a different deposit is not on the same axis. An incomparable value is a wrong value; the fix is to compute it against the fitted deposit, not to leave it out.
>>> full = {f: [0.0, 0.0] for f in AGGREGATE_FEATURES} >>> full["pres"] = [2.0, 0.1] >>> bool(aggregate_score(full)[0] > aggregate_score(full)[1]) True
- Parameters:
imputed_out (list | None)
cls (str)
species (str)
mode (str)
- Return type:
np.ndarray
- mhcmatch.rank.probability(scores, prevalence=0.06016260162601626)[source]#
Calibrated
P(response)from aggregate log-odds, anchored on a pool prevalence.aggregate()is fitted without a shared intercept – each screen got its own, unpenalised, so prevalence and candidate generation stayed out of the slopes. That makes the score a well-behaved ranking and leaves it one additive constant short of a probability. This supplies that constant, by the only rule that needs no new data: pickbso that the mean fitted probability over the pool equals the prevalence the caller declares,\[\frac{1}{n}\sum_i \sigma(s_i + b) = \pi ,\]then report \(\sigma(s_i + b)\). The left side is strictly increasing in
bfrom 0 to 1, so the root exists, is unique, and bisection finds it – no SciPy, no fitting.This is a prior shift, not a recalibration. It preserves the ranking exactly (
bis additive and \(\sigma\) is monotone), so nothing about the ordering is being claimed. What it buys is portability: a raw-score cut-off is meaningless across cohorts whose base rates differ by four orders of magnitude, and “P >= 0.2 at an assumed 6 % pool prevalence” is a statement another cohort can be held to.prevalenceis yours to set, and the answer is only as good as it is. Halving it roughly halves every probability; it does not move a single rank.>>> p = probability([3.0, 0.0, -3.0], prevalence=0.25) >>> [round(v, 4) for v in p] [0.6579, 0.0874, 0.0047] >>> round(sum(p) / 3, 6) 0.25
- Parameters:
prevalence (float)
- Return type:
list
- mhcmatch.rank.POOL_PREVALENCE: float = 0.06016260162601626#
37 immunogenic of 615 tested candidates on TESLA (Wells et al., Cell 2020), the community benchmark whose whole design is “these are the candidates a pipeline nominated; which of them respond”. It is a prior, not a measurement of your cohort, and it is the single number the emitted probability is most sensitive to.
The nine screens behind the fit span 0.0060 % positive (Neopep, 19 of 318,197 candidates, an exhaustive scan) to 59.7 % (ITSNdb, 89 of 149, a curated positive-enriched set) – four orders of magnitude – so no default can be right for every pool. What
--prevalencefixes is that a threshold on the raw score is not portable between them and a probability is: two cohorts scored at the same prevalence are on the same axis. Anchors worth knowing: a ranked shortlist that has already been through a cassette selection responds at ~19 % per unit (41 of 216 assayed units, 13 patients; Sahin et al., Nature 2026;651:1088-1096), and an unfiltered exhaustive scan is three orders of magnitude below this default.- Type:
Default pool prevalence for
probability()
- mhcmatch.rank.AGGREGATE_FEATURES: tuple = ('binder', 'log10a', 'expr_lvl', 'expr_norm', 'C_phys_buried', 'C_phys_charge', 'C_corpus_thymus', 'C_corpus_self', 'C_corpus_viral')#
The features the shipped aggregate expects, in order. Read it rather than typing the list. The fitted term set moves whenever the model is refitted, and a hardcoded copy would go stale silently;
tests/test_aggregate_terms.pyasserts this tuple equals the artifact’s ownfeatures, so the two cannot drift.
- mhcmatch.rank.AGGREGATE_COLUMNS: tuple = ('C_phys_buried', 'C_phys_charge', 'C_corpus_thymus', 'C_corpus_self', 'C_corpus_viral')#
The aggregate’s recognition features, emitted whenever the aggregate is what scored. A model emits the features it used and refuses to run without them.
The
C_physcolumns are computed here (mhcmatch.complement.burial()– a matrix product against a published residue vector, free). The threeC_corpuschannels are the only features a caller has to supply, because they need a reference deposit; seeaggregate().
- mhcmatch.rank.EXPR_COLUMNS: tuple = ('expr_lvl', 'expr_norm', 'expr_floor', 'expr_floor_pooled')#
The fitted expression terms, emitted beside
expression/expr_pctso a row carries every feature that produced its score. They are not recoverable from what else is emitted –expressionislog1p(TPM)and says nothing about the floor it was divided by, andexpr_normis a reference lookup the row does not otherwise report.log10ais not here for the opposite reason: it is exactlylog10(occ/(1-occ))on the emittedoccupancy.expr_flooris that floor, in TPM, andexpr_floor_pooledis 1 when it is the artifact’s pooled fallback rather than the context’s own. Both are reported, never fitted: the two terms are only comparable across runs that divided by the same number, and the number alone does not identify itself – GTEx Liver’s floor is 0.1800 TPM against a pooled 0.180005.Separate from
AGGREGATE_COLUMNS, which is also the required-input check for the recognition channels a caller must supply. These are computed inside the ranker, so a row that has not reached that point simply has no value and gets an empty cell.
- mhcmatch.rank.FEATURES_ONLY_PATHOGEN: dict = {'mhc1': {'features': ['C_phys_buried', 'C_phys_charge', 'C_corpus_thymus', 'C_corpus_self'], 'model': '', 'version': None}, 'mhc2': {'features': ['C_phys_buried', 'C_phys_charge', 'C_corpus_thymus', 'C_corpus_self'], 'model': '', 'version': None}}#
the columns that mode can be fitted on, and no others. Not
FEATURES_ONLYminus a few names as a convenience – the three it drops are dropped for the reasons inTERMS_PATHOGEN_EXPECTED, and asking for them here is an error rather than a wider net. Somhcmatch rank pairs --epitope pathogen --score featurescomputes exactly what the pathogen fit fits, needs no--expr-floor, and builds thethymusandselfcorpus tables without touching theviralone. Keyed by class, because the two classes admit different host corpus channels.Class II carried none until 2026-09-20, and the reason given was wrong about the tables. This said “there is no class-II counterpart – the mimicry tables are a k-mer lookup over a class-I TCR face”, so listing them would name always-NaN columns. But
corpus_countsresolves a per-row class-II register on the reference side and has done since the class-II cells were added:corpus_tables.npzshipsmhc2|thymus|human|3,mhc2|self|human|3and their mouse twins, all built over class-II registers. The columns were computable all along and this list is what declined to compute them.thymusat human class II now reads the pooled HLA Ligand Atlas class-II corpus – 132,818 peptides over 29 benign tissues, seemimics.CLS_REFS– so the pair is displayed self against bulk displayed self, which is the contrast the channel is read on.- Type:
The same stand-in for a
pathogenfit
- mhcmatch.rank.stand_in(mode='neoantigen', cls='mhc1')[source]#
The
--score featurespseudo-artifact for(mode, cls): what to compute when nothing scores.--score featureshas no fitted artifact to read afeatureslist off, so it needs one supplied. This is the single place that answers it, read by bothcli._aggregate_channels(which corpus tables to build) and the bannerrankprints (which columns it filled) – two answers that must agree, and did not when each carried its own list.This is the ADMISSIBLE set, not the shipped one. A pathogen fit may be fitted on any subset of these; the four released ones use two or three each. Narrowing this to whatever the current artifacts happen to use would make the next arm unfittable, which is the bootstrap
FEATURES_ONLYexists to avoid – a column that cannot be measured cannot be fitted.clsis required forpathogenbecause class II has no corpus block to compute.- Parameters:
mode (str)
cls (str)
- Return type:
dict
- mhcmatch.rank.AGGREGATE_BLOCKS: tuple = (('presentation', ('binder', 'log10a')), ('expression', ('expr_lvl', 'expr_norm')), ('physchem', ('C_phys_buried', 'C_phys_charge')), ('corpus', ('C_corpus_thymus', 'C_corpus_self', 'C_corpus_viral')))#
The hierarchy the aggregate was fitted as, in pipeline order. Blocks are entered one on top of the last, so a recognition coefficient is what it is worth after presentation and expression rather than in competition with them. Reported by
bench/results/epic_recognition_terms.md; carried here because a consumer grouping the emitted columns should not have to re-derive the grouping.
- mhcmatch.rank.CHANNEL_COLUMNS: tuple = ('C_corpus_thymus', 'C_corpus_self', 'C_corpus_viral')#
The subset of
AGGREGATE_COLUMNSthatchannels()has to return. TheC_physpair is not in it: the library can always compute them, so making the caller pass them would be ceremony.
- mhcmatch.rank.PHYS_COLUMNS: dict = {'C_phys_buried': 'Rose', 'C_phys_charge': 'ATCHLEY:AF5'}#
C_physcolumn -> the residue scalemhcmatch.complement.burial()reads for it. One place, because the column name and the scale it means are two halves of the same fact. Each is averaged over the TCR face bycomplement.burial().C_phys_buriedis the Rose 1985 burial propensity – the mean fraction of a residue’s surface area occluded on folding – andC_phys_chargeis Atchley 2005’s electrostatic-charge factor. The two are orthogonal by measurement, r = +0.008 at the peptide level, which is the whole reason the block carries them rather than a hydropathy scale: hydropathy is burial measured a second way (r = -0.837 peptide-level) and the pair was not identified. Charge also has the lowest cysteine loading of the 141 scales swept, -0.0028, which matters because the Chowell-family corpora carry a 12.5x mass-spectrometry cysteine gradient that a selected basis can learn. bench/results/physchem_pc.md.
- mhcmatch.rank.WT_COLUMNS: tuple = ('d_occupancy', 'wt_absent', 'agretopicity', 'wt_peptide')#
The three columns a pathogen epitope cannot have. Each is a contrast against a germline counterpart, and a peptide the host does not encode has none – so in
pathogenmode they are not merely missing, they are constant:agretopicityis always NaN,wt_absentalways 1.0, andd_occupancyalways equalsoccupancyexactly, becaused_occupancy()falls back to the mutant’s own occupancy when there is no wild type. Emitting three constants invites a reader to interpret them; dropping them says what is true. None is fitted in any artifact, so nothing downstream loses a term.log10ais deliberately NOT here. It islog10([P]/Kd)of the candidate itself (_logit10()) and needs no wild type at all – it leaves the pathogen fit for collinearity withbinder(r = +0.812), which is a different reason and belongs in the artifact’sfeatureslist rather than here.
- mhcmatch.rank.expr_percentile(rows)[source]#
Ranked.expression’s percentile withinrows, in (0, 1);0.5where absent.Emitted beside the model’s own expression term, not read by it. The fitted term is
expr_level(), which keeps the spacing between candidates that a rank discards. This column is kept because it is useful in its own right – it says where a candidate stands in the list it was submitted with, which is the question a shortlist actually poses, and it is unit-free and needs no floor.Two properties worth knowing:
The unit stops mattering. TPM, FPKM and raw counts give the same column, because a percentile is invariant to any monotone rescaling of abundance.
Missing needs no imputation constant. A row with no expression value sits at
0.5, which is what “no information” means on a percentile scale.
The cost, stated: the column is batch-relative. One peptide’s percentile depends on what else was scored with it, so it is a statement about standing in the submitted list rather than about the peptide.
A batch with fewer than two finite values – including one that is entirely absent – is all
0.5: one point has no percentile.- Return type:
list
- mhcmatch.rank.expr_norm_level(rows, floor, tumor=None, tissue=None, path=None, species='human')[source]#
The second expression term: the same gene’s level in normal tissue, on the same floor.
For a candidate whose source gene sits at
r_NTPM in the tumour’s matched normal tissue:expr_norm = log2(1 + r_N / c)
Two free terms, not a ratio.
expr_lvlis what this candidate is transcribed at andexpr_normis what its gene runs at in healthy tissue, both on one floor, so the model can represent a tumour-versus-normal contrast as equal-and-opposite coefficients if that is what the data supports. It does not: fitted, both are positive, so the normal-tissue level carries signal of its own rather than acting as a denominator. Imposing the ratio would have asserted the result instead of measuring it.r_Nis the gene’s median over the tumour’s matched normal tissue(s), and where no tissue resolves it is the gene’s own pan-tissue median – never missing. A missing value here becomes a batch constant, a constant cannot reorder anything, and a model that still spends weight on it has taken that weight from the abundance term. That is not hypothetical: it is what put the term below its own baseline on the one deposit in the fitting corpus carrying neither a gene name nor a tumour type.A candidate whose gene is not in the reference at all gets
nan, which is a different fact from a gene that is silent, and the two must not be collapsed.- Parameters:
floor (float)
tumor (str | None)
tissue (str | None)
path (str | None)
species (str)
- Return type:
list
- mhcmatch.rank.rank_pairs(store, rows, cls='mhc1', *, keep=None, tissue=None, tumor=None, refs=None, gate=None, score='aggregate', channels=None, prevalence=0.06016260162601626, species='human', expr_floor=None, expr_prefilter=0.0, mode='neoantigen', alleles=None, best_by='presentation', allele_policy=None)[source]#
Rank
(peptide, wt_peptide, allele)triples – the shape a benchmark or a variant table has.rank_fasta()needs mutation-spanning windows andrank_table()needs another tool’s.scored.csv; neither is what you have when a caller has already given you the mutant k-mer, its germline counterpart and the restricting allele. That third shape is the one every neoantigen screen is distributed in, so scoring one used to mean reimplementing this function outside the package – and a reimplementation is a second model nobody benchmarked.rowsis an iterable of mappings withpeptideandallele;wt_peptide,geneandtpmare optional. A row with no wild type is not an error:wt_absentcarries it, agretopicity is undefined rather than zero, and occupancy falls back to the mutant’s own – which is the case a frameshift or a fusion product is in, and imputing a germline peptide for it would report a number for a quantity that does not exist.Rows are grouped by allele and each group scored in one
mhcmatch.predict.binder_ranks()call, so the per-allele calibrator background is paid once per allele rather than once per row. Output order is the input’s, not the grouping’s.The reported ``allele`` is the one scored against, which is not always the one supplied. A cell naming several alleles is split (
split_alleles()) and the best presenter stands for the row, so every number in that row is against the reported allele. A caller joining the output back to its input must key on the peptide, not on the allele.allelesreplaces the row’s own cell with one candidate set for every row – the input’s restriction is emitted verbatim inalleleand is not read for scoring, andallele_scoredcarries the groove that won. That is a different measurement, not a wider one: abinderminimised over 203 grooves sits systematically lower than one taken at the recorded groove, so a model fitted under one must not score under the other.allele_policyis what says which, and_finish()refuses the mismatch.best_bypicks the criterion – see_wins().- Parameters:
cls (str)
tissue (str | None)
tumor (str | None)
refs (dict | None)
gate (dict | None)
score (str)
prevalence (float)
species (str)
expr_floor (float | None)
expr_prefilter (float)
mode (str)
alleles (list[str] | None)
best_by (str)
allele_policy (dict | None)
- Return type:
list[Ranked]
- mhcmatch.rank.split_alleles(cell, cls='mhc1')[source]#
Alleles named by one restriction cell, in input order, without repeats.
A screen that did not resolve which of a donor’s alleles restricts a candidate writes the whole genotype into the cell (
'HLA-A*01:01,HLA-B*07:02,HLA-C*07:02'). That string is not an allele name: it resolves to nothing, the calibrator then builds a background that does not depend on any allele, and the row comes backNaN. Splitting it is the difference between 1,076 keys and the 79 alleles they actually name. Names the pseudosequence tables do not know are dropped, so the caller can tell “no allele resolved” from “scored badly”.- Parameters:
cls (str)
- Return type:
list[str]
- mhcmatch.rank.species_of(cell)[source]#
'human'/'mouse'/None, from a restriction cell alone.Read off the first name the cell carries – a genotype names one donor, so its parts do not disagree about the genus.
Nonemeans the MHC is not one this package models, which is a different statement from “the allele is unknown”:'HLA class I'is human and unresolvable,'DLA-88*501:01'is resolvable in principle and not human.- Return type:
str | None