Structure and visualization#
The optional structure-based head, ligand context, and motif logos.
mhcmatch.structure module#
Structure-based pMHC Miyazawa–Jernigan contact energy and WT/MT ΔΔG (optional tcren dep).
Threads a peptide onto a template pMHC crystal groove and sums the MJ contact potential (all shipped
by tcren). For a query allele with no own template we borrow the groove-closest template
(the same pseudosequence kernel the diffusion uses). For a single-mutation WT/MT pair the
ΔΔG = MJ(mut) − MJ(wt) on one backbone is a physics-based differential affinity estimator
(adaptive double threading, Jojic et al. 2006) – the structural complement to the sequence-based
mhcmatch.affinity.
On measured HLA-A*02:01 the MJ energy tracks log-IC50 at Spearman ≈ 0.55 (see
bench/affinity/bench_structure.py), at ~0.02 ms/peptide after a one-time template build.
Needs the [structure] extra:
pip install 'mhcmatch[structure]' # pulls tcren
Template structures are not vendored (they live in tcren’s Canonical2026 set, which the
tcren wheel deliberately does not ship). Resolution order for the templates: the structure_dir
argument, then $MHCMATCH_STRUCTURES, then tcren’s own data dir ($TCREN_DATA_DIR, or an
editable tcren checkout’s data/) under Canonical2026.
- class mhcmatch.structure.StructureScorer(structure_dir=None, templates=None, pseudoseq=None, cutoff=5.0)[source]#
Bases:
objectMJ contact-energy scorer over template pMHC structures.
pseudoseq(amhcmatch.Pseudoseq) enables borrowing the groove-closest template for alleles without their own; omit it to restrict to exact-allele templates.- template_for(allele, length)[source]#
(pdb, chains)of the best template foralleleat peptidelength: exact allele if present, else the groove-closest templated allele (needspseudoseq).Noneif none.
- mj_energies(peptides, allele)[source]#
{peptide: MJ energy}for equal-lengthpeptidesonallele’s template (lower = stronger binding). Empty dict if no matching-length template. One batch swap.
mhcmatch.ligand module#
Full ligand spans: extend a binding core to the peptide that is actually presented.
Given a 9-mer binding core located in its source protein, presented_span() returns the most
likely observed eluted-ligand span around it – the peptide a wet lab would synthesise, rather
than the bare core. Three tiers of evidence, weakest last:
observed– a reference ligand in the panel that contains the core and occurs in the protein. A real eluted span: the gold standard when it exists.modeled– the highest-scoring feasible span underSpanModel, a flank/context model fit to mass-spectrometry ligandome data.fixed– caller-specified flank sizes, clipped at the protein termini (fixed_span()).
This is not a cleavage predictor, and not an immunogenicity predictor. MHC-II peptides are
generated bind-first-trim-later: the groove protects the core while exopeptidases erode the flanks,
so there is no strong sequence-specific endoprotease step to simulate (Paul et al. 2018,
PMID 30127785 – a dedicated MHC-II cleavage motif reaches AUC 0.767 on ligands and has zero
predictive power on CD4 epitopes). What this models is P(observed ligand span | source protein),
a convolution of protease specificity, HLA-DM editing, binding, stability and mass-spectrometry
detection bias. Context/flank models are known to improve ligand prediction while degrading
CD4 T-cell epitope benchmarks (Reynisson et al. 2020, PMID 32406916). Use this to enumerate and
choose ligands to synthesise or model structurally – never to rank epitopes by immunogenicity.
For MHC-I the peptide is the ligand: there is nothing to extend, so there is no span function.
processing_score() instead scores an 8-11mer’s source-protein context, the shape MHCflurry-2.0
uses for antigen processing (PMID 32711842). Class I and class II are deliberately different
entry points – a 9-mer class-II core is always <=11 residues and would silently misroute through
any length-based class inference.
- mhcmatch.ligand.PAD = '-'#
Out-of-protein context position (the span abuts a protein terminus). Modelled, not dropped: a ligand ending exactly at the protein’s C-terminus is evidence about where spans end.
- mhcmatch.ligand.LIGAND_KEYS = ('ligN+1', 'ligN+2', 'ligN+3', 'ligC-3', 'ligC-2', 'ligC-1')#
The ligand’s own terminal residues – inside the detected peptide, so subject to MS bias.
- mhcmatch.ligand.FLANK_KEYS = ('flankN-3', 'flankN-2', 'flankN-1', 'flankC+1', 'flankC+2', 'flankC+3')#
The residues flanking the ligand in the source protein – never in the detected peptide.
- mhcmatch.ligand.CTX_KEYS = ('flankN-3', 'flankN-2', 'flankN-1', 'ligN+1', 'ligN+2', 'ligN+3', 'ligC-3', 'ligC-2', 'ligC-1', 'flankC+1', 'flankC+2', 'flankC+3')#
3 upstream + 3 ligand-N + 3 ligand-C + 3 downstream. This is the NetMHCIIpan
-contextwindow (PMID 30446001); half the signal sits inside the ligand.- Type:
All 12 context positions
- mhcmatch.ligand.STRUCTURE_FLANK = 2#
across the 93 pMHC-II crystals of the Canonical2026 set the resolved peptide has a median length of 13 with a median of 2 flanking residues on each side, and only 13% resolve <=11 residues. So the core ± 1 (11mer) that TCRmodel2 and the fine-tuned-AlphaFold pipelines feed their networks is an input convention, not a statement about what is ordered – it discards real density in most structures. Reproduce with
bench/pdb_flanks.py.- Type:
Flank size for structure prediction (core ± 2 = a 13mer). Measured, not guessed
- mhcmatch.ligand.ASSAY_FLANK = 6#
Flank size for a synthesised assay peptide (core ± 6 = a 21mer). What matters for a CD4 assay is not hitting the eluted boundaries exactly – the APC re-trims whatever you give it – but that the peptide contains the natural ligand. Measured on held-out eluted ligands, the fraction of cores whose full observed ligand is contained in the emitted peptide: 13mer 11%, 15mer 31%, 17mer 52%, 19mer 67%, 21mer 80%. Longer also tracks the MHC-II affinity optimum of ~18-20 aa (O’Brien et al. 2008, PMID 19036163). The conventional 15mer covers only 31%.
- class mhcmatch.ligand.Span(peptide, start, end, core, core_start, source, score=0.0, n_alternatives=0, clipped=(0, 0), support=0)[source]#
Bases:
objectA ligand span located in its source protein.
- Parameters:
peptide (str)
start (int)
end (int)
core (str)
core_start (int)
source (str)
score (float)
n_alternatives (int)
clipped (tuple)
support (int)
- peptide: str#
- start: int#
- end: int#
- core: str#
- core_start: int#
- source: str#
- score: float = 0.0#
- n_alternatives: int = 0#
- clipped: tuple = (0, 0)#
- support: int = 0#
- property flanks#
(n_left, n_right)residues flanking the core within this span.
- class mhcmatch.ligand.SpanModel(ctx, lens, padbg=0.02, background='markov', _m1=None)[source]#
Bases:
objectLigandome-fit flank/context model:
P(observed ligand span | source protein).ctxis allele-agnostic by construction: exopeptidase trimming is a property of the proteolytic machinery, not of the groove. That is measured, not assumed – per-allele context PWMs sit within JSD 0.003-0.010 of the pooled one for MHC-II – and pooling also unlocks the ~70% of class-II eluted-ligand records whose restriction is only a placeholder.lensis a ligand-length prior, not a core-relative flank-length prior: defining an N-/C- flank length requires a binding core, and the allele-agnostic register is tied across >=2 frames on ~66% of real ligands, so such a prior would encode a tie-breaking rule rather than biology. N/C asymmetry is instead carried by the context positions, which are fit independently per side.The span score is a plain log-likelihood,
log P(L) + context log-odds, with no free parameters. A tuned weight on the length prior was tried – it looked better on the training fold and did not transfer (held-out set-recall 0.155 vs 0.158 unweighted, within noise), so it was dropped rather than shipped.- Parameters:
ctx (dict)
lens (dict)
padbg (float)
background (str)
_m1 (dict)
- ctx: dict#
- lens: dict#
- padbg: float = 0.02#
- background: str = 'markov'#
- context_score(protein, start, end, flank_only=False)[source]#
Log-odds of the context around span
[start, end)vs the proteome null.- Parameters:
protein – the source protein sequence.
start – 0-based span start.
end – 0-based exclusive span end.
flank_only – score only the 6
FLANK_KEYS. The 6 ligand-internal positions carry the peptide’s own anchor signal, so including them partly measures binding; the flank-only score is the honest processing signal.
- Returns:
Summed log-odds. Positions outside the protein score against
padbg.
- best_span(protein, core_start, core_len=9, delta=1.0)[source]#
Highest-scoring feasible span containing the core, as
(start, end, score, n_alt).Only spans that are real substrings of
proteinare enumerated, so the result never runs off a terminus.n_altcounts other spans withindeltalog-odds of the best – nested sets mean several spans are often legitimately correct.The binding term is identical for every span sharing this core, so it cancels in the argmax and is omitted: ranking is driven purely by the length prior and the flank context. (Do not substitute
AnchorModel.score()here – it is a max over register frames and so grows with peptide length, which would just select the longest span.)
- mhcmatch.ligand.load_span_model(cls='mhc2', background='markov')[source]#
Load the vendored
SpanModelforcls("mhc1"|"mhc2").Fit from IEDB mass-spectrometry eluted ligands against UniProt reference proteomes; see
src/mhcmatch/data/PROVENANCE.mdandbench/train_spans.py.
- mhcmatch.ligand.observed_spans(core, protein, corpus, core_start=None)[source]#
Reference ligands that contain
coreand occur inprotein, best first.corpusis any iterable of peptide strings (e.g.store._panel['mhc2'].epitopes). Because the caller supplies the protein,ligand in proteinis the provenance check – no source accession is needed. Occurrences that do not bracket the core are rejected: a ligand may also appear elsewhere in the protein.This is a lookup, not a prediction. Never fold its hit rate into a prediction metric, and never report a training ligand as a novel result – check
Span.source.
- mhcmatch.ligand.fixed_span(core, protein, left, right, strict=False, core_start=None)[source]#
Extend
corebyleft/rightresidues, clipped at the protein termini.A requested flank that runs off the protein is reported, not silently shortened: the shortfall lands in
Span.clipped.- Parameters:
core – the binding core (must occur in
protein).protein – the source protein sequence.
left – residues requested upstream of the core.
right – residues requested downstream of the core.
strict – raise instead of clipping when the flank does not fit.
core_start – 0-based start of
core; defaults to its first occurrence.
- Raises:
ValueError – if
coreis not inprotein, orstrictand the flank does not fit.
- mhcmatch.ligand.presented_span(core, protein, model=None, corpus=None, mode='auto', flanks=(3, 3), core_start=None)[source]#
The most likely presented ligand span around an MHC-II binding
core.- Parameters:
core – the 9-mer binding core, located in
protein.protein – the source protein sequence.
model – a
SpanModel; defaults to the vendored MHC-II model.corpus – reference ligands for the
observedtier (e.g. panel epitopes). Optional.mode –
"auto"(observed -> modeled), or force one of"observed"|"modeled"|"fixed". Benchmarks must pass"modeled": leavingobservedon turns the metric into a coverage statistic.flanks –
(left, right)formode="fixed".core_start – 0-based start of
core; defaults to its first occurrence. Pass it explicitly when the core repeats in the protein.
- Returns:
A
Span, orNonewhenmode="observed"and no reference ligand contains the core – itself informative: the core has never been eluted.- Raises:
ValueError – if
coreis not a substring ofprotein, or is not 9 residues.
Warning
Do not pick a peptide to synthesise from the ``modeled`` span alone. Held-out, it places both boundaries within +-1 residue only 31% of the time and within +-2 only 47% – barely better than simply centring a 15mer on the core (28% / 50%), which actually wins at +-3 (79% vs 62%). Its real edge is the exact-span hit rate (0.158 vs 0.069), i.e. the question “what was actually eluted?” – not “what should I make?”. For the latter use
fixed_span()withASSAY_FLANK(a 21mer, which contains the true ligand 80% of the time) orSTRUCTURE_FLANK(a 13mer, the median resolved crystal). Seebench/results/spans_mhc2_human.md.
- mhcmatch.ligand.processing_score(peptide, protein, model=None, flank_only=False, start=None)[source]#
Source-protein context log-odds of an MHC-I
peptide– a score, never a span.For MHC-I the peptide is the ligand, so there is nothing to extend. This scores how ligand-like its context looks (the antigen-processing signal MHCflurry-2.0 models, PMID 32711842) and composes into ranking, not into an emitted peptide.
- Parameters:
peptide – the 8-11mer, located in
protein.protein – the source protein sequence.
model – a
SpanModel; defaults to the vendored MHC-I model.flank_only – score only the 6 flanking positions. The ligand’s own termini carry its anchor signal, so the full 12-position score partly measures binding rather than processing.
start – 0-based start of
peptide; defaults to its first occurrence.
- Raises:
ValueError – if
peptideis not inprotein, or is not 8-11 residues.
mhcmatch.logo module#
Per-allele motif logos (information content) + length distributions.
motif() returns the numeric logo (PWM, per-position bits, length histogram) – pure-Python,
always available. render() draws it with logomaker (optional [logo] extra). MHC-I
logos use peptides of a fixed length (default the modal length); MHC-II uses register-anchored
9-mer cores. See the theory appendix §6.
- mhcmatch.logo.motif(store, allele, cls, length=None)[source]#
Logo data for
allele’s presented peptides.Returns
{allele, cls, width, n, pwm, bits, length_hist}wherepwm[i]is a residue->freq dict (sums to 1),bits[i]the information content (log2(20) - entropy) in [0, log2(20)], andlength_hista length->count dict over all the allele’s peptides.