Structure and visualization#

The optional structure-based head, ligand context, and motif logos.

mhcmatch.structure module#

Structure-based pMHC Miyazawa–Jernigan contact energy and WT/MT ΔΔG (optional tcren dep).

Threads a peptide onto a template pMHC crystal groove and sums the MJ contact potential (all shipped by tcren). For a query allele with no own template we borrow the groove-closest template (the same pseudosequence kernel the diffusion uses). For a single-mutation WT/MT pair the ΔΔG = MJ(mut) − MJ(wt) on one backbone is a physics-based differential affinity estimator (adaptive double threading, Jojic et al. 2006) – the structural complement to the sequence-based mhcmatch.affinity.

On measured HLA-A*02:01 the MJ energy tracks log-IC50 at Spearman ≈ 0.55 (see bench/affinity/bench_structure.py), at ~0.02 ms/peptide after a one-time template build.

Needs the [structure] extra:

pip install 'mhcmatch[structure]'      # pulls tcren

Template structures are not vendored (they live in tcren’s Canonical2026 set, which the tcren wheel deliberately does not ship). Resolution order for the templates: the structure_dir argument, then $MHCMATCH_STRUCTURES, then tcren’s own data dir ($TCREN_DATA_DIR, or an editable tcren checkout’s data/) under Canonical2026.

class mhcmatch.structure.StructureScorer(structure_dir=None, templates=None, pseudoseq=None, cutoff=5.0)[source]#

Bases: object

MJ contact-energy scorer over template pMHC structures.

pseudoseq (a mhcmatch.Pseudoseq) enables borrowing the groove-closest template for alleles without their own; omit it to restrict to exact-allele templates.

template_for(allele, length)[source]#

(pdb, chains) of the best template for allele at peptide length: exact allele if present, else the groove-closest templated allele (needs pseudoseq). None if none.

mj_energies(peptides, allele)[source]#

{peptide: MJ energy} for equal-length peptides on allele’s template (lower = stronger binding). Empty dict if no matching-length template. One batch swap.

mj_energy(peptide, allele)[source]#

MJ contact energy of peptide on allele’s template (nan if untemplated).

ddg(wt, mut, allele)[source]#

MJ(mut) − MJ(wt) on the shared backbone (the WT/MT differential; <0 = mutant binds better). wt and mut must be the same length. nan if untemplated.

mhcmatch.ligand module#

Full ligand spans: extend a binding core to the peptide that is actually presented.

Given a 9-mer binding core located in its source protein, presented_span() returns the most likely observed eluted-ligand span around it – the peptide a wet lab would synthesise, rather than the bare core. Three tiers of evidence, weakest last:

  • observed – a reference ligand in the panel that contains the core and occurs in the protein. A real eluted span: the gold standard when it exists.

  • modeled – the highest-scoring feasible span under SpanModel, a flank/context model fit to mass-spectrometry ligandome data.

  • fixed – caller-specified flank sizes, clipped at the protein termini (fixed_span()).

This is not a cleavage predictor, and not an immunogenicity predictor. MHC-II peptides are generated bind-first-trim-later: the groove protects the core while exopeptidases erode the flanks, so there is no strong sequence-specific endoprotease step to simulate (Paul et al. 2018, PMID 30127785 – a dedicated MHC-II cleavage motif reaches AUC 0.767 on ligands and has zero predictive power on CD4 epitopes). What this models is P(observed ligand span | source protein), a convolution of protease specificity, HLA-DM editing, binding, stability and mass-spectrometry detection bias. Context/flank models are known to improve ligand prediction while degrading CD4 T-cell epitope benchmarks (Reynisson et al. 2020, PMID 32406916). Use this to enumerate and choose ligands to synthesise or model structurally – never to rank epitopes by immunogenicity.

For MHC-I the peptide is the ligand: there is nothing to extend, so there is no span function. processing_score() instead scores an 8-11mer’s source-protein context, the shape MHCflurry-2.0 uses for antigen processing (PMID 32711842). Class I and class II are deliberately different entry points – a 9-mer class-II core is always <=11 residues and would silently misroute through any length-based class inference.

mhcmatch.ligand.PAD = '-'#

Out-of-protein context position (the span abuts a protein terminus). Modelled, not dropped: a ligand ending exactly at the protein’s C-terminus is evidence about where spans end.

mhcmatch.ligand.LIGAND_KEYS = ('ligN+1', 'ligN+2', 'ligN+3', 'ligC-3', 'ligC-2', 'ligC-1')#

The ligand’s own terminal residues – inside the detected peptide, so subject to MS bias.

mhcmatch.ligand.FLANK_KEYS = ('flankN-3', 'flankN-2', 'flankN-1', 'flankC+1', 'flankC+2', 'flankC+3')#

The residues flanking the ligand in the source protein – never in the detected peptide.

mhcmatch.ligand.CTX_KEYS = ('flankN-3', 'flankN-2', 'flankN-1', 'ligN+1', 'ligN+2', 'ligN+3', 'ligC-3', 'ligC-2', 'ligC-1', 'flankC+1', 'flankC+2', 'flankC+3')#

3 upstream + 3 ligand-N + 3 ligand-C + 3 downstream. This is the NetMHCIIpan -context window (PMID 30446001); half the signal sits inside the ligand.

Type:

All 12 context positions

mhcmatch.ligand.STRUCTURE_FLANK = 2#

across the 93 pMHC-II crystals of the Canonical2026 set the resolved peptide has a median length of 13 with a median of 2 flanking residues on each side, and only 13% resolve <=11 residues. So the core ± 1 (11mer) that TCRmodel2 and the fine-tuned-AlphaFold pipelines feed their networks is an input convention, not a statement about what is ordered – it discards real density in most structures. Reproduce with bench/pdb_flanks.py.

Type:

Flank size for structure prediction (core ± 2 = a 13mer). Measured, not guessed

mhcmatch.ligand.ASSAY_FLANK = 6#

Flank size for a synthesised assay peptide (core ± 6 = a 21mer). What matters for a CD4 assay is not hitting the eluted boundaries exactly – the APC re-trims whatever you give it – but that the peptide contains the natural ligand. Measured on held-out eluted ligands, the fraction of cores whose full observed ligand is contained in the emitted peptide: 13mer 11%, 15mer 31%, 17mer 52%, 19mer 67%, 21mer 80%. Longer also tracks the MHC-II affinity optimum of ~18-20 aa (O’Brien et al. 2008, PMID 19036163). The conventional 15mer covers only 31%.

class mhcmatch.ligand.Span(peptide, start, end, core, core_start, source, score=0.0, n_alternatives=0, clipped=(0, 0), support=0)[source]#

Bases: object

A ligand span located in its source protein.

Parameters:
  • peptide (str)

  • start (int)

  • end (int)

  • core (str)

  • core_start (int)

  • source (str)

  • score (float)

  • n_alternatives (int)

  • clipped (tuple)

  • support (int)

peptide: str#
start: int#
end: int#
core: str#
core_start: int#
source: str#
score: float = 0.0#
n_alternatives: int = 0#
clipped: tuple = (0, 0)#
support: int = 0#
property flanks#

(n_left, n_right) residues flanking the core within this span.

class mhcmatch.ligand.SpanModel(ctx, lens, padbg=0.02, background='markov', _m1=None)[source]#

Bases: object

Ligandome-fit flank/context model: P(observed ligand span | source protein).

ctx is allele-agnostic by construction: exopeptidase trimming is a property of the proteolytic machinery, not of the groove. That is measured, not assumed – per-allele context PWMs sit within JSD 0.003-0.010 of the pooled one for MHC-II – and pooling also unlocks the ~70% of class-II eluted-ligand records whose restriction is only a placeholder.

lens is a ligand-length prior, not a core-relative flank-length prior: defining an N-/C- flank length requires a binding core, and the allele-agnostic register is tied across >=2 frames on ~66% of real ligands, so such a prior would encode a tie-breaking rule rather than biology. N/C asymmetry is instead carried by the context positions, which are fit independently per side.

The span score is a plain log-likelihood, log P(L) + context log-odds, with no free parameters. A tuned weight on the length prior was tried – it looked better on the training fold and did not transfer (held-out set-recall 0.155 vs 0.158 unweighted, within noise), so it was dropped rather than shipped.

Parameters:
  • ctx (dict)

  • lens (dict)

  • padbg (float)

  • background (str)

  • _m1 (dict)

ctx: dict#
lens: dict#
padbg: float = 0.02#
background: str = 'markov'#
context_score(protein, start, end, flank_only=False)[source]#

Log-odds of the context around span [start, end) vs the proteome null.

Parameters:
  • protein – the source protein sequence.

  • start – 0-based span start.

  • end – 0-based exclusive span end.

  • flank_only – score only the 6 FLANK_KEYS. The 6 ligand-internal positions carry the peptide’s own anchor signal, so including them partly measures binding; the flank-only score is the honest processing signal.

Returns:

Summed log-odds. Positions outside the protein score against padbg.

best_span(protein, core_start, core_len=9, delta=1.0)[source]#

Highest-scoring feasible span containing the core, as (start, end, score, n_alt).

Only spans that are real substrings of protein are enumerated, so the result never runs off a terminus. n_alt counts other spans within delta log-odds of the best – nested sets mean several spans are often legitimately correct.

The binding term is identical for every span sharing this core, so it cancels in the argmax and is omitted: ranking is driven purely by the length prior and the flank context. (Do not substitute AnchorModel.score() here – it is a max over register frames and so grows with peptide length, which would just select the longest span.)

mhcmatch.ligand.load_span_model(cls='mhc2', background='markov')[source]#

Load the vendored SpanModel for cls ("mhc1" | "mhc2").

Fit from IEDB mass-spectrometry eluted ligands against UniProt reference proteomes; see src/mhcmatch/data/PROVENANCE.md and bench/train_spans.py.

mhcmatch.ligand.observed_spans(core, protein, corpus, core_start=None)[source]#

Reference ligands that contain core and occur in protein, best first.

corpus is any iterable of peptide strings (e.g. store._panel['mhc2'].epitopes). Because the caller supplies the protein, ligand in protein is the provenance check – no source accession is needed. Occurrences that do not bracket the core are rejected: a ligand may also appear elsewhere in the protein.

This is a lookup, not a prediction. Never fold its hit rate into a prediction metric, and never report a training ligand as a novel result – check Span.source.

mhcmatch.ligand.fixed_span(core, protein, left, right, strict=False, core_start=None)[source]#

Extend core by left/right residues, clipped at the protein termini.

A requested flank that runs off the protein is reported, not silently shortened: the shortfall lands in Span.clipped.

Parameters:
  • core – the binding core (must occur in protein).

  • protein – the source protein sequence.

  • left – residues requested upstream of the core.

  • right – residues requested downstream of the core.

  • strict – raise instead of clipping when the flank does not fit.

  • core_start – 0-based start of core; defaults to its first occurrence.

Raises:

ValueError – if core is not in protein, or strict and the flank does not fit.

mhcmatch.ligand.presented_span(core, protein, model=None, corpus=None, mode='auto', flanks=(3, 3), core_start=None)[source]#

The most likely presented ligand span around an MHC-II binding core.

Parameters:
  • core – the 9-mer binding core, located in protein.

  • protein – the source protein sequence.

  • model – a SpanModel; defaults to the vendored MHC-II model.

  • corpus – reference ligands for the observed tier (e.g. panel epitopes). Optional.

  • mode – "auto" (observed -> modeled), or force one of "observed" | "modeled" | "fixed". Benchmarks must pass "modeled": leaving observed on turns the metric into a coverage statistic.

  • flanks – (left, right) for mode="fixed".

  • core_start – 0-based start of core; defaults to its first occurrence. Pass it explicitly when the core repeats in the protein.

Returns:

A Span, or None when mode="observed" and no reference ligand contains the core – itself informative: the core has never been eluted.

Raises:

ValueError – if core is not a substring of protein, or is not 9 residues.

Warning

Do not pick a peptide to synthesise from the ``modeled`` span alone. Held-out, it places both boundaries within +-1 residue only 31% of the time and within +-2 only 47% – barely better than simply centring a 15mer on the core (28% / 50%), which actually wins at +-3 (79% vs 62%). Its real edge is the exact-span hit rate (0.158 vs 0.069), i.e. the question “what was actually eluted?” – not “what should I make?”. For the latter use fixed_span() with ASSAY_FLANK (a 21mer, which contains the true ligand 80% of the time) or STRUCTURE_FLANK (a 13mer, the median resolved crystal). See bench/results/spans_mhc2_human.md.

mhcmatch.ligand.processing_score(peptide, protein, model=None, flank_only=False, start=None)[source]#

Source-protein context log-odds of an MHC-I peptide – a score, never a span.

For MHC-I the peptide is the ligand, so there is nothing to extend. This scores how ligand-like its context looks (the antigen-processing signal MHCflurry-2.0 models, PMID 32711842) and composes into ranking, not into an emitted peptide.

Parameters:
  • peptide – the 8-11mer, located in protein.

  • protein – the source protein sequence.

  • model – a SpanModel; defaults to the vendored MHC-I model.

  • flank_only – score only the 6 flanking positions. The ligand’s own termini carry its anchor signal, so the full 12-position score partly measures binding rather than processing.

  • start – 0-based start of peptide; defaults to its first occurrence.

Raises:

ValueError – if peptide is not in protein, or is not 8-11 residues.