Interface feature reference#
tcren features turns a set of TCR–pMHC structures into one flat table — one row per
structure, one interface descriptor per column — and tcren recognize turns that table into
scores. Features and scores are two commands because they are two jobs: the feature pass reads
structures and is the expensive half, the scoring pass is arithmetic over a table and can be
repeated for nothing.
# the four default families (--all adds potts and kinetics):
tcren features -s structures/ -i placement,interface,topology,energetics -o feats.tsv
# one family only -- and only that family is computed:
tcren features -s structures/ -i topology -o shape.tsv
# scores from the table, without re-reading a single structure:
tcren recognize --features feats.tsv -o scores.tsv
tcren recognize -s structures/ reads the structures itself and writes the descriptor table
(--full for the CDR3-frame layer, --mechanics for the kinetics terms). The scores Q,
T and S come from --features, which is the two-command route above.
Output is TSV. The first column complex.id is the structure-file stem (the SHA-256 TCR_hash
for the modelled sets), which is the join key to labels and AlphaFold confidences. Degenerate or
undefined terms are NaN. Every
feature is also available programmatically from tcren.recognition.recognition_features() (pass
full=True); the column-name tuples are tcren.recognition.RECOGNITION_FEATURES,
CDR3_FRAME_FEATURES and FULL_FEATURES.
tcren.recognition.DESCRIPTORS gives every column’s family and whether the receptor enters its
definition; tcren.recognition.descriptors() selects on both (see Families, and which descriptors involve the receptor).
Metadata that ships with a structure set#
Descriptors are computed from coordinates; the binding label, the epitope and allele, and the
generator’s own confidence (iptm, plddt, ranking_confidence) are not in the coordinates
and travel beside them, in a metadata.tsv keyed by id = the structure stem — the same value
tcren features writes into complex.id. tcren features joins it automatically when one is
present (--no-metadata to skip), and tcren.metadata.join_metadata() does the same in
Python:
from tcren.metadata import join_metadata
table = join_metadata(table, "vdjdb_free_pool") # no-op if the set ships none
The key must be the stem, not a bare receptor hash. A receptor that appears once as a positive
and once as a mispaired negative shares its hash across both rows, so a hash-keyed table is not
unique and the join silently drops rows: the shipped vdjdb_binder_benchmark table was keyed that
way and matched 523 of 1,089 structures, losing all 566 negatives.
tcren.metadata.read_metadata() raises on a duplicate id for that reason.
Conventions used below: tp = TCR↔peptide interface, tm = TCR↔MHC, pm = peptide↔MHC.
Phi_* is a raw interface energy Φ and dPhi_* its poly-alanine reference difference
ΔΦ = Φ(sequence) − Φ(reference) — the d is that difference, never a derivative, and the only
dd quantity in the package is ddG, the change in binding free energy on mutation. Every
energy column is named Phi, because the potential is fixed by the interface rather than chosen
per column: TCR↔peptide uses the TCRen potential, TCR↔MHC and peptide↔MHC the
Miyazawa–Jernigan (MJ) potential. Energies are in dimensionless statistical-potential units (more negative = more
favourable); they are log-odds ratios of contact frequencies and are not in kT.
Note
Both commands annotate the MHC chains, which needs the MHC allele reference. It is built from
IMGT on demand rather than bundled in the wheel, so run tcren build-mhc-ref once after a
pip install (see The MHC allele reference — build it once).
Options#
tcren features:
option |
default |
what it does |
|---|---|---|
|
required |
structure file, directory, |
|
|
the per-structure descriptor table |
|
|
comma-separated families; only what you ask for is computed |
|
off |
every family, |
|
|
Cα thresholds (Å) for the footprint flag complex, so the |
|
|
organism passed to the TCR annotator |
|
|
worker processes for featurisation ( |
|
on |
also search mouse, to catch a mis-declared organism; |
tcren recognize takes -s, -o (default recognize.tsv), --organism, -t and
--autodetect-species with the same meanings, plus:
option |
default |
what it does |
|---|---|---|
|
unset |
score a table already written by |
|
off |
add the 18 CDR3-frame descriptors and the intra-peptide terms |
|
off |
append the |
--features and -s are the two ways in. With --features the output is the score table
alone — complex.id, Q, T and S; with -s
it is the descriptor table.
A table read through --features is checked against the installed descriptor catalogue before it
is scored (tcren.provenance), and the command refuses a table written under a different
one rather than scoring it. Every table tcren features writes carries a
<name>.provenance.json sidecar recording the version, the invocation and a digest of the
catalogue, so a stale table cannot quietly produce a number that a fresh run would not reproduce.
Columns the reliability score reads#
tcren.reliability.s_score() is the recommended single-structure score and reads three blocks
off this table. Two of them come from tcren features directly, and the third has to be joined:
block |
source |
what to request |
|---|---|---|
|
|
|
|
|
|
|
|
join on |
The energy block reads n_contacts, the Potts count of available pairs
that engaged, so its full input is -i placement,interface,topology,potts. The column belongs to
the potts family and to no other: through 2.19.0 the footprint wrote its CDR-loop tally under
the same name — a different quantity on the same structure, 66 against 29 on 1ao7 — and since the
topology pass runs before the Potts one, the emitted column meant whichever family the caller
happened to ask for. That tally is n_loop_contacts now, and a table that still carries the old
name is refused rather than standardized
against the wrong population. Without n_contacts the contact term drops out and is reported as
n/a rather than imputed.
So -i placement,interface,topology,energetics is what tcren assess requires (see
Reliability: scoring one modelled structure), and
without the joined neg_energy column assess emits the two-block form and says so in its
report rather than imputing the missing block.
Families, and which descriptors involve the receptor#
Every emitted column is catalogued in tcren.recognition.DESCRIPTORS as
name -> (family, involves_tcr), and selected with tcren.recognition.descriptors() or with
tcren features -i. The families split by what each quantity is invariant under, which is
also the axis along which they carry independent evidence:
family |
what it is |
examples |
|---|---|---|
|
where the receptor sits in the groove frame. Frame-dependent. |
|
|
how much contact there is and of what chemical kind. SE(3)-invariant. |
|
|
the shape of the contact set, free of its size, plus the two faces read as height fields. SE(3)-invariant, so these need no canonical orientation. |
|
|
statistical-potential interface energies and their poly-alanine references. |
|
|
the same contact energy read against the partition function instead of a poly-alanine
interface, under the coupled contact-map model ( |
|
|
the interface as a network of breakable springs. Off unless asked for. |
|
Which reference to use is decided by the task, not by preference. energetics subtracts a
poly-alanine interface in the same pose, which is right when every candidate carries its own pose
— ranking peptides for a fixed receptor. potts subtracts every contact map the geometry admits,
which is right when the pose is shared and what varies is capacity — ranking receptors for a fixed
epitope. Each is at or near chance on the other’s task.
The gap between the two faces#
tcren.topology.surface.surface_map() rasterises a face as a height field on a grid in a
groove frame refit from the structure – x groove width, y peptide N->C, z toward the
TCR. Build it twice, once per side, and the two share a grid, so the space between them is a
subtraction rather than new geometry:
gap(x, y) = h_tcr(x, y) - h_pmhc(x, y) # Angstrom, per cell
The sign is the whole point, and it is not the one intuition suggests. Over 60 Native2026 crystals the median gap is -1.7 A and 71 % of cells are interdigitated: the receptor’s lowest point in a cell lies below the groove’s highest point in the same cell. The faces interlock rather than stack. A gap that grows is a receptor riding on a few high points.
Because the two signs mean opposite things, pooling them loses the distinction, so the field is reported three ways:
sc_gap_mean,sc_gap_sdthe pooled moments, in Angstrom.
sc_gap_vol,sc_interlockthe field integrated over the contact plane with the signs apart, in Angstrom^3 – the void and the interdigitated volume. On 1AO7 these are 315 and 1,089.
sc_interlock_frac,sc_gap_depth,sc_gap_height,sc_gap_asymthe field read by sign: what share of cells interlock (0.771 on 1AO7), how far in over those alone, how far off over the rest alone, and the balance of the two volumes in [-1, 1].
sc_gap_meanisinterlock_fracweighting the middle two, so these are the three numbers it collapses into one.
Measured over 4,907 labelled benchmark structures, sc_gap_mean and sc_gap_sd carry the
largest binder/non-binder contrasts of any published interface descriptor tested (Cohen’s d
-0.651 and -0.681) at an R^2 on all 141 pre-2.30 descriptors of 0.131 and 0.255 – a channel
nothing else in the catalogue reaches. sc_shape, which is Lawrence & Colman’s Sc on a raster
instead of a dot surface, is the more familiar quantity and the less novel: R^2 0.445, nearest
neighbour m_erank_tm at rho 0.414.
Two views of the same descriptors#
The six families are what computes the descriptors: one pass of one command each, and
tcren features -i <family> is how you ask for one. The invariance classes are what they
mean, and the two do not line up.
Every descriptor, by what computes it and by what it is invariant under. Edge labels are descriptor counts and edge width tracks them.#
Read the thick edges. placement is geometric throughout – 31 of 31 – and it is the only
family that describes the docking in the sense of a quantity preserved by distance-preserving
transformations. The topology family is the largest at 70 and mostly compositional: 34 of
its columns are diversity, coverage or labelling measures over cells, positions and chemistry, 22
are geometric (every surface height-field quantity is built from Angstroms on a metric grid) and
14 are invariants of the interface surface under continuous deformation. interface
contributes 23 counts, two continuous quantities and one categorical.
That matters when a block is built from a family rather than from a class. Q is named for
interface geometry and carries one continuous quantity of four – burial, an area – with no
angle, distance or height in it; T is named for shape and carries one topological invariant of
five. Seven of the nine terms across both blocks are compositional, which is why two blocks with
different names read the same evidence.
tcren.recognition.descriptors filters on either axis, and they compose:
descriptors("topology", invariance="topological") # the interface surface, 14 columns
descriptors(invariance="geometric", tcr_only=True) # the docking
See also
Every descriptor, with its units lists all 164 with their units and definitions.
The alanine scan, on both sides#
dPhi_tcr_pep and dPhi_pep_mhc are aggregate references: the whole peptide replaced by
poly-alanine at once. To see which residue earns the energy, scan one at a time.
tcren.ddg.alanine_scan() walks the peptide and tcren.ddg.tcr_alanine_scan() the
receptor’s contacted CDR residues. Both truncate one residue to alanine in 3D through
tcren.refine.substitute.substitute_residues(), recompute the contact map and rescore, so a
side chain that was the only thing bridging to its partner loses those contacts. Alanine is the
target this needs no rotamer for: its heavy atoms are exactly backbone plus Cβ, so truncating at
Cβ is the alanine.
$ tcren ddg -s complex.pdb --native LLFGYPVYV --alanine-scan --side both -o ala.csv
--side takes peptide (default), tcr or both; --virtual takes the fast path that
re-indexes the mutant on the native map with no atoms moved, and is peptide-side only, because
truncating a receptor side chain without moving atoms would leave every contact it made in place.
Output is long — one row per scanned residue, with side, pos, wt_aa and ddG, plus
chain.id, chain.type, region.type and pos_index on the receptor rows. ddG is
E(native) - E(Ala@residue) throughout, so a positive value marks a residue that was earning
its place.
tcren.ddg.tcr_alanine_reference() folds a receptor scan into four per-structure numbers —
dPhi_ala_cdr12, dPhi_ala_cdr3a, dPhi_ala_cdr3b and their total dPhi_ala_tcr. Each is
the sum of the per-residue ΔΔGs of that loop, not the energy of mutating the whole loop in one
pass; those differ once atoms move, because truncating every side chain at once loses contacts each
residue alone retains.
Warning
Before 2.25.0 the scan’s structural path threaded the whole peptide through
substitute_peptide(), which truncates every residue to
backbone plus Cβ. Each position was therefore read against a poly-stub baseline rather than the
native — on 1ao7 the native sequence threaded back through it keeps 14 of 29 TCR:peptide
contacts — and the resulting offset appeared at every position, including positions with no
contacts at all. substitute_peptide remains correct for the poly-alanine reference, where
every residue genuinely is mutated.
placement and interface were a single geometry family until 2026-08-24, and
energetics was physics. Both retired names still resolve in
descriptors(), so descriptors("geometry") returns the pooled
placement + interface set.
kinetics measures unbinding rather than nativeness, and it is the most expensive family, so it
is not computed unless asked for.
The catalogue holds descriptors only. The fitted composites that used to sit beside them — the
real-vs-shuffled recognizers, the frozen forced-pose logistic, the fitted binder score and the
cohort posterior — were removed in 2.26.0: their coefficients were frozen against training sets
that no longer exist, which made them the one part of the package a reader could not reproduce.
Scoring is tcren.reliability over this table.
involves_tcr is False for five columns — Phi_pep_mhc, dPhi_pep_mhc, mhc_class_bin,
Phi_pep_int and n_pep_int — each computed from the peptide and MHC alone. Two structures of the same epitope
on the same allele share their values whatever the receptor, so such a column carries cohort
identity rather than interface physics, and a model given one can reach a cohort-level label
without learning anything about recognition. Any analysis whose question is about receptors should
select tcr_only=True:
from tcren.recognition import descriptors
descriptors("energetics", tcr_only=True)
# ('Phi_tcr_pep', 'Phi_tcr_mhc', 'Phi_cdr12', 'Phi_cdr3a', 'Phi_cdr3b', 'dPhi_tcr_pep')
Core recognition descriptors (40)#
The interface block tcren.descriptors.compute computes itself
(tcren.recognition.RECOGNITION_FEATURES), emitted by tcren recognize with no extra flags.
These 40 are a subset of the catalogue, not the catalogue: tcren.recognition.DESCRIPTORS
carries 164 columns across the six families, the rest of them produced by
tcren.topology, tcren.energetics, tcren.potts and tcren.mechanics.
Coverage & burial#
Column |
Unit |
Description |
Source |
|---|---|---|---|
|
count |
Distinct TCR interface residues contacting the pMHC (interface size). |
contact map |
|
ratio [0, 0.5] |
|
contact map |
|
count |
Number of TCR↔peptide residue–residue contacts. |
contact map |
|
count |
Distinct peptide residues contacted by the TCR. |
contact map |
|
count |
Number of TCR↔MHC residue–residue contacts. |
contact map |
|
Ų |
Interface ΔSASA = SASA(TCR) + SASA(pMHC) − SASA(complex), biopython Shrake–Rupley. |
biopython SASA |
Docking geometry#
Column |
Unit |
Description |
Source |
|---|---|---|---|
|
degrees |
Incident (tilt) angle of the TCR over the pMHC groove — a clean structural angle. |
|
|
degrees |
TCR crossing (scanning) angle relative to the groove long axis. |
|
|
Å |
MHC-stub → TCR-stub rigid-body separation (native TCRdock geometry). |
|
|
radians |
Rigid-body dihedral of the TCR about the MHC stub (circular; wraps at ±π). |
|
|
unit |
y/z components of the TCR stub unit vector in the MHC frame. |
|
|
unit |
y/z components of the MHC stub unit vector. |
|
Interface energies#
Column |
Unit |
Description |
Source |
|---|---|---|---|
|
TCRen |
Raw TCR↔peptide interface energy (whole interface, all TCR regions). |
|
|
MJ |
Raw TCR↔MHC interface energy. |
|
|
MJ |
Raw peptide↔MHC interface energy. |
|
|
TCRen |
TCR↔peptide energy over the CDR1+CDR2 loops only. |
|
|
TCRen |
TCR↔peptide energy over the CDR3α loop only. |
|
|
TCRen |
TCR↔peptide energy over the CDR3β loop only. |
|
|
TCRen |
Poly-alanine reference delta of the TCR↔peptide energy (geometry-normalized ΔΦ). |
|
|
MJ |
Poly-alanine reference delta of the peptide↔MHC energy. |
|
Contact-type tallies#
Per-interface counts of contacts of each chemical type (tp = TCR↔peptide, tm = TCR↔MHC),
from tcren.contact_types.contact_type_counts().
Columns |
Unit |
Description |
|---|---|---|
|
count |
Salt-bridge contacts on each interface. |
|
count |
Hydrogen-bond contacts on the TCR↔MHC interface. (The TCR↔peptide count is emitted once,
as |
|
count |
Aromatic (π-stacking) contacts on each interface. |
|
count |
Hydrophobic contacts on each interface. |
|
count |
Remaining contacts on each interface. |
|
count |
Hydrogen-bond contacts on the TCR↔peptide interface; a term of Eq. Q. |
MHC class#
Column |
Unit |
Description |
|---|---|---|
|
0/1 |
1 if any MHC chain is MHC class II, else 0 (from MHC annotation). |
Interface quality — clashes & contact stability#
Coordinate-only reads of forced-pose quality, always emitted by recognize (not part of the 34
model features): a steric-clash burden (tcren.clashes) and TCR:peptide contact fragility
(tcren.mechanics.stability). Both are computed natively (_geom).
Column |
Unit |
Description |
Source |
|---|---|---|---|
|
count |
Peptide↔partner heavy-atom pairs overlapping by more than 0.4 Å (Bondi vdW radii); the signature of a non-physical / wrong-register pose. |
clashes |
|
Å |
Summed overlap of all clashing pairs — a total steric-burden measure. |
clashes |
|
count |
Expected TCR:peptide contacts lost under a 1 Å isotropic shift, |
stability |
|
Å |
Mean contact margin |
stability |
|
ratio [0, 1] |
Fraction of TCR:peptide contacts with margin ≥ 1 Å (robust to a 1 Å shift). |
stability |
CDR3-frame descriptors (18) — --full#
The FramePose layer the whole-TCR features miss: each CDR3 loop projected onto the pMHC groove
frame (basis u, w, n; origin = peptide Cα centroid). Computed by
tcren.recognition.recognition_features() with full=True
(CDR3_FRAME_FEATURES). The cdr3b_* strain terms are the load-bearing
signal for forced-pose / hallucination detection. Columns are prefixed by loop (cdr3a_ for TRA,
cdr3b_ for TRB):
Suffix |
Unit |
Description |
|---|---|---|
|
Å |
Distance from loop Cα centroid to the groove origin (how far the loop reaches out). |
|
unit |
Projection of the (centroid − origin) direction onto |
|
unit |
Projection of the loop N→C axis onto |
|
Å |
Minimum Cα–Cα distance from the loop to the peptide — engagement depth. |
|
Å |
Loop end-to-end extension |
The intra-peptide term (--full)#
The three interface energies above all sum over contacts between two different chains, so a
peptide held in its bound conformation by its own side chains costs the same as one that is not.
tcren.intra_peptide_energy() is that omitted term, and recognize --full emits it
(PEPTIDE_INTERNAL_FEATURES). It is computed over
tcren.peptide_internal_contacts() — heavy atoms within 5 Å, sequence separation ≥ 3 — under a
symmetrised potential, since an intra-chain pair has no from/to orientation to respect.
The 5 Å cutoff is the same contact definition the rest of the package uses, so an internal contact
and an interface contact mean the same thing. The separation floor is what does the work: it
excludes pairs that touch because they are bonded. Over the 17 deposited complexes in
tests/assets/pdb the count is 18 contacts at |i−j| ≥ 3 and 134 at |i−j| ≥ 2, and that
sevenfold jump is the i/i+2 pairs of an extended chain — geometry rather than folding.
That also sets expectations for the term’s size: a canonical extended class-I 9-mer makes zero to two internal contacts, so this separates candidates only where the peptide is genuinely bulged or packed against itself. It is off everywhere by default.
Column |
Unit |
Description |
|---|---|---|
|
MJ |
The peptide’s contact energy with itself, symmetrised potential. Lower = more favourable, as everywhere in tcren. |
|
count |
How many such contacts the peptide makes. |
As a scoring term rather than a descriptor, it is opt-in at each layer, weighted by w and
added to the energy it accompanies (w = 0, the default, computes nothing and leaves every score
byte-identical):
$ tcren score -s c.pdb -c candidates.txt --intra-weight 0.5 # score = Φ_TP + w·E_intra
$ tcren scoring -s c.pdb --intra-weight 0.5 # reports Phi_pep_int, folds w·it into Phi_total
from tcren import ContactMap, intra_peptide_energy, score_peptides
from tcren.potential import mj, tcren
cm = ContactMap.from_structure(structure, peptide_internal=True) # required for the term
intra_peptide_energy(cm, mj()) # the native peptide's own energy
intra_peptide_energy(cm, mj(), peptide="KQWLVWLFL") # a candidate threaded onto its pose
score_peptides(cm, candidates, tcren(), intra_weight=0.5, intra_potential=mj())
The term’s potential defaults to MJ, not TCRen: TCRen is derived from TCR↔peptide contacts and says nothing about the contacts a chain makes with itself.
Scores#
tcren recognize --features turns a feature table into the scores the method proposes. Every one
of them is fit-free and defined for a single structure.
Column |
What it discriminates |
Model |
|---|---|---|
|
Binder vs non-binder, and a real interface vs a manufactured one. The recommended score (Reliability: scoring one modelled structure). |
|
|
Interface quality for a single structure, against the shipped crystal reference. |
Fit-free |
|
Footprint shape, free of footprint size — the channel that holds up when the generator had no template to copy. |
Fit-free |
|
Crystal-natural against generated-forced pose. Not a column: 2.26.0 removed the
|
Fit-free |
Note
Cohort-relative scores live in tcren.cohort and are computed by tcren, not
downstream: pass the whole cohort you are ranking. Evaluation is the other side of that line —
ROC/PR, bootstrap intervals, and anything that consumes a binding label stay outside the
library, because this package is built to score without one.