Structure summary (oracle)#
tcren.summarize_structure() is the one-call entry point the paper notebooks use:
it takes a single TCR–peptide–MHC structure and returns a bundle of ready-to-tabulate
frames by composing the pipeline (S1+S2), the percentile rank (S3) and the alanine scan
(S4). Nothing is re-derived — the facade only orchestrates those functions, so its
scores frame is byte-identical to tcren.pipeline.run()’s scores.
API#
The full signature and argument reference live with the module autodoc:
tcren.oracle.summarize_structure() (re-exported as tcren.summarize_structure).
The five returned frames:
Key |
Source |
Contents |
|---|---|---|
|
S1+S2 ( |
One row of per-interface energies ( |
|
S3 ( |
One row: the native peptide’s energy and its |
|
S4 ( |
Per-position alanine scan ( |
|
S1+S2 ( |
The per-residue region-markup table. |
|
S1+S2 ( |
The annotated residue-contact table. |
Example#
The script below turns a PDB into the five summary CSVs (default: the bundled 1ao7
fixture):
"""Turn a TCR-pMHC PDB into the five tcren summary tables.
Runs the :func:`tcren.summarize_structure` facade on a structure and writes each of the
five returned frames to a CSV, demonstrating the one-call entry point the paper notebooks
use. The default structure is the bundled ``1ao7`` test fixture.
Usage::
python scripts/summarize_structure_example.py [STRUCTURE.pdb] [OUT_DIR]
"""
from __future__ import annotations
import sys
from pathlib import Path
from tcren import summarize_structure
_DEFAULT_PDB = Path(__file__).resolve().parents[1] / "tests" / "assets" / "pdb" / "1ao7.pdb"
def main() -> None:
pdb = Path(sys.argv[1]) if len(sys.argv) > 1 else _DEFAULT_PDB
out_dir = Path(sys.argv[2]) if len(sys.argv) > 2 else Path("summary")
out_dir.mkdir(parents=True, exist_ok=True)
# One call composes S1-S4: pipeline scores + markup + contacts, percentile rank, and
# the per-position alanine scan. superimpose=False skips canonical orientation here.
tables = summarize_structure(pdb, superimpose=False, background=1000, alanine=True)
for name, frame in tables.items():
path = out_dir / f"{name}.csv"
frame.write_csv(str(path))
print(f"{name:9s} {frame.shape[0]:>4d} x {frame.shape[1]:<2d} -> {path}")
if __name__ == "__main__":
main()
Run it with the activated tcren environment:
$ python scripts/summarize_structure_example.py complex.pdb summary/
scores 1 x 6 -> summary/scores.csv
rank 1 x 5 -> summary/rank.csv
ddg 9 x 3 -> summary/ddg.csv
markup 605 x 7 -> summary/markup.csv
contacts 512 x 18 -> summary/contacts.csv
Command line#
The rank and ddg frames are also available as standalone CLI subcommands.
tcren rank — percentile-rank a peptide’s TCRen energy against a random pMHC
background. With no -c/--candidates it ranks each structure’s own native peptide:
$ tcren rank -s complex.pdb -o rank.csv
$ tcren rank -s complex.pdb -c candidates.txt --background 5000 --seed 1 -o rank.csv
The output carries complex.id, peptide, score (native energy), rank_pct
(fraction of background scoring at least as well — lower energy is a better binder, so a
small rank_pct flags a strong binder) and n_background. --background-source
points at a FASTA/text file of epitopes to sample the background from instead of drawing
it uniformly at random.
tcren ddg — fast ΔΔG of peptide mutations (virtual-matrix path; no atoms move).
ddG = E(native) - E(mutant), and lower energy binds better, so a positive value is
stabilising – the mutant improves on the native:
$ tcren ddg -s complex.pdb --native LLFGYPVYV --alanine-scan -o ddg.csv
$ tcren ddg -s complex.pdb --native LLFGYPVYV --mutant LLFGYPVYA --mutant LLFAYPVYV -o ddg.csv
Pass exactly one of --alanine-scan (one row per position mutated to alanine) or one
or more --mutant (neoantigen mode). Both subcommands share the --interface,
--regions, -p/--potential and --cutoff options with tcren score.