The chunk format#

A chunk is one publication, stored as chunks/PMID_<id>.txt with one record per row. A record reports paired chains: the alpha and the beta of one clone are columns of the same row. chains is derived from that, never the other way round.

Two rows in two different chunks are independent reports, never duplicates, even when every field matches. The motif stage is tuned against that replication (Motif denoising and tuning).

The declared column order#

The 33 columns below, in this order, are what template.tsv carries, and both are generated from the field registry so neither can describe a column set the build does not read. Chunk columns keep their dotted names: the tidy tables ship underscore_case, but a chunk header is the submission contract.

This is an order, not a gate. A submitted chunk may order its columns however it likes and may omit the optional ones - what is checked is that every column it does name is one of these, plus the three curation columns a chunk may carry and the template does not: submitter, comment and method.pairing.

Column

Title

chunk.id

Chunk id

cdr3.alpha

CDR3 alpha

v.alpha

V alpha

j.alpha

J alpha

cdr3.beta

CDR3 beta

v.beta

V beta

d.beta

D beta

j.beta

J beta

species

Species

mhc.a

MHC A

mhc.b

MHC B

mhc.class

MHC class

antigen.epitope

Epitope

antigen.gene

Epitope gene

antigen.species

Epitope species

reference.id

Reference

method.identification

Identification method

method.frequency

Frequency

method.frequency.count

Frequency count

method.frequency.total

Frequency total

method.singlecell

Single cell

method.sequencing

Sequencing

method.verification

Verification

meta.study.id

Study id

meta.cell.subset

Cell subset

meta.subset.frequency

Subset frequency

meta.subject.cohort

Subject cohort

meta.subject.id

Subject id

meta.replica.id

Replica id

meta.clone.id

Clone id

meta.epitope.id

Epitope id

meta.tissue

Tissue

meta.donor.MHC

Donor MHC

meta.donor.MHC.method

Donor MHC method

meta.structure.id

Structure id

The rest of this page is how to fill each one, which is the part no generated table can carry.

Complex information columns (required)#

These columns describe the TCR:peptide:MHC complex and are mandatory in any submission.

column name

description

cdr3.alpha

TCR alpha CDR3 amino acid sequence. Give the complete sequence, starting with C and ending with F/W, where possible. Trimmed sequences are fixed at the build stage when sufficient V/J germline parts are present

v.alpha

TCR alpha Variable (V) segment id, to the best resolution available (TRAVX*XX, e.g. TRAV7, TRAV7*01, TRAV7*02…). Strictly IMGT nomenclature. May be left blank if unknown.

j.alpha

TCR alpha Joining (J) segment id

cdr3.beta

TCR beta CDR3 amino acid sequence

v.beta

TCR beta V segment id

j.beta

TCR beta J segment id

species

TCR parent species (HomoSapiens, MusMusculus,…)

mhc.a

First MHC chain allele, to the best resolution available, HLA-X*XX:XX, e.g. HLA-A*02:01

mhc.b

Second MHC chain allele (B2M for MHCI)

mhc.class

MHCI or MHCII

antigen.epitope

Amino acid sequence of the epitope

antigen.gene

Parent gene of the epitope sequence (e.g. pp24). A property of the peptide, so within one chunk and one antigen.epitope it should be constant. QC reports a dense run of prefix+integer under counter in antigen.gene: dragging a cell down a spreadsheet column increments it, and that turned one epitope’s Eef2 into Eef2..Eef188, holding 65 clonotypes apart because the column is part of the deduplication key. The same rule covers antigen.species, mhc.a and mhc.b, and deliberately not meta.clone.id or meta.subject.id, where a counter is the content

antigen.species

Parent species of the antigen, to the best clade resolution available (e.g. HIV-1, HIV-1*HXB2)

reference.id

Pubmed id, doi, etc

submitter

Name of submitting person/organization

Notes:

If a record represents a clonotype whose alpha or beta sequence is unknown, leave the missing CDR3/V/(D)/J fields blank.

V/(D)/J fields may be left blank, in which case the CDR3 fixing and verification procedure is skipped for that record.

Every record must have at least one of cdr3.alpha and cdr3.beta filled.

Method information columns (optional)#

These columns are optional to fill, but should be present in the table header. They set the confidence ranking of an entry: a single confidence score is computed from factors such as the fraction of a given TCRab sequence among the tetramer+ clones sequenced and the verification experiments performed.

column name

description

method.identification

tetramer-sort, dextramer-sort, pelimer-sort, pentamer-sort, etc. for sorting-based identification. For molecular assays use antigen-loaded-targets (T cell specificity analysed against cells incubated with antigenic peptide) or antigen-expressing-targets (T cell specificity analysed against cells transformed with an antigenic organism, protein or peptide, e.g. BCL transformed with EBV). For magnetic cell separation use beads. Add cultured-T-cells or limiting-dilution-cloning if T cells were cultured before sequencing, since method.frequency then has a different meaning. For UMI-tagged multimers use tetramer-umi, etc. Separate phrases with a comma.

method.frequency

Frequency in the isolated epitope-reactive population, reported as X/X where possible, e.g. 7/30 if a given V/D/J/CDR3 is encountered in 7 out of 30 tetramer+ clones. The population is defined by the pMHC the sort used, not by the antigen it came from - a tetramer is one epitope on one allele. The formats X%, X.X%, X.X and 1e-04 are also supported. Measured over the corpus: 43,231 records write a ratio, 17,700 a percentage, 2,601 a float.

method.frequency.count

Optional, and preferred over encoding the count in the string above (#696). Reads, UMIs or cells supporting this clonotype, as an integer. A clonotype supported by 3 reads is better evidence than one supported by 1, and a frequency erases exactly that difference: at any realistic depth both are ~1/total, and 1/33921 against 3/33921 is 2.9e-5 against 8.8e-5, which a two-significant-figure export makes the same number.

method.frequency.total

Optional. The sample total method.frequency.count is out of. Where this pair is left blank the build parses it out of method.frequency when that is an unambiguous x/X; where it is given, the submitted value wins and nothing is parsed. Where all three are present they must agree - vdjdb qc reports frequency disagrees with its count and total and does not repair it, because which of the three the paper supports is a curation question.

method.singlecell

yes if single cell sequencing was performed, blank otherwise

method.sequencing

Sequencing method: sanger, rna-seq or amplicon-seq

method.verification

tetramer-stain, dextramer-stain, pelimer-stain, pentamer-stain, etc. for methods that include TCR cloning and re-staining with multimers. For magnetic cell separation use beads. restimulation, co-culture, antigen-loaded-targets, antigen-expressing-targets for molecular assays that validate the specificity of cloned T-cell receptors. direct if the affinity of the TCR of a specific T cell to the pMHC is quantified directly. Several comma-separated verification methods may be given.

Notes:

If method.identification is left blank, the record is assigned the lowest confidence score possible.

For special cases such as CD8-null tetramers, which use HLA with mutated residues that abrogate CD8 binding, specify cd8null-tetramer in method.identification rather than using the mhc.a field.

The build collapses the columns above into a JSON string held in a single method column, e.g.:

{
   "identification":"tetramer-sort",
   "frequency":"5/13",
   "sequencing":"sanger",
   "verification":"antigen-loaded-targets"
}

Meta-information columns (optional)#

column name

description

meta.study.id

Internal study id

meta.cell.subset

T-cell subset, free style, e.g. CD8+, CD4+CD25+

meta.subset.frequency

Frequency of a given TCR sequence in the specified cell subset, e.g. 5% means the TCR sequence represents an expanded clone occupying 5% of CD8+ cells

meta.subject.cohort

Subject cohort, free style, e.g. healthy or HIV+. Where possible, specify to what extent a healthy donor is healthy, e.g. CMV-seronegative.

meta.subject.id

Subject id (e.g. donor1, donor2,…)

meta.replica.id

Replicate sample coming from the same donor, also used for different time points, etc (e.g. 5mo)

meta.clone.id

T-cell clone id

meta.epitope.id

Epitope id (e.g. FL10)

meta.tissue

Tissue used to isolate T-cells: PBMC, spleen, etc. or TCL (T-cell culture) if isolated from re-stimulated T-cells

meta.donor.MHC

Donor MHC list if available, blank otherwise. IMGT nomenclature (e.g. HLA-A*02:01) is preferable. Allele group names (e.g. A02, B18) are also accepted (do not use an asterisk in such cases). Use a comma to separate alleles.

meta.donor.MHC.method

Donor MHC typing method if available, blank otherwise

meta.structure.id

PDB structure ID if one exists, blank otherwise. A record with associated structural data gets the highest confidence score, so the field is not free text: a figure or table reference here awards that score for evidence the reader cannot check. QC reports anything that is not a four-character PDB entry id under structure id is not a PDB id.

comment

Plain text comment, maximum 140 characters

Note:

These columns are optional, but the subject identifier, replica identifier and the other id fields above are used when scanning a submission for duplicates. Duplicate records, those with identical complex information columns, are not allowed, but they are not treated as duplicates when they have distinct id fields.

The build collapses the columns above into a JSON string held in a single meta column, e.g.:

{
   "cell.subset":"CD8+",
   "subject.cohort":"HSV-2+",
   "subject.id":12,
   "clone.id":46,
   "tissue":"PBMC"
}

Condition association columns (for extended database, TBA)#

Condition metadata:

column name

description

condition.name

natural language terms like T1D, pollen allergy, BRCA or YF vaccination

condition.id

ICD-11:5A10 for T1D in ICD-11 or OMIM:114480 for breast cancer in OMIM

condition.type

infection, vaccination, cancer, allergy or autoimmune

condition.subtype

natural language terms like acute or poor prognosis or grade II

Association metadata:

column name

description

condition.freq

fraction of samples matching the entry

condition.count

number of samples matching the entry (can be blank)

population.freq

fraction of controls matching the entry, or Pgen computed by OLGA/IgOR

population.count

number of controls matching the entry (can be blank)

association.pvalue

Association P-value, e.g. enrichment P-value for Fisher’s exact test

association.test

Fisher, TCRNET, ALICE or another statistical method

Ambiguous antigens (for extended database, TBA)#

Columns for peptide pools, long peptides used in T-cell culture expansion, and non-peptide ligands.

column name

description

antigen.epitope.long

encompassing protein sequence containing the epitope

antigen.peptide.pool

e.g. MIRA COVID19 TBD

antigen.nonpeptide

α-GalCer or KRN7000 TBD

Non TRAB columns (for extended database, TBA)#

Columns for non alpha-beta T-cells, CAR-T, etc.

column name

description

v.delta

ID of Variable segment in delta chain

cdr3.delta

CDR3 of delta chain

…

…

v.heavy.shm

CIGAR string of hypermutations in the heavy chain Variable segment

…

…