The chunk format#
A chunk is one publication, stored as chunks/PMID_<id>.txt with one record per row. A
record reports paired chains: the alpha and the beta of one clone are columns of the same
row. chains is derived from that, never the other way round.
Two rows in two different chunks are independent reports, never duplicates, even when every field matches. The motif stage is tuned against that replication (Motif denoising and tuning).
The declared column order#
The 33 columns below, in this order, are what
template.tsv
carries, and both are generated from the field registry so neither can describe a column set the
build does not read. Chunk columns keep their dotted names: the tidy tables ship
underscore_case, but a chunk header is the submission contract.
This is an order, not a gate. A submitted chunk may order its columns however it likes and may omit
the optional ones - what is checked is that every column it does name is one of these, plus the
three curation columns a chunk may carry and the template does not: submitter, comment and
method.pairing.
Column |
Title |
|---|---|
chunk.id |
Chunk id |
cdr3.alpha |
CDR3 alpha |
v.alpha |
V alpha |
j.alpha |
J alpha |
cdr3.beta |
CDR3 beta |
v.beta |
V beta |
d.beta |
D beta |
j.beta |
J beta |
species |
Species |
mhc.a |
MHC A |
mhc.b |
MHC B |
mhc.class |
MHC class |
antigen.epitope |
Epitope |
antigen.gene |
Epitope gene |
antigen.species |
Epitope species |
reference.id |
Reference |
method.identification |
Identification method |
method.frequency |
Frequency |
method.frequency.count |
Frequency count |
method.frequency.total |
Frequency total |
method.singlecell |
Single cell |
method.sequencing |
Sequencing |
method.verification |
Verification |
meta.study.id |
Study id |
meta.cell.subset |
Cell subset |
meta.subset.frequency |
Subset frequency |
meta.subject.cohort |
Subject cohort |
meta.subject.id |
Subject id |
meta.replica.id |
Replica id |
meta.clone.id |
Clone id |
meta.epitope.id |
Epitope id |
meta.tissue |
Tissue |
meta.donor.MHC |
Donor MHC |
meta.donor.MHC.method |
Donor MHC method |
meta.structure.id |
Structure id |
The rest of this page is how to fill each one, which is the part no generated table can carry.
Complex information columns (required)#
These columns describe the TCR:peptide:MHC complex and are mandatory in any submission.
column name |
description |
|---|---|
cdr3.alpha |
TCR alpha CDR3 amino acid sequence. Give the complete sequence, starting with C and ending with F/W, where possible. Trimmed sequences are fixed at the build stage when sufficient V/J germline parts are present |
v.alpha |
TCR alpha Variable (V) segment id, to the best resolution available ( |
j.alpha |
TCR alpha Joining (J) segment id |
cdr3.beta |
TCR beta CDR3 amino acid sequence |
v.beta |
TCR beta V segment id |
j.beta |
TCR beta J segment id |
species |
TCR parent species ( |
mhc.a |
First MHC chain allele, to the best resolution available, |
mhc.b |
Second MHC chain allele ( |
mhc.class |
|
antigen.epitope |
Amino acid sequence of the epitope |
antigen.gene |
Parent gene of the epitope sequence (e.g. |
antigen.species |
Parent species of the antigen, to the best clade resolution available (e.g. |
reference.id |
Pubmed id, doi, etc |
submitter |
Name of submitting person/organization |
Notes:
If a record represents a clonotype whose alpha or beta sequence is unknown, leave the missing CDR3/V/(D)/J fields blank.
V/(D)/J fields may be left blank, in which case the CDR3 fixing and verification procedure is skipped for that record.
Every record must have at least one of
cdr3.alphaandcdr3.betafilled.
Method information columns (optional)#
These columns are optional to fill, but should be present in the table header. They set the confidence ranking of an entry: a single confidence score is computed from factors such as the fraction of a given TCRab sequence among the tetramer+ clones sequenced and the verification experiments performed.
column name |
description |
|---|---|
method.identification |
|
method.frequency |
Frequency in the isolated epitope-reactive population, reported as |
method.frequency.count |
Optional, and preferred over encoding the count in the string above (#696). Reads, UMIs or cells supporting this clonotype, as an integer. A clonotype supported by 3 reads is better evidence than one supported by 1, and a frequency erases exactly that difference: at any realistic depth both are |
method.frequency.total |
Optional. The sample total |
method.singlecell |
|
method.sequencing |
Sequencing method: |
method.verification |
|
Notes:
If
method.identificationis left blank, the record is assigned the lowest confidence score possible.
For special cases such as CD8-null tetramers, which use HLA with mutated residues that abrogate CD8 binding, specify
cd8null-tetramerinmethod.identificationrather than using themhc.afield.
The build collapses the columns above into a JSON string held in a single method column, e.g.:
{
"identification":"tetramer-sort",
"frequency":"5/13",
"sequencing":"sanger",
"verification":"antigen-loaded-targets"
}
Meta-information columns (optional)#
column name |
description |
|---|---|
meta.study.id |
Internal study id |
meta.cell.subset |
T-cell subset, free style, e.g. |
meta.subset.frequency |
Frequency of a given TCR sequence in the specified cell subset, e.g. |
meta.subject.cohort |
Subject cohort, free style, e.g. |
meta.subject.id |
Subject id (e.g. |
meta.replica.id |
Replicate sample coming from the same donor, also used for different time points, etc (e.g. |
meta.clone.id |
T-cell clone id |
meta.epitope.id |
Epitope id (e.g. |
meta.tissue |
Tissue used to isolate T-cells: |
meta.donor.MHC |
Donor MHC list if available, blank otherwise. IMGT nomenclature (e.g. HLA-A*02:01) is preferable. Allele group names (e.g. |
meta.donor.MHC.method |
Donor MHC typing method if available, blank otherwise |
meta.structure.id |
PDB structure ID if one exists, blank otherwise. A record with associated structural data gets the highest confidence score, so the field is not free text: a figure or table reference here awards that score for evidence the reader cannot check. QC reports anything that is not a four-character PDB entry id under |
comment |
Plain text comment, maximum 140 characters |
Note:
These columns are optional, but the subject identifier, replica identifier and the other id fields above are used when scanning a submission for duplicates. Duplicate records, those with identical complex information columns, are not allowed, but they are not treated as duplicates when they have distinct id fields.
The build collapses the columns above into a JSON string held in a single meta column, e.g.:
{
"cell.subset":"CD8+",
"subject.cohort":"HSV-2+",
"subject.id":12,
"clone.id":46,
"tissue":"PBMC"
}
Condition association columns (for extended database, TBA)#
Condition metadata:
column name |
description |
|---|---|
condition.name |
natural language terms like |
condition.id |
|
condition.type |
|
condition.subtype |
natural language terms like |
Association metadata:
column name |
description |
|---|---|
condition.freq |
fraction of samples matching the entry |
condition.count |
number of samples matching the entry (can be blank) |
population.freq |
fraction of controls matching the entry, or Pgen computed by OLGA/IgOR |
population.count |
number of controls matching the entry (can be blank) |
association.pvalue |
Association P-value, e.g. enrichment P-value for Fisher’s exact test |
association.test |
|
Ambiguous antigens (for extended database, TBA)#
Columns for peptide pools, long peptides used in T-cell culture expansion, and non-peptide ligands.
column name |
description |
|---|---|
antigen.epitope.long |
encompassing protein sequence containing the epitope |
antigen.peptide.pool |
e.g. |
antigen.nonpeptide |
|
Non TRAB columns (for extended database, TBA)#
Columns for non alpha-beta T-cells, CAR-T, etc.
column name |
description |
|---|---|
v.delta |
ID of Variable segment in delta chain |
cdr3.delta |
CDR3 of delta chain |
… |
… |
v.heavy.shm |
CIGAR string of hypermutations in the heavy chain Variable segment |
… |
… |