CDR3 fixing#
At this stage the build checks each CDR3 sequence against the reported V and J segments:
For canonical CDR3 sequences, those starting with a conserved
Cand ending withF/W: checks whether the 5’ and 3’ germline parts match the corresponding V/J segment sequences.For truncated CDR3 sequences: adds the conserved
C/F/Wresidues. Further missing residues are added when a relatively large contiguous V/J germline match is present.When an excessive germline part is reported (e.g.
FGXGinstead of justFat the CDR3 3’ part), the excess residues are removed.Mismatches in the V/J germline regions are corrected when a reliable non-contiguous V/J match is found.
Repertoire sequencing (RepSeq) data processing software reports canonical clonotype sequences, while a high number of epitope-specific TCR sequences in the literature are reported inconsistently. Fixing brings both into the same form, so that RepSeq data can be annotated with database records.
When the V/J germline match is good and the CDR3 sequence contains errors, the database carries the fixed sequence in place of the original. The fixer’s report is stored in the cdr3fix.alpha and cdr3fix.beta columns, e.g.
{
"fixNeeded":true,
"good":false,
"cdr3":"CASSQDVGTGGVFALYF",
"cdr3_old":"CASSQDVGTGGVFALY",
"jFixType":"FixAdd",
"jId":"TRBJ1-6*01",
"jCanonical":true,
"jStart":14,
"vFixType":"FailedBadSegment",
"vId":null,
"vCanonical":true,
"vEnd":-1
}
and
{
"fixNeeded":true,
"good":true,
"cdr3":"CASSLSRGGNQPQYF",
"cdr3_old":"CASSLSRGGNQPQY",
"jFixType":"FixAdd",
"jId":"TRBJ1-5*01",
"jCanonical":true,
"jStart":9,
"vFixType":"NoFixNeeded",
"vId":"TRBV14*01",
"vCanonical":true,
"vEnd":4
}
Field descriptions:
field |
description |
|---|---|
|
|
|
|
|
Fixed CDR3 sequence |
|
Original CDR3 sequence |
|
Type of fix applied to CDR3 J germline part |
|
|
|
J segment identifier |
|
A 0-based index of first CDR3 amino acid that belongs to J segment |
|
Type of fix applied to CDR3 V germline part |
|
|
|
V segment identifier |
|
A 0-based index of the last CDR3 amino acid of V segment plus one |
Note:
The V and J fix types are
NoFixNeeded,FixAdd,FixReplace,FixTrim,TruncatedGermline,FailedReplace(too many mismatches),FailedBadSegment(bad segment entry) andFailedNoAlignment(no alignment at all).
TruncatedGermlineis the one that needs reading twice, because it reports a success and a caveat together. IMGT ships some allele records as partial sequences that stop inside the anchor region, so the germline is correct as far as it goes but does not reach the anchor. The boundary placed from one is therefore a lower bound: residues past it are unattributed rather than known non-templated. The repair itself ran normally and the record counts asgood. It is on 1,144 chains of the current corpus, 1,132 beta and 12 alpha, andTRBV11-2*02alone is 850 of them; before it existed all 1,144 reportedFailedBadSegmentwithvEndat -1, which is how a placeable boundary came to be thrown away (antigenomics/arda#135).