arda — Antigen Receptor Domain Annotation#
Fast FR/CDR region annotation for TCR and BCR nucleotide and amino-acid
sequences. arda builds a pre-aligned IgBLAST reference database once, then
maps queries with MMseqs2 and transfers the region markup through the alignment
in a small C++ hot path — producing AIRR-formatted output that matches IgBLAST.
Contents
- Installation
- Usage
- D segments and tandem D-D
- Somatic hypermutation
- Error correction
- The abundance test
- Quality: the evidence abundance does not have
- Where the abundance model stops working: the ladder and the cliff
- The quality-directed rescue
- Modes:
--ec-mode - The clonotype key:
--clonotype-key - What the options actually do
- Calibration: the MIGEC spike-ins
- What a junction disagreement means
- Run QC, verbosity and logging
- Pipeline integration
- Running on a cluster (SLURM)
- Use cases and analysis guidelines
- Bulk RNA-seq: extract a repertoire from a transcriptome
- Targeted amplicon / RepSeq
- Monoclonal QC: is my cell line what it says it is?
- Negative controls
- Low-frequency variants: spike-ins, MRD, minor clones
- Somatic hypermutation
- Comparing arda against another tool
- The one test that needs no external truth
- Reference database build
- Exporting the reference
- API reference