Output files and columns#
What a run writes, and what is in it. This is the reference to reach for when a column’s meaning matters.
Reading arda’s files#
All of arda’s tabular output is tab-separated, uncompressed, with a single header line and quoting turned off.
Important
Read it with quoting off too. A reader using the usual default turns a blank field into the
two-character value "", which silently makes every clonotype look chimeric and every empty
flag look set.
import polars as pl
df = pl.read_csv("SAMPLE.clones.tsv", separator="\t", quote_char=None)
import pandas as pd
df = pd.read_csv("SAMPLE.clones.tsv", sep="\t", quoting=3) # csv.QUOTE_NONE
Coordinates are 1-based and closed, in query space, as both AIRR and GFF3 are. A region’s
*_start and *_end both name bases that belong to the region, so
sequence[start-1:end] is that region’s sequence. Anything 0-based half-open must convert.
An empty field means not evaluated, never zero and never false. A locus no read spanned has no
minimum junction length rather than a minimum of zero; a read that never reached both the junction
and FR4 has an empty productive rather than F.
Files a run writes#
arda rnaseq and arda amplicon write these under --out-dir, named from
--out-prefix (or from the sample id when one is given):
File |
Contents |
|---|---|
|
One AIRR Rearrangement row per mapped read. Stage 1’s output. The largest file, and the one to keep if you want to re-derive anything. |
|
One row per clonotype. Stage 2’s output, with Stage 3’s assembled clones folded in. This is the table most analyses start from. |
|
Stage 3’s assembled long-CDR3 contigs, as AIRR rows. Written when |
|
The run report: what was read, what mapped, why the rest did not, wall time, peak RSS, and the versions of arda, MMseqs2 and the reference. |
|
Run QC in long format. The numbers that decide whether a sample is usable, without re-reading the FASTQ. |
arda cells writes a different set, described under arda cells output.
<prefix>.airr.tsv — one row per mapped read#
A spec-valid AIRR Rearrangement
file; it passes airr.schema validation. The same schema is what arda annotate writes.
Identity and locus#
Column |
Meaning |
|---|---|
|
The read or record name, verbatim from the input. |
|
The query sequence, re-oriented to the plus strand if it mapped reversed. |
|
|
|
|
|
The cell barcode, present when |
Gene calls#
Column |
Meaning |
|---|---|
|
IMGT allele names. A comma-joined tie list when the alignment cannot separate several germlines — standard AIRR practice, and an exact statement of what the read supports. |
|
The D germline, and the second D of a tandem D-D. VDJ loci only; empty on a VJ locus, which is not a miss. |
|
The Karlin–Altschul E-value the D call was accepted on. Lower is better. Re-thresholdable
with |
|
The constant-region gene, from a hit on a J+C scaffold or from a read’s constant-region mate. |
|
The IGH isotype class — |
|
Added by |
Regions#
For each region r in fwr1, cdr1, fwr2, cdr2, fwr3, cdr3, fwr4:
Column |
Meaning |
|---|---|
|
1-based closed coordinates in query space. |
|
The region’s nucleotide sequence. |
|
Its translation. |
A region absent from the read is empty rather than guessed: a V-only query yields fwr1
through fwr3 and nothing else, and a J-only query yields fwr4 alone. There is no coverage
filter.
The junction#
Column |
Meaning |
|---|---|
|
Cys104 through Phe/Trp118 inclusive of both anchors. This is AIRR’s |
|
The IMGT CDR3: the same stretch excluding both anchors, two residues shorter. |
|
The non-templated stretches: V→D, D→D2, D→J. |
|
The anchor bases the called germlines template into each end of the junction. |
|
How many bases were imputed from the germline to finish a truncated junction, under
|
Warning
junction and cdr3 are not interchangeable. VDJdb’s cdr3 column holds what arda
calls junction. Conflating them shifts every coordinate by one residue at each end.
A junction is emitted only when the read actually reaches its Cys104 anchor. A read that lands mid-V and has no anchor gets an empty junction rather than a shortened one.
Alignment and mutation#
Column |
Meaning |
|---|---|
|
The aligned query and germline strings. |
|
Per-segment CIGAR against the germline. |
|
Segment boundaries in query space. |
|
The same boundaries in germline space. Filter on the V pair when you need to know how much V a read actually covered. |
|
Fraction identity to the called V germline, scoped outside the junction by default.
|
|
Per-segment substitutions in germline coordinates, comma-joined. See Somatic hypermutation. |
|
The Phred behind each mutation entry, comma-joined as integers, one for one with the
mutation list. Requires |
|
Phred over exactly the bases of |
Note
The two quality columns use different encodings on purpose: junction_quality is a Phred+33
string, v_mutation_quality is comma-joined integers. Reading one as the other yields
plausible numbers off by 33.
Productivity#
Column |
Meaning |
|---|---|
|
|
|
|
|
The conjunction of the two. |
All three are empty when the read never reached both the junction and FR4 — unevaluable rather than non-productive. On real bulk RNA-seq that is around 72 % of mapped reads. They are flags: no stage drops a non-productive read. Worked examples in Productivity: productive, stop_codon, vj_in_frame.
Search diagnostics#
Columns prefixed mmseqs2_ carry the raw search result — score, E-value, identity, and query
and target spans. They are diagnostics rather than AIRR fields; ignore them unless you are
debugging a call.
<prefix>.clones.tsv — one row per clonotype#
Twenty-one columns, in this order, plus umi_count when the input supports it:
Column |
Meaning |
|---|---|
|
The clonotype’s junction, anchors included. |
|
The calls the reads agreed on. |
|
Reads encompassing the junction. arda’s abundance estimate. Invariant across every
|
|
Distinct fragments. Use this when you need molecules rather than abundance. |
|
Distinct UMIs. Present only under |
|
As in the AIRR table. |
|
The non-templated stretches. |
|
Segment boundaries within the junction. |
Note
A V/J boundary disagreement inside the junction is not an error. Exonuclease chew-back and N/P addition mean the V-end / N-D-N / J-start partition is often not identifiable from sequence alone, so the ground truth for it is unknown. What is checkable is the junction’s outer bounds, the gene calls, and whether a junction was invented without an anchor.
Grouping is controlled by --clonotype-key (full = locus, V, J, junction; or junction
alone) and --call-level (allele or gene).
<prefix>.arda.json — the run report#
Top level: arda_version, mmseqs_version, reference (path, size, mtime),
wall_seconds, and one object per stage.
{
"arda_version": "2.30.1",
"mmseqs_version": "18-8cc5c",
"wall_seconds": 3.394,
"map": {
"total_reads": 1320,
"mapped_reads": 453,
"mapped_fraction": 0.3431818181818182,
"per_locus": {"IGH": 104, "IGK": 107, "IGL": 104, "TRB": 88, "TRA": 45, "TRG": 5},
"read_length_mean": 100.0,
"paired": true,
"threads": 4,
"peak_rss_mb": 352.1,
"unmapped": {
"prefilter_rejected": 788,
"no_hit": 8,
"constant_only": 54,
"below_min_score": 17,
"accounted": 1320
}
},
"assemble": {"seeds": 58, "contigs": 25, "contigs_complete": 10, "reads_rescued": 6},
"correct": {"clonotypes_in": 49, "clonotypes_out": 49, "reads": 57, "collapsed": 0}
}
Two fields deserve naming.
unmapped accounts for every read that produced no row, so a low mapped fraction has a
cause: rejected by the k-mer prefilter, no search hit, a constant-region-only hit, or a hit below
--min-score. accounted equals total_reads.
fast_fraction appears under map when the two-pass search ran. It is the fraction of reads
that hit both a V and a J segment, and it predicts whether the amplicon speed configuration pays
better than the library’s name does.
<prefix>.stats.tsv — run QC#
Four columns — scope, key, metric, value — one value per cell, so a metric can be
grepped, join-ed across samples or plotted without reshaping.
|
|
What it holds |
|---|---|---|
|
|
The run report verbatim: total and mapped reads, FASTQ bytes, read length, paired, threads, wall time, peak RSS. |
|
— |
Library-wide totals, junction lengths and quality, SHM rate, V/J gene coverage. |
|
|
Per locus, reads and clonotypes: functional, non-functional, stop codons, truncated junctions, junction length min/max/mean, junction quality, SHM rate, chimeras. |
|
|
Reads and clonotypes per germline gene. |
|
|
A recurrent, high-quality V mutation, with its frequency and mean Phred. A shortlist, not a genotype call. |
|
|
The distributions, keyed |
$ awk -F'\t' '$1=="chain" && $2=="IGH"' SAMPLE.stats.tsv
chain IGH reads 104
chain IGH reads_truncated_junction 1
chain IGH junction_nt_mean 48.75
chain IGH junction_quality_mean 35.6718
chain IGH shm_rate 0.0413
chain IGH clonotypes_chimeric 2
The distribution scopes exist because a mean cannot show a shape, and shape is most of the diagnosis. A bimodal junction length is two primer sets in one tube; a read-length cliff is an adapter left on; a clone-size distribution with no singletons is a library amplified before it was sequenced. Each of those reads as an unremarkable mean.
Cohort roll-up and the dashboard are in Quality control.
arda cells output#
File |
Contents |
|---|---|
|
The assembled per-cell contigs. |
|
Their AIRR annotation, with |
|
One row per |
|
One row per cell, including the two columns the doublet scatter uses. |
|
The same QC scopes as the bulk modes, so one cohort table holds every mode. |
Column-level detail and the doublet criteria are in Single cell.
Reference exports#
arda export-ref writes the reference itself, in three kinds × four formats — scaffolds,
collapsed per-allele segments, or per-allele CDR3 anchors, as TSV, FASTA, GFF3 or AIRR. Because
coordinates are already 1-based closed, GFF3 passes through unchanged. See
Exporting the reference.