Pipeline integration#
arda is a plain CLI over named files, so it embeds in any workflow engine without glue code. This page covers the one-shot mode commands and the ready-made Nextflow module.
One-shot commands#
arda rnaseq (bulk) and arda amplicon (targeted RepSeq) each run map, assemble and
correct in a single call — the three steps a pipeline almost always wants together — with the
speed and denoising configuration that regime needs:
arda rnaseq --r1 R1.fq.gz --r2 R2.fq.gz --out-prefix SAMPLE --out-dir results/
arda amplicon --r1 R1.fq.gz --r2 R2.fq.gz --out-prefix SAMPLE --out-dir results/
It writes four files under --out-dir:
|
corrected clonotype table ( |
|
mapped reads, AIRR Rearrangement schema |
|
Stage-3 long-CDR3 reads rescued by contig assembly
(omitted with |
|
merged run report (map + assemble + correct: reads mapped, per-locus counts, isotype/constant fragments, timing, peak RSS) |
Single-end input drops --r2. The defaults match the individual commands (--min-score 75,
--kmer 12 for ~300 MB peak RSS, complete-junction clonotypes, D mapping on); use map,
assemble, correct and shm separately when you need to tune their individual knobs, and
--no-map-d to skip D in all three stages.
Warning
arda rnaseq run was removed in 2.16.0. It was the entry point for amplicon libraries as
well, with the regime spelled out as four loose flags that do not compose — and the one flag it
exposed for four releases, --two-pass, is a loss in both regimes (0.762× on bulk, 0.87× on
an IGH amplicon). The regime is now the command name. --exact turns every speedup off and
reproduces the pre-2.16.0 default output; note that arda rnaseq without it enables
--prefilter, which costs ~0.15 % of mapped reads.
Nextflow module#
A drop-in, nf-core-style local module ships in integrations/nextflow/arda/ (main.nf,
environment.yml, Dockerfile, nextflow.config, README.md). It wraps one
arda <mode> call per sample, emits a versions.yml, and publishes to
${params.outdir}/arda/.
Requirements#
arda (PyPI arda-mapper) plus the mmseqs2 binary, both declared in environment.yml.
-profile conda builds the environment automatically; for -profile docker/singularity,
build the image from the provided Dockerfile and push it to your registry.
Drop into an nf-core/rnaseq pipeline#
The module consumes the same per-sample FASTQ channel the aligners do, so the sample sheet is unchanged. Five edits, each mirroring how an existing tool is wired:
Copy
integrations/nextflow/arda/tomodules/local/arda/in your pipeline checkout.Include and call it in
workflows/rnaseq/main.nfon the trimmed, aligner-independent channelch_strand_inferred_filtered_fastq, and mixARDA.out.versionsintoch_versions.includeConfig "../../modules/local/arda/nextflow.config"fromworkflows/rnaseq/nextflow.config(setsext.argsand the publishDir).Register a
run_arda = falseparam innextflow.configandnextflow_schema.json(strict schema validation).For container profiles, pin
withName: 'ARDA' { container = '<registry>/arda-mapper:2.6.0' }in your deployment config.
Then run with --run_arda. The module’s README.md has copy-paste snippets and a standalone
one-process test harness.
Runtime & resources#
arda is CPU-bound — the MMseqs2 search dominates — so give it cores. Measured on bulk tumor
RNA-seq at 32 cores, arda rnaseq --assemble (map + assemble + correct):
reads |
cores |
wall time |
throughput |
peak RSS |
clonotypes |
|---|---|---|---|---|---|
104.9 M (2×150) |
32 |
46 min |
~38,300 reads/s |
2,707 MB |
28,444 |
153.9 M (2×100) |
32 |
51 min |
~50,300 reads/s |
511 MB |
1,519 |
138.5 M (2×100) |
32 |
45 min |
~51,900 reads/s |
314 MB |
30 |
Throughput scales roughly linearly with cores; a typical full-depth bulk RNA-seq sample takes ~45 min
on 32 cores. The Nextflow module is labelled process_high; raise it with
withName: 'ARDA' { cpus = 32 } for full-depth data. --threads follows task.cpus.
Peak memory tracks repertoire richness, not read depth. The third sample above has more reads
than the first yet uses 9× less RAM — it carries 30 clonotypes against 28,444. Mapping alone is flat
and small (~300–400 MB at any depth); it is Stage 3 (assembly + error correction) that holds the clone
set in memory. Budget ~4 GB for a lymphocyte- or B-cell-rich sample and ~1 GB for a cold one. Under
a hard memory cap, --no-assemble restores the flat mapping-only profile — at the cost of the long
CDR3s that no single read spans.
Organism follows the genome#
The shipped nextflow.config reads params.genome — the iGenomes assembly key nf-core sets from
--genome — to both gate ARDA and choose its reference: GRCh38 runs with --organism human,
GRCm39 with --organism mouse, and any other (or unset) genome skips ARDA while the rest of the
pipeline completes. arda ships full references only for human and mouse, so nothing else needs wiring.
Tuning#
Extra arda flags pass through the module’s ext.args (--reconstruct, --min-score 0,
--kmer 11 …); a static ext.args string overrides the genome-driven --organism default.
--threads is wired to task.cpus automatically.