Commands#
Five commands move reads, and three read no reads at all. That count is fixed: a sixth pipeline command needs a failing benchmark that the existing five cannot pass, because every flag and every stage is a thing that can disagree with another one.
The pipeline#
migec checkout reads.fq.gz --bc-pattern '^NNNNNNNN' -o co/ # find and cut out the barcode
migec refine co/S1.fq.gz -o ref/ # fix the errors IN the barcode
migec assemble ref/S1.fq.gz -o asm/ # collapse each molecule
command |
what it does |
the number it decides |
|---|---|---|
reads the barcode layout off the data |
the pattern you paste into |
|
finds the pattern, cuts the barcode out, demultiplexes |
how many reads carry a usable barcode |
|
corrects errors in the barcode itself, calls cells |
how many molecules there were |
|
collapses each molecule’s reads into one consensus |
the sequence, and the quality it is allowed to claim |
|
keeps every read of a fraction of the barcodes |
how many molecules a smaller copy of the library holds |
suggest and subsample are outside the pipeline in different directions: one runs before it
to tell you what to type, the other cuts a library down to something a laptop can iterate on.
Reading the output#
command |
what it does |
reads no reads because |
|---|---|---|
draws twenty QC panels from the TSVs the stages wrote |
it computes nothing; a figure that cannot be redrawn from a committed table will eventually disagree with the report |
|
|
writes or explains a barcode table, lists the presets |
it only ever touches the layout |
|
prints what a |
it reads headers |
Every stage writes plain TSV next to its output and a .json summary beside that, so the
figures and any downstream script see the same numbers the report printed.