Skip to content

BAM cluster format + streaming-based demultiplexing

Latest

Choose a tag to compare

@bentyeh bentyeh released this 31 Oct 22:43
· 17 commits to main since this release

Use BAM files to represent clusters, instead of the previous text-based cluster file format.

  • All reads (whether from genomic DNA or bead oligo) use the following tags:
    • CB: barcode. '.'-delimited string of tags, ordered from terminal tag to the first ODD tag; sample name is appended to the end of the barcode. (example: NYStgBot_1-A1.OddBot_70-F10.EvenBot_46-D10.OddBot_33-C9.EvenBot_11-A11.OddBot_17-B5.sample1)
    • RT: read type. name of the DPM tag or BEAD tag. (example: BEAD_AB1-A1)
    • dc: PCR duplicate multiplicity
  • Genomic DNA (chromatin) reads are aligned to the genome. gDNA-specific tags include the following:
    • RG: read group. Target name assigned upon demultiplexing. (example: AB1-A1)
  • Bead oligo reads are unmapped (0x4 flag is set). Bead oligo-specific tags include the following:
    • YG: read group. (Within a cluster, should be the same as the RG value for genomic DNA reads.)
    • RX: bead oligo UMI sequence
    • QX: bead oligo UMI quality scores
    • YC: number of genomic DNA reads in the cluster

Streaming-based demultiplexing: All reads are first sorted by the cluster (bead) barcode. Then, only 1 cluster of reads needs to be fully loaded in memory at a time to assign a target label.

Pipeline counts is re-implemented to leverage Snakemake for parallelization.

Additional changes

  • Update conda environment files.
  • Simplify Snakemake workflow profiles.
  • Update documentation.

Credits

  • Albert Yang: initial proof-of-concept implementation of the BAM-based cluster format (commit e208534)

Full Changelog: v2.0.0...v3.0.0