Skip to content

v2.8.0: fiberhmm-call is the CLI; fiberhmm-run removed

Choose a tag to compare

@mtcicero26 mtcicero26 released this 16 Apr 21:46
· 131 commits to main since this release

⚠️ Reverse-strand MM/ML parsing bug — UPGRADE TO v2.9.0

Versions 2.6 – 2.8.2 had a parser bug that shifted m6A/m5C modification positions on reverse-aligned reads (~24 bp median), causing ns/nl/as/al/MA/AQ tags to be wrong on reverse reads. Forward reads unaffected. DAF with IUPAC (R/Y) encoding unaffected.

Upgrade: pip install --upgrade fiberhmm (v2.9.0+). See v2.9.0 release notes for details.


Primary CLI consolidation

`fiberhmm-call` (introduced in v2.7.0, region-parallel added in v2.7.1) is now the recommended tool for all production workflows. It runs nucleosome/MSP + TF calling fused in a single Python process — 2–9× faster than the old `fiberhmm-run` subprocess chain.

Two-mode pattern

Input shape Mode Flag
Coordinate-sorted + indexed BAM Region-parallel (fastest, near-linear scaling) `--region-parallel`
Unaligned / unsorted / stdin Streaming (supports `-i -`, `-o -`, pipes to `ft fire`) default

Breaking: `fiberhmm-run` removed

The old wrapper ran apply + recall + fire as separate subprocess stages connected by streaming pipes. Each pipe cost a full BAM record re-serialization and re-parse on both sides — the dominant throughput cost on the old pipeline.

The `fiberhmm-run` entry point remains as a migration stub: running it prints a clear error and migration example, then exits. Scripts using `fiberhmm-run` will fail loudly instead of silently using slower code.

Migration

```bash

Old

fiberhmm-run -i in.bam -o out.bam --enzyme hia5 --seq pacbio --fire -c 8

New — sorted + indexed input (recommended for production)

fiberhmm-call -i in.bam -o recalled.bam --enzyme hia5 --seq pacbio
-c 8 --io-threads 16 --region-parallel --skip-scaffolds
ft fire recalled.bam out.bam # only if you need FIRE scoring

New — unaligned input, pipe to fire

fiberhmm-call -i in.bam -o - --enzyme hia5 --seq pacbio
-c 8 --io-threads 16
| ft fire - out.bam
```

Correctness validation

Tag-level equivalence against `fiberhmm-apply | fiberhmm-recall-tfs` reference, tested on DddB, DddA, and Hia5 datasets:

Enzyme Reads Byte-identical Notes
DddB DAF (22913 reads) 22913 100% all tags identical
DddA DAF (4800 reads) 4796 99.92% 4 reads differ only in MA (phantom pass-through annotation)
Hia5 PacBio (1226 reads) 1190 99% 36 reads differ only in MA (pre-tagged input edge case)

The differences are all in the recall-tfs "pass-through" path: when MM/ML extraction fails in the old pipeline but the input read has a pre-existing ns tag, recall-tfs emits an MA annotation based on that tag. `fiberhmm-call` fused doesn't replicate this (arguably cleaner — if extraction fails, no synthetic annotation). Does not affect production workflows.

Minor change: default `--min-read-length` for `fiberhmm-call`

Changed from 0 → 1000 to match `fiberhmm-apply` behavior. Prevents unexpectedly large ns/as counts on short-read contamination.

Install

```bash
pip install --upgrade fiberhmm
```

🤖 Generated with Claude Code