Skip to content

Releases: IPNP-BIPN/bwa-mem4

bwa-mem4 v4.3.3

Choose a tag to compare

@github-actions github-actions released this 15 Aug 21:00

bwa-mem4 4.3.3

The crate becomes dependable: needletail's xz, BAM/CRAM output and mimalloc are
now features, all on by default, so the binary is unchanged while the library
target builds with no C toolchain, no lzma link conflict and no global allocator
imposed on an embedder.


Download bwa-mem4 4.3.3

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.3.3-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.3.3-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.3.3-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.3.3-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.3.3-linux-x86_64.tar.gz
./bwa-mem4-v4.3.3-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).

bwa-mem4 v4.3.2

Choose a tag to compare

@github-actions github-actions released this 14 Aug 21:14

bwa-mem4 4.3.2

The command implementations build as a library target as well as a binary, so an
embedder can align in process on the same code path, and MemArgs derives Default
so the struct can be built without naming every field.


Download bwa-mem4 4.3.2

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.3.2-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.3.2-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.3.2-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.3.2-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.3.2-linux-x86_64.tar.gz
./bwa-mem4-v4.3.2-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).

bwa-mem4 v4.3.1

Choose a tag to compare

@github-actions github-actions released this 14 Aug 20:58

bwa-mem4 4.3.1

Reader off the critical path, FIFO input fixed, parallel mate-file reads, one
allocation per read, length-sorted mate-rescue batch, C libsais as the default
suffix-array backend with an optional OpenMP build, the Rust-mechanics
documentation layer on all ten crates, and the macos-14 arm64 PGO training step
fixed.


Download bwa-mem4 4.3.1

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.3.1-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.3.1-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.3.1-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.3.1-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.3.1-linux-x86_64.tar.gz
./bwa-mem4-v4.3.1-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).

bwa-mem4 v4.3.0

Choose a tag to compare

@github-actions github-actions released this 31 Jul 17:45

bwa-mem4 4.3.0

Completes the bwa-mem2 option surface. -x, -f and -1 were the three
letters this port did not accept; all three land here.

-x is not a straight port, and that is deliberate. bwa-mem2 prints WARNING: bwa-mem2 doesn't work well with long reads or contigs; please use minimap2 instead. before running any preset, and it is right: the long-read presets
only retune a short-read seed-and-extend design and cannot turn it into a
long-read mapper. So -x pacbio, -x pbref and -x ont2d are mapped by
rammap (pure Rust, minimap2-equivalent,
MIT) and produce its output, rather than faithfully reproducing a result bwa's
own code tells you not to use. The substitution is never silent: a banner on
stderr names the mapper and the preset, and the SAM header records rammap and
its version rather than bwa-mem4. -x intractg is intra-species contig
alignment, not long reads, and stays on the byte-identical path.

Everything else in this binary remains byte-identical to bwa-mem2 2.3.

bwa-mem4 index --mmi <preset|all> builds rammap's minimizer index up front;
otherwise mem builds and caches it beside the reference on first use. The
five bwa index files are untouched and stay byte-identical to bwa-mem2 index
either way.

Gates: check.sh (fmt, clippy -D warnings, workspace tests, x86_64 cross-check,
x86_64 tests under Rosetta), opt_parity.sh 64 passed / 0 failed, and the new
longread_parity.sh, the only harness here whose oracle is not bwa-mem2: it
requires identical records against the rammap binary for all three presets,
plus thread-count invariance and a check that intractg still comes out of bwa's
own code.


Download bwa-mem4 4.3.0

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.3.0-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.3.0-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.3.0-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.3.0-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.3.0-linux-x86_64.tar.gz
./bwa-mem4-v4.3.0-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).

bwa-mem4 v4.2.0

Choose a tag to compare

@github-actions github-actions released this 30 Jul 08:05

bwa-mem4 4.2.0

2.07x faster than bwa-mem2 2.3 on giab-4m, GRCh38, -t16, default -K, 3 reps:
91.01s against 188.24s, while staying byte-identical to it on 8,052,432 records.

Highlights: padding-free column range and two-row blocking in the mate-rescue
kernels, the Apple Silicon P-core cap made opt-in, a vectorised .pac unpack,
zlib-rs inflate by default, and new AVX-512BW seed-extension kernels.


Download bwa-mem4 4.2.0

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.2.0-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.2.0-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.2.0-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.2.0-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.2.0-linux-x86_64.tar.gz
./bwa-mem4-v4.2.0-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).

bwa-mem4 v4.1.1

Choose a tag to compare

@github-actions github-actions released this 23 Jul 08:23

bwa-mem4 4.1.1

Memory

  • Per-batch RAM no longer scales with the batch's seeding intermediates: the
    seeding chunk is capped and sized per thread, so the reads held concurrently
    are a fixed budget regardless of -t. Peak RSS -38% at -t16 (30.0 -> 18.5 GB on
    GRCh38 / 4M pairs), byte-identical.
  • With a fixed -K, peak RSS is now flat across thread counts: ~11.8 GB at
    -t4/-t8/-t16 with -K 10000000 (vs 12.9/15.2/18.4 at the -t-scaled default).
    The default -K = 10M x threads is unchanged, so default output stays
    byte-identical to bwa-mem2 (whose PE output itself depends on -t).

Packaging

  • Aligns the workspace's internal dependency version fields, which 4.1.0 left at
    4.0.1.

Output is byte-identical to bwa-mem2; only the SAM version string changes.
Binaries attached for linux-x86_64, linux-aarch64, macos-arm64 and macos-x86_64.


Download bwa-mem4 4.1.1

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.1.1-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.1.1-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.1.1-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.1.1-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.1.1-linux-x86_64.tar.gz
./bwa-mem4-v4.1.1-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).

bwa-mem4 v4.1.0

Choose a tag to compare

@github-actions github-actions released this 23 Jul 06:36

bwa-mem4 4.1.0: libsais-by-default index build (byte-identical, ~2.6x faster) + x86 AVX2/AVX-512 mate-rescue kernels


Download bwa-mem4 4.1.0

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.1.0-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.1.0-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.1.0-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.1.0-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.1.0-linux-x86_64.tar.gz
./bwa-mem4-v4.1.0-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).

bwa-mem4 v4.0.1

Choose a tag to compare

@github-actions github-actions released this 22 Jul 13:05

bwa-mem4 4.0.1

Performance

  • Mate rescue: the local-SW kernel's F-recurrence is reassociated so the serial
    per-column chain drops from three ops to two (F no longer waits on H). +14% on
    the rescue kernel, byte-identical, across the NEON u8, NEON i16 and AVX2 paths.
    Measured end-to-end at 4M pairs on GRCh38, -t8.

Packaging

  • Bioconda recipe plus build fixes (MSRV, macOS deployment target, libclang).

Output is byte-identical to bwa-mem2 as in 4.0.0; only the SAM version string
changes from 4.0.0 to 4.0.1. Binaries are attached for linux-x86_64,
linux-aarch64, macos-arm64 and macos-x86_64.


Download bwa-mem4 4.0.1

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.0.1-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.0.1-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.0.1-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.0.1-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.0.1-linux-x86_64.tar.gz
./bwa-mem4-v4.0.1-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).

bwa-mem4 v4.0.0

Choose a tag to compare

@github-actions github-actions released this 22 Jul 01:55

bwa-mem4 4.0.0

The first public release. bwa-mem4 is a from-scratch Rust reimplementation of
bwa-mem2 2.3 whose acceptance criterion, from the very first commit, is
byte-identity: the index files and the SAM/BAM/CRAM records it writes are
identical, byte for byte, to the patched bwa-mem2 2.3 oracle. Verified on a real
32.9x human WGS (GIAB HG002, 2x150), whole genome: byte-identical index,
byte-identical SAM over 353,517,767 single-end and 707,312,349 paired-end
records, ALT contigs included.

Output is not affected by anything in this release: it is the same bytes
bwa-mem2 produces, by construction and by gate. What changed here is the name,
the resident memory, and the paired-end speed. @PG is the one deliberate
difference (we stamp our own identity, never bwa-mem2's), and it is excluded
from every parity gate.

The name bwa-mem3 was already taken: fg-labs/bwa-mem3 is on bioconda and its
Rust bindings are on crates.io, all predating this project by about three months.
@nh13 asked us to disambiguate and he was right to. The binary, the ten
crates.io packages (all under the bwa-mem4 prefix), and the @PG PN:/VN:
stamped into every output file are renamed together, because a binary that fixes
bwa-mem2 bugs while announcing itself as bwa-mem3 would mislabel the data it
writes. The Rust library names are unchanged, so no use statement moves.
References to fg-labs/bwa-mem3 are deliberately left as-is: they point at that
project and keep its name.

With a real tab in -R, bwa-mem2's @PG line carried seven fields instead of
four and two ID: tags, one from the header and one swallowed from the read
group into CL:. That is an invalid SAM header; strict parsers such as noodles
reject it. Tabs are now escaped to the two characters \t, which is exactly the
spelling bwa accepts on the -R command line, so the field stays a faithful
record of what the user typed. Costs no byte-identity: @PG is excluded from
the parity gate.

Measured on the whole-genome index, -8 GB, byte-identical (index files unchanged
on disk; only what is loaded changed). Two changes that compose:

  • The reference is read from the 2-bit .pac (~775 MB), unpacked one base at a
    time, instead of memory-mapping the one-byte-per-base .0123 (~6.2 GB). The
    reverse-complement half is reconstructed rather than stored.
  • The BWT is read into aligned buffers instead of memory-mapped and copied. The
    mmap-then-copy held the ~10 GB file and its copy resident at once, and that
    transient, not the reference, was the real peak. read leaves the source in
    the kernel page cache, which is not counted in the process's resident set.

Neither alone moves the peak: the first is masked by the BWT transient, the
second would leave .0123 resident. Together they take the peak below
bwa-mem2's 16.8 GB.

The mate-rescue local-SW kernel (the majority of paired-end time) now computes
its substitution score with one biased table lookup on target XOR query
instead of three candidate adds and four selects, and detects dead cells with a
single high-bit test. ~15 inner-loop operations become ~11. Byte-identical
(a 2000-job property test against per-job ksw_align2, plus the whole-genome
paired-end gate). Measured +3.6% on rescue cost. The i16 path (rare on real
reads) and the AVX2 path keep the previous form; porting AVX2 is a follow-up and
is byte-identical either way.

Single-end 2.7x, paired-end 1.9x on GIAB HG002 (M4 Max). The ratio is not a
constant: it falls as thread count rises and rises with batch count, so it is
always quoted with both.

  • The i16 and AVX2 mate-rescue kernels do not yet carry the score-table change;
    they are byte-identical without it, just not as fast.
  • SA compression is not ported: it would break index byte-identity.
  • A GIAB hap.py/vcfeval concordance gate (showing byte-identity translates into
    variant-calling concordance) is future work.

Drop-in. Same command line as bwa-mem2's mem, same index format (an index
built by either tool is readable by the other). The only observable change is
the @PG line naming bwa-mem4, and lower memory. cargo install bwa-mem4 gets
the binary; prebuilt binaries for Linux and macOS, x86_64 and arm64, are
attached below with a SHA256SUMS file.

  • @nh13 (Nils Homer): the NEON SW backend and Apple Silicon tuning ported from
    his bwa-mem2#288 and his fg-labs/bwa-mem3 fork, which took the SW path from
    1.48x to ~2.2x; the ungapped fast path; the mate-rescue score-table technique
    in this release, read from his kswv; and the request to rename, which he was
    right about. Where his refused or separate work identified a trap the merged
    version avoids, that counts too.
  • The bwa-mem2 authors (Vasimuddin et al.) and Heng Li (bwa): the algorithm and
    the oracle this is measured against.

Download bwa-mem4 4.0.0

Every binary below was required to rebuild the committed testdata/tiny index
byte-identically on its own platform, and to align 500 pairs into exactly 1000
records, before this release was allowed to appear. The artifact is checked, not just
the source it came from.

File Platform
bwa-mem4-v4.0.0-linux-x86_64.tar.gz Linux x86_64
bwa-mem4-v4.0.0-linux-aarch64.tar.gz Linux arm64
bwa-mem4-v4.0.0-macos-arm64.tar.gz Apple Silicon macOS (11.0+)
bwa-mem4-v4.0.0-macos-x86_64.tar.gz Intel macOS (10.12+)
SHA256SUMS checksums for all four
tar xzf bwa-mem4-v4.0.0-linux-x86_64.tar.gz
./bwa-mem4-v4.0.0-linux-x86_64/bwa-mem4 --version

Or build from source: cargo build --release, Rust 1.96, no system packages
needed (htslib is vendored).