Skip to content

Releases: kevinjohncutler/opencodecs

opencodecs 0.7.2

Choose a tag to compare

@kevinjohncutler kevinjohncutler released this 05 Oct 23:02

opencodecs 0.7.2 is about speed. We measured opencodecs against Apple's ImageIO, NVIDIA's GPU codecs, and the Intel and AMD media engines. Most of the gaps those measurements exposed were in opencodecs itself, and they are closed here for every platform. Every decode returns exactly the same output as 0.7.1.

Faster everywhere

Change Example
AVIF decode converts YUV to RGB on several threads 4096 x 3072 4:2:0: 149 to 58 ms
PNG decode inflates in one libdeflate call, and the defilters stop waiting on the byte just written 4096 x 4096 16-bit gray: 155 to 79 ms
gzip decode goes through libdeflate 256 MiB on macOS: 438 to 367 ms (text), 807 to 646 ms (less compressible data)
New opencodecs.decode_batch decodes many zstd, deflate or LZ4 chunks in a few native calls 256 MiB of 64 KiB zstd chunks: 28 ms, against 179 ms one Python call per chunk
Local zarr and OME-Zarr arrays of zstd chunks read, decompress and place in native code, in parallel by default full 256 MiB read: 511 to 45 ms
HEIF decode gives libheif's HEVC decoder its thread count untiled 4096 x 3072: 288 to 78 ms
HTJ2K decode splits a multi-tile codestream and decodes the tiles in parallel 16 tiles: 118 to 13.5 ms

AVIF encode now defaults to speed 6. Before, it used libavif's library default, which is libaom's slowest search. The new default is about 50 times faster: 0.5 s instead of 22 s for a 12-megapixel photo. Files are about 10% larger at equal perceived quality. SSIMULACRA2 and Butteraugli scores, on photographs and on microscopy, show no perceptible difference. libavif's own avifenc also defaults to 6. imagecodecs keeps speed 0, so this is a deliberate difference from it; pass speed=0 to match imagecodecs.

Opt-in hardware backends. None of these is ever used unless you ask for it by name:

  • backend="nvimgcodec" runs JPEG, JPEG 2000 and HTJ2K on an NVIDIA GPU. An HTJ2K lossless decode is 9x faster, or 16x into pinned memory, and lossless decodes match the CPU path exactly.
  • backend="nvvideocodec" runs HEIF decode on NVIDIA's video decoder (NVDEC) and AVIF/HEIF encode on its video encoder (NVENC). Decoded pixels match libheif exactly, except in 10- and 12-bit 4:2:0 and 4:2:2 files: there about 1 sample in 100,000 can differ by 1 from the libheif 1.21.0 these wheels bundle, which rounds one default color matrix; newer libheif agrees with the GPU exactly. Tested on Ada (RTX 4090) and Blackwell (RTX 5060 Ti), which also decodes HEVC 4:2:2. The encodes are a fast mode: 14 ms for AVIF, with larger files.
  • backend="imageio" decodes HEIF on Apple's media engine on macOS: 10 to 19x faster on large single-picture files.

Install the NVIDIA dependencies with pip install 'opencodecs[gpu]'. Importing opencodecs loads none of these packages. A backend that cannot run raises an error instead of falling back to the CPU, and the first call in a process costs about a second.

Measured and not added: Intel Quick Sync and AMD VCN through VA-API, and NVIDIA's nvCOMP. The reasons are in CHANGES.rst.

The full list is in CHANGES.rst.

opencodecs 0.7.1

Choose a tag to compare

@kevinjohncutler kevinjohncutler released this 04 Oct 11:29

opencodecs 0.7.1 is about files that open in other software, not just in opencodecs.

Every encoder is now tested against a decoder that is not ours. Each of the 46 codecs that can encode writes a stream, and an independent library reads it back:

  • the format's own Python binding where one exists (zstandard, lz4, brotli, blosc2, cramjam, lerc, zfpy, pcodec, qoi, pillow-heif, the standard library);
  • imagecodecs;
  • format readers (tifffile, mrcfile, nibabel, pydicom, numcodecs, numpy).

Every writer's files are opened the same way: TIFF with each compression, pyramids, NDTiff, OME-Zarr, MRC, NIfTI, RGBE, JPEG XL animation and CZI. CZI files are additionally checked with libCZI (pylibCZIrw), czifile and aicspylibczi. A codec that gains an encoder without such a test fails the suite.

Fixes this found

  • CZI files written with the default metadata could not be opened by libCZI, so neither by czicompress nor by ZEN. The writer now always roots the metadata at ImageDocument, which libCZI requires. czi_recompress keeps the source metadata unchanged.
  • OME-Zarr stores written with compressor="blosc2" opened only in opencodecs. They now carry the metadata of ocf-blosc2, the numcodecs plugin for Blosc2, so zarr opens them with that plugin installed. Stores written the old way still read.
  • TiffWriter and the NDTiff writer raised for JPEG XL with compression_level. JPEG XL now takes level with imagecodecs' meaning: 100 or below is lossy at that quality, above 100 is lossless.
  • On Windows, a closed CZI reader raised ValueError: mmap closed or invalid instead of its own closed-reader error.

The full list, including why the CZI zstd level stays at 3, is in CHANGES.rst.

opencodecs 0.6.0

Choose a tag to compare

@kevinjohncutler kevinjohncutler released this 03 Oct 10:37

Cancelling czi_recompress

czi_recompress takes an optional should_continue: a zero-argument
callable polled once per sub-block, before that sub-block is read. When it
returns false the rewrite raises the new CziCancelled and the partial
output is removed. Existing callers are unaffected; the default is None.

It is a plain callable rather than a threading.Event so a caller can pass
a deadline, a signal flag or a cancellation token just as easily, and nothing
in the writer needs to know which. CziCancelled subclasses
CziWriterError, so code already catching that keeps working while code
that cares can tell "the user stopped it" from "the file is broken" without
matching on message text.

Why it exists: a caller replacing a subprocess (ZEISS's czicompress) with
this function loses the ability to kill a process mid-file, and a Cancel
button that only takes effect between files is a visible regression on a
slide scan. Polling per sub-block takes effect within roughly workers
sub-blocks, which on a 481-sub-block scan is sub-second.

An interrupted rewrite now removes dst on any exception, not only
cancellation: a failed verification or a KeyboardInterrupt too. A CZI
missing sub-blocks still has a valid header and directory, so readers accept
it and the loss is silent; truncated-but-plausible is worse than absent.

CZI reading

  • CziReader(path, index=...) reads pixels from a payload index (offsets,
    sizes, compressions, plane shape, pixel type) without reading the file
    header, the directory or any sub-block header, and payload_index()
    returns one from an open reader. A caller that stores the index with its
    own metadata reads planes with no metadata round trips over a network
    share.
  • read(indices=[...]) reads chosen sub-blocks in parallel into one
    stack. attachments() and read_attachment(name) return embedded
    files (Thumbnail, TimeStamps, Label ...) exactly as stored.
  • Files on disk are read with positional reads instead of a memory map: a
    cold plane off a network share was paged in a cluster at a time, and one
    exact-range read (split into concurrent pieces of about 1 MiB when only a
    few planes are read) fetches it in a fraction of the time. The file
    header is fetched when first needed, the directory comes in one request,
    and a sub-block's header and payload in one. The directory and metadata
    are read on first use, so a corrupt directory now raises CziError on
    first use rather than in the constructor.
  • Uncompressed stacks are read on several threads. They were split into 8
    MiB tasks, the size zstd reads use, so nine 2 MiB planes made three
    tasks, which the worker policy ran on one thread; they now go one plane
    per 2 MiB task, and two tasks are enough to share. Warm, that stack reads
    1.63x faster than a reader copying from an mmap on 128-core Linux, where
    it had been 0.59x as fast (1.51x and 0.88x on a 20-core Mac). Slide
    regions that need two or three tiles decode them in parallel for the same
    reason.
  • JPEG XR (compression 4, most slide scans) decodes through the copy of
    jxrlib maintained inside ZEISS's libCZI (BSD-2-Clause, Microsoft), which
    ZEISS reworked for speed. No compiler flags brought the upstream 2019.10.9
    release within 10% of aicspylibczi on x86-64 Linux; with libCZI's copy a
    slide region read inside one tile takes 0.99x aicspylibczi's time there
    and is 1.08x faster on a 20-core Mac (upstream: 0.90x and 0.93x). The
    wheels build it static, from a pinned libCZI commit, with NDEBUG
    (upstream's Makefile left 68 assert() calls in the decode loop) and
    link-time optimization. A jxrlib built by bench/build_codec_libs.sh
    is now linked by path, so a system jxrlib earlier on the library path
    can no longer replace it. Upstream jxrlib (Windows wheels, distribution
    packages) still works.
  • The hi/lo byte unshuffle of zstd sub-blocks uses a byte loop for 2-byte
    pixels under GCC and Clang, which GCC vectorizes better: 1.46x faster for
    a 1000 x 1000 plane and 1.11x for 2000 x 2000 on x86-64 Linux. Clang ties;
    MSVC keeps the word form.
  • Payload buffers are shared and reused, and the read pool holds at most
    the 32 threads one read uses (it was twice the CPU count).

TIFF reading

  • A multi-page TIFF decodes each page straight into its slice of one
    output, instead of decoding pages separately and stacking them, a second
    copy of the whole stack. A stack stored as plain pixels, end to end in
    the file (what writers produce for a contiguous series), is read as one
    span by parallel positioned reads: a 23-page 2000 x 2000 uint16 stack in
    10.8 ms where tifffile takes 33.8 ms (128-core Linux, warm); page by page
    it took 59.6 ms. TiffPage.asarray takes out=.
  • Decode workers are sized by whichever asks for more: 1 MiB of output each,
    or 256 KiB of compressed input each. Output alone held a deflate image
    with 6.8 MB of compressed data in 31 strips to 7 workers.

opencodecs 0.5.0

Choose a tag to compare

@kevinjohncutler kevinjohncutler released this 02 Oct 20:42

opencodecs 0.5.0 brings two things: speed, above all on Windows, and a compatibility pass over every codec opencodecs shares with imagecodecs.

Speed. The speed changes are the same code on every operating system and compiler. Against the 0.4.0 wheels:

  • Delta and XOR decode, distances 2 to 4: 4.3 to 5.1x on Windows, 6.1 to 6.7x on Linux and 2.2 to 2.5x on the Mac.
  • LZW decode: 2.8 to 3.2x.
  • Bit-packed unpack: 1.7 to 4.6x.
  • Byte shuffle and bitshuffle on Windows: 1.7 to 2.3x.
  • BC7 decode: 1.14 to 1.55x.
  • Linux wheels now build their codec libraries from source. Earlier Linux wheels shipped AlmaLinux 8's zstd 1.4.4, lz4 1.8.3, libwebp 1.0 and openjpeg 2.4, plus a libdeflate built by GCC 8. Serial reads of deflate TIFFs on Linux are 1.2 to 1.5x faster. With the old libraries, 58 compatibility tests failed against the Linux wheel; they now pass.

The README's tables compare opencodecs with tifffile and with imagecodecs 2026.8.16 on macOS, Linux and Windows.

Compatibility. An audit found 54 cases where opencodecs and imagecodecs disagreed, and all of them are fixed under one rule:

  • A published specification wins.
  • A library's own format is written bare, with no opencodecs header, and files that earlier versions wrote still read.
  • Where no specification decides, opencodecs matches imagecodecs.
  • No parameter is silently ignored and no data is silently lost.

Several of these fixes change the bytes opencodecs writes: floatpred now implements TIFF predictor 3, float delta works on bit patterns (pass legacy_float=True to read old streams), and rcomp, aec, pcodec, SZ3 and SPERR drop their private headers. CHANGES.rst names each one.

Removed: opencodecs.tifffile_patch. tifffile 2026.8.23 and later refuse codec functions from other modules. Use opencodecs.read, opencodecs.tiff_imwrite and opencodecs.TiffWriter instead.

Known issues

  • BC3 decode is 0.80x of 0.4.0 on the Mac, because its alpha now rounds as the Khronos specification defines.
  • Built from source against libheif 1.23, HEIF alpha stored at a different depth from the color is rescaled, and its low bits are lost. The wheels ship libheif 1.21, which raises instead.
  • Built from source against OpenJPH older than 0.31, HTJ2K refuses a quality level. Older system OpenJPH and SZ3 builds also write bytes that differ from the wheels.
  • opencodecs.read cannot open lossless YCbCr JPEG TIFFs; it raises.
  • SZ3 can crash on a damaged payload.
  • On Linux, SZ3 and SPERR output can differ in its bytes from imagecodecs' own builds of the same libraries. Both are lossy float compressors.
  • The Linux wheels decode AVIF and HEIF but do not encode them, as in earlier releases.

The full list is in CHANGES.rst.

opencodecs 0.4.0

Choose a tag to compare

@kevinjohncutler kevinjohncutler released this 29 Sep 11:04

How opencodecs spends threads, especially when several of your threads call it at once, and one place for each thing it does differently by operating system. Figures compare against 0.3.1 on a 20-core arm64 Mac and a 64-core x86-64 Linux workstation: both builds in fresh processes, alternating, each process one sample, identical decoded pixels required.

  • One shared pool. run_batched, map_batches and map_bounded (under TIFF, DICOM, EER, FITS, HDF5, CZI, Zarr and more) built a ThreadPoolExecutor per call. They now share one process-wide pool. A 144-tile TIFF reads 1.34x faster on Linux (9.2 to 6.9 ms) and 1.11x on the Mac.
  • Concurrent callers share instead of multiplying. A call counts as in flight for its whole duration, and an automatic worker count is its share of what the calls in flight divide: the pool helpers' workers, which hand the GIL back and forth, share 32 between calls (each call still takes at most 16), and native codec threads share the CPU count. A lone call keeps full width and an explicit numthreads is honored as given. With the TIFF work below, two and eight threads reading a tiled deflate TIFF at once: 3.46x and 4.63x on Linux, 3.61x and 4.01x on the Mac; eight JPEG XL readers 1.86x on Linux and 1.03x on the Mac.
  • JPEG XL sizes its threads to the image. One-shot decode used libjxl's default of one thread per hardware thread, all created for each call whatever the image size. It now reads the size from the header and uses about one thread per 64K pixels, up to 20. 512 x 512: 2.54x faster on Linux (12.9 to 5.1 ms, 5 threads instead of 129) and 1.24x on the Mac; 2048 x 2048: 1.10x for one caller and 1.56x for eight on Linux, parity on the Mac.
  • AVIF decode uses up to 8 threads by default instead of one per core: AV1 decodes in parallel across tiles, and past 8 threads there was nothing left to gain. 1.07 to 1.13x on Linux, 1.00 to 1.03x on the Mac. numthreads=0 still means every core.
  • TIFF segments decode, un-predict and land in one native call. A tile used to be inflated in one call, un-predicted in a second and copied into the output with numpy, with a GIL handoff between each, which is what capped concurrent readers. Now a batch of segments goes through all three under one GIL release (_tiff.decode_segments_into), for no compression, deflate, zstd, LZW and PackBits with predictors 1 to 3, and a whole strip decompresses straight into its output rows. zlib and zstd are reached through a C table their own modules export, so _tiff links neither. Separate planes, bilevel and byte-swapped files keep the general path, as does any segment the native call rejects, which also keeps its exact error. Against 0.3.1, a tiled 4096 x 4096 uint16 image with the horizontal predictor: one reader 2.94x on the Mac and 2.38x on Linux, eight readers 3.65x on both; a 144-tile image 1.98x and 1.96x; the same pixels in strips 1.28x and 1.37x for one reader, 1.59x and 1.94x for eight. Serial reads are unchanged or faster.
  • Uncompressed TIFF strips laid out end to end go straight into the output (core.io.read_file_into): positioned reads split across the shared pool, where one thread used to copy them out of the file mapping. One reader 2.29x on the Mac and 3.79x on Linux; with numthreads=1, 1.65x on the Mac and level on Linux; eight readers 1.66x on Linux and level on the Mac, where both versions copy from the mapping. One choice depends on the kernel, and is named for it: macOS copies one thread's worth of bytes out of the reader's file mapping faster than it reads them, and Linux reads faster, alone and with eight readers alike, so that is what each does.
  • The README compares TIFF reads directly with tifffile: 2.7 to 4.3x faster on the Mac and 4.0 to 7.7x on Linux for one reader of a tiled compressed image, up to 13.4x for eight concurrent readers.
  • Ultra HDR runs its kernels on the shared pool and its encode and decode steps on a persistent pool of their own, instead of building up to five pools per call: encode 1.06x on the Mac and 1.30x on Linux, decode 1.09x on both.
  • Each operating-system difference is named once. setup.py states its platform conventions in one place and links the library files it finds by path on every platform; descriptor reads and writes, Windows' binary-mode flag and uncached opens each have one implementation; capabilities are probed instead of inferred from the platform. A test lists the nine operating-system branches left in the package and setup.py, each with the reason it cannot be a probe, and fails on a new one.
  • A CZI file's compression can be rewritten without losing its container. czi_recompress(src, dst) re-encodes every sub-block and keeps its dimensions, scene, mosaic index and pyramid level, and the metadata XML; write_frame and write_many take a sub-block's own dimension list, and subblock_dims builds one from a reader's entry. libCZI reads the output of a 481 sub-block slide scan and agrees with the input. 185 MB takes 122 ms against 340 ms for ZEISS's czicompress, and 297 ms with verify=True, which decodes every sub-block back before it is written.
  • Windows wheels use libdeflate, as macOS and Linux wheels already did. conda-forge names its import library deflate.lib, which the old probe did not recognize, so Windows fell back to zlib. On a Windows VM, against 0.3.1: deflate encode 2.79x and decode 1.79x, PNG encode 2.82x (decode stays on zlib, level), a tiled deflate TIFF read 1.33x; output at the default level is 0.7% (deflate) to 1.5% (PNG) larger, the same trade macOS and Linux make.
  • import opencodecs on macOS ran mount once per extension, about 4.5 ms each; the mount table is now read once, about 200 ms per import.
  • Fix: the DICOM codec dropped numthreads, so numthreads=1 still decoded on several threads.
  • Fix: the NDTiff writer counted the buffers after a partial writev twice, so its position ran ahead of the file, and tiff_reader took one os.pread as complete even when it returned short.
  • Fix (building from source): with a per-user cache of source-built libraries present, a Windows build handed MSVC .so paths and -Wl,-rpath flags; the MozJPEG extension on Linux had no rpath to its own library, so the loader could take the system's libturbojpeg for it; bench/build_codec_libs.sh passed -mcpu=apple-m1 on Intel Macs.
  • Fix: pools created before os.fork() hung in the child, which inherits the pool object but none of its threads (the first submit in a forked DataLoader-style worker waited forever). A forked child now builds fresh pools.

opencodecs 0.3.1

Choose a tag to compare

@kevinjohncutler kevinjohncutler released this 28 Sep 08:02

A CZI read scheduling fix, and a clear refusal for the files that cannot stack.
Figures compare against 0.3.0 on the same machines.

CZI whole-stack reads hand out work by bytes

read derived its batch size from the worker count, giving each worker
ceil(n/workers) consecutive sub-blocks. That makes the read wait on that many
sub-blocks even when most workers have already finished: nine 8 MiB sub-blocks
over eight workers is two rounds, and measured 15.1 ms against 7.9 ms for nine
one-sub-block tasks. Tasks are now sized by output bytes, about 8 MiB each, so a
large sub-block gets a task to itself and small ones ride together. That keeps
the reason batching existed, which is that a 0.7 MiB tile cannot pay for its own
future, without paying the rounding.

_READ_MAX_WORKERS goes 8 to 32: the old value was measured against the
dispatch above, where extra workers could not help. With byte-sized tasks, 20 to
32 measured fastest on both a 20-core arm64 Mac and a 128-core x86-64 Linux
host, and past 32 the Linux host gives it back to memory-bandwidth contention. A
semaphore keeps the resolved worker count a real bound when the byte sizing
yields more tasks than that.

Eight files of 2000 x 2000 uint16 ZSTDHDR sub-blocks, read back to back: 22.1 to
about 15 ms per file on macOS, which moves a whole-stack read from a little
behind aicspylibczi to about 1.3x ahead of it.

A mixed sub-block CZI says so, instead of failing inside the codec

read stacks every sub-block into one array, and sized that array from the first
sub-block, which is wrong for any file whose sub-blocks differ: a pyramidal slide
scan interleaves down-scaled sub-blocks with full-resolution ones, and a mosaic
can clip its edge tiles. The decode destination check then failed with an element
count the caller never chose, naming neither the file's shape mix nor a way
forward. It now refuses up front, listing the shapes it cannot stack and the APIs
that do work (CziPyramidReader.read_region, entries_at_level, read_tile and
iter_tiles), and the new CziReader.is_uniform lets a caller branch before
asking. Decoding was never at fault: JPEG XR sub-blocks are pixel-exact against
aicspylibczi.

v0.3.0

Choose a tag to compare

@kevinjohncutler kevinjohncutler released this 28 Sep 04:33

Mostly a speed release, plus JPEG XR support for Zeiss slide scans. There are no breaking changes: nothing public was removed or renamed and no default changed.

Faster

Measured against 0.2.0 on the same inputs, both versions in fresh processes, counted only where both returned identical pixels. Medians on a 20-core Apple silicon Mac and a 64-core x86-64 Linux workstation.

Workload Mac speedup Linux speedup
CZI: build the pyramid reader 51x 35x
CZI: open a slide and read the first crop 6.5x 4.6x
CZI: 50 disjoint crops with read_regions 5.5x 5.9x
CZI: read a 20 x 20 tile mosaic 2.4x 3.0x
CZI: write 8 zstd frames with write_many 4.1x 4.6x
Zarr: whole 4096 x 4096 array, zstd, num_workers=8 3.4x 3.7x
Zarr: same, zlib 3.1x 4.7x
blosc2 from 8 threads, decode / encode 6.5x / 6.6x 6.2x / 7.1x

The full table, with times, is in the README. Zarr regional reads stay serial unless you pass num_workers.

New

  • JPEG XR in CZI. Most Zeiss slide scans store their tiles as JPEG XR (CZI compression 4). A new _jpegxr extension over jxrlib decodes them, and every tile of a 20684 x 32751 Axioscan slide matches czifile exactly. It ships in every wheel, statically linked, so there is nothing extra to install.
  • Parallel Zarr regional reads with num_workers on read_region or OmeZarrArray.
  • CZI: read(out=...) accepts any array including numpy.memmap; read_regions decodes each tile shared by several boxes once, straight into every box; an opt-in decoded-tile cache; write_many compresses on workers while the file stays byte-identical to sequential writes; and opt-in CziWriter(background_encode=True) compresses the current frame while you prepare the next.
  • Caller-owned outputs. Native codecs decode into out= buffers and sinks you provide, seven byte codecs have stateful streaming iterators, and PNG has native row sessions.
  • One bounded pipeline for readers and writers. opencodecs.core.pipeline.map_bounded schedules parallel work in order under a byte budget and a shared worker budget, with per-worker scratch and opt-in bit-exact verification of written segments.

Fixed

  • CZI pyramids on real slides. CziPyramidReader took sub-block positions as level pixels, but Zen stores them in full-resolution slide coordinates, often far from zero and negative, so a real slide reported every level as (N, 0). Levels now count from reader.origin, matching czifile, and positions are divided by each level's scale.
  • CZI mosaic overlaps were composed in directory order. They now compose in mosaic-index order, as libCZI and czifile do; on the corpus slide that changes 8.9 million overlap pixels at level 0.
  • blosc2 encoding under threads selected its compressor through process-global state, so concurrent encodes could use another thread's compressor. Each call now uses its own context. Chunks with item size 1 and shuffle carry slightly different header flags; old and new chunks decode in both versions.

Build

  • Every wheel ships the same 40 compiled extensions: the 39 of 0.2.0 plus _jpegxr.
  • SZ3 3.3.2, since upstream deleted the 3.3.1 tag.

Full detail in CHANGES.rst.

v0.2.0

Choose a tag to compare

@kevinjohncutler kevinjohncutler released this 10 Sep 11:09

Breaking

avif.decode() on an image sequence and webp.decode() on an animation now return every frame as a stack, not the first image. That is what gif has always done here and what imagecodecs returns for the same files; returning frame 0 silently discarded the rest. A still image is untouched, and out= is refused for a sequence rather than filled with frame 0. heif deliberately does not follow: its multi-image files are unrelated images, not a time sequence.

New

  • Whole-slide pyramids. Readers for DICOM VL Whole Slide Microscopy (a slide is a series of instances, ordered by total pixel matrix size) and for Olympus VSI/ETS, both behind the shared PyramidReader interface.
  • Offset reads. Readers take a byte range instead of swallowing the whole file, so a 300 GB slide no longer has to fit anywhere.
  • URLs everywhere. Every codec and reader accepts an http(s) URL alongside a path, bytes, memoryview, mmap, BytesIO or an open file.
  • Parallel decode wherever the pieces were already independent, with a cfitsio thread-safety fix underneath it (three file-scope statics needed thread-local storage before HCOMPRESS tiles could be threaded at all).
  • AVIF tiling and encoder choice: tile_cols_log2 / tile_rows_log2 / auto_tiling, yuv_format, codec and codec_options. Tiling defaults to 4x4 once the long axis reaches 1024 px, measured ~2x faster encode for +4.4% bytes at unchanged PSNR.

Fixed

  • Five codecs were holding the GIL through their decode. _jpeg2k, _mozjpeg, _charls, _zfp and _openjph all declared nogil on the declaration, which means "safe to call without the GIL", not "called without it". They now release it, so threading them actually helps.
  • Four advertised codecs shipped in no wheel. jpegls (CharLS), snappy, gif and deflate(backend="isal") are in the README codec table, and reading the published 0.1.13 wheels back shows _charls, _isal and _snappy absent on every platform, _gif present only on macOS. Their libraries appeared in no CI install path, so the header probe dropped each extension, pytest.importorskip removed the tests that would have noticed, and both the build and the suite stayed green. Every platform now ships an identical 39 extensions, and ci/check_wheel_contents.py fails the build if any of them goes missing again.
  • macOS wheels were unbuildable and nothing reported it, because Homebrew bottles follow whichever image GitHub calls macos-latest and that moved to macOS 26, leaving delocate to reject a macosx_15_0 wheel. The runner is pinned to macos-15 so the supported floor stays a decision rather than drifting upward with each image roll.
  • _plio bypassed the shadow-copy loader and hung the test suite on macOS.
  • Vendored cfitsio was missing two upstream bounds checks.

Also

Our own EER decoder (1.7x faster) and TIFF LZW encoder (1.4x faster); RGBE decode 1.7x faster after auditing every vendored source against upstream; no imagecodecs-derived Cython source remains, with attribution restored where it does apply.

Full detail in CHANGES.rst.