Skip to content

Release 4.14.0

Choose a tag to compare

@FrancescAlted FrancescAlted released this 28 Sep 11:37
· 90 commits to main since this release

Changes from 4.13.1 to 4.14.0

Python-Blosc2 4.14.0 adds remote tables, shared caches, native HDF5 range reads,
and indexed queries over remote data. It bundles C-Blosc2 3.3.5.

Remote arrays, tables, and hierarchies

  • blosc2.open() returns RemoteArray, RemoteCTable, or RemoteStore
    according to the selected object. path= selects a node inside a container;
    dataset= remains an alias. Remote Zarr v2/v3 groups now open as stores.
  • RemoteCTable reads B2Z CTables, single-file Parquet, and PyTables tables
    on demand. Column projection avoids unrelated data; Parquet reads and caches
    individual physical fields by row group. lazy=False imports Parquet eagerly.
  • Multi-column reads overlap independent requests, bounded by max_concurrency
    and transport-buffer budgets. Decoding and cache publication stay serialized.
  • RemoteObject provides the common source, attributes, traffic, cache,
    reference-saving, and lifetime API. RemoteStore, TreeStore, and DictStore
    provide .info summaries of their contents without opening leaf readers.
  • cache_dir= retains metadata and accessed data across sessions. Add
    shared_cache=True for simultaneous processes; all users of that cache
    directory must enable sharing. Persistent caches also support local sources.
  • save() writes a reference with retained cache data, without fetching missing
    data. include_cache=False creates a cold reference. materialize() produces
    independent local data; table copy(), to_b2z(), and to_b2d() do likewise.
  • TreeStore can persist RemoteStore mounts. Reference exports preserve them;
    materialization expands them, preserves group attributes, and rebuilds indexes.
  • Standalone arrays/tables and root stores support refresh(). Successful refresh
    invalidates previously obtained child handles; callers must retrieve them again.

HDF5, PyTables, and indexes

  • Native HDF5 metadata indexes replace Kerchunk translation. Remote reads no
    longer require Kerchunk, Zarr, or numcodecs. Uncompressed, deflate/gzip,
    shuffle, and Blosc2 chunks are decoded directly; other filters use h5py.
    Cached metadata avoids repeated scans on warm opens.
  • Local and remote PyTables tables open as RemoteCTable without requiring
    PyTables installed. Supported persisted PyTables indexes are imported as
    Blosc2 OPSI indexes and cached for subsequent queries.
  • Remote CTable queries use persisted SUMMARY, FULL, PARTIAL, OPSI, BUCKET,
    and list-membership indexes. Unsupported query shapes fall back to scans.
    Index construction remains a local operation.

Lists and Arrow interchange

  • ListArray supports nullable items and recursive nesting, up to 32 list levels.
    contains() and overlaps() provide row predicates; optional
    kind="membership" indexes accelerate flat scalar-list queries.
  • New list schemas default to batch_rows=2048. Pass None for caller-managed
    batching. Existing schemas retain their stored batching behavior.
  • Arrow/Parquet exports default to 65,536 rows per batch, improving throughput
    at the cost of larger temporary allocations and default Parquet row groups.
    Imports and persisted list batches retain the 2,048-row default.
  • Parquet list imports now default to MessagePack, matching Arrow imports.
    Use list_serializer="arrow" or CLI --list-serializer arrow to opt in.
    Dense local Arrow-backed list exports keep their Arrow buffers where possible.

Bug fixes

  • Fixed asarray() data corruption for arrays larger than 16 MB with block
    partitions that divide chunks but are not C-contiguous (#723, PR #724).
    Thanks to @jeandet.
  • Fixed concurrent NDArray reads sharing mutable decompression state, including
    views, and released the GIL while waiting for read/frame locks (#555, PR #713).
    Thanks to @Johnny-Kao.
  • Fixed stale remote index handles, cache cleanup, failed-open resource leaks,
    double-closing fsspec sessions, and Windows path/refresh handling.
  • Hardened HDF5 metadata validation and bounded deflate decoding. Preserved
    query strings in dataset URLs and NumPy attributes in metadata round trips.
  • Fixed Parquet list-cell shape preservation, list predicates, and cached-read
    exception handling. Arrow imports preserve nested item nullability.
  • Preserved metadata-only groups during tree traversal and materialization;
    handled root entries in store summaries without looping.
  • Added tested CTable query and sorting examples (#650, PR #711).
    Thanks to @armutlutost.

Compatibility notes

  • hdf5_index replaces refs in blosc2.open() and RemoteArray. It accepts
    a native index dictionary or JSON index path; legacy Kerchunk maps are rejected.
    The hdf5 extra now requires only h5py and hdf5plugin.
  • RemoteCTable.save() now saves a reference and returns its path. Use
    materialize() for the former full-copy behavior. Local CTable.save() is
    unchanged. Replacing a reference destination requires overwrite=True.
  • Shared cache factories default to a 256 MiB aggregate compressed-payload
    budget
    , matching ordinary opens. Pass max_cache_bytes=None for unlimited
    retention. This does not limit peak RAM or total disk use.
  • Shared caches require lazy access; explicit lazy=False is rejected.
    Ordinary table/store disk caches remain exclusive to one owner at a time.