Repository navigation
Release 4.14.0
Changes from 4.13.1 to 4.14.0
Python-Blosc2 4.14.0 adds remote tables, shared caches, native HDF5 range reads,
and indexed queries over remote data. It bundles C-Blosc2 3.3.5.
Remote arrays, tables, and hierarchies
blosc2.open()returnsRemoteArray,RemoteCTable, orRemoteStore
according to the selected object.path=selects a node inside a container;
dataset=remains an alias. Remote Zarr v2/v3 groups now open as stores.RemoteCTablereads B2Z CTables, single-file Parquet, and PyTables tables
on demand. Column projection avoids unrelated data; Parquet reads and caches
individual physical fields by row group.lazy=Falseimports Parquet eagerly.- Multi-column reads overlap independent requests, bounded by
max_concurrency
and transport-buffer budgets. Decoding and cache publication stay serialized. RemoteObjectprovides the common source, attributes, traffic, cache,
reference-saving, and lifetime API.RemoteStore,TreeStore, andDictStore
provide.infosummaries of their contents without opening leaf readers.cache_dir=retains metadata and accessed data across sessions. Add
shared_cache=Truefor simultaneous processes; all users of that cache
directory must enable sharing. Persistent caches also support local sources.save()writes a reference with retained cache data, without fetching missing
data.include_cache=Falsecreates a cold reference.materialize()produces
independent local data; tablecopy(),to_b2z(), andto_b2d()do likewise.TreeStorecan persistRemoteStoremounts. Reference exports preserve them;
materialization expands them, preserves group attributes, and rebuilds indexes.- Standalone arrays/tables and root stores support
refresh(). Successful refresh
invalidates previously obtained child handles; callers must retrieve them again.
HDF5, PyTables, and indexes
- Native HDF5 metadata indexes replace Kerchunk translation. Remote reads no
longer require Kerchunk, Zarr, or numcodecs. Uncompressed, deflate/gzip,
shuffle, and Blosc2 chunks are decoded directly; other filters use h5py.
Cached metadata avoids repeated scans on warm opens. - Local and remote PyTables tables open as
RemoteCTablewithout requiring
PyTables installed. Supported persisted PyTables indexes are imported as
Blosc2 OPSI indexes and cached for subsequent queries. - Remote CTable queries use persisted SUMMARY, FULL, PARTIAL, OPSI, BUCKET,
and list-membership indexes. Unsupported query shapes fall back to scans.
Index construction remains a local operation.
Lists and Arrow interchange
- ListArray supports nullable items and recursive nesting, up to 32 list levels.
contains()andoverlaps()provide row predicates; optional
kind="membership"indexes accelerate flat scalar-list queries. - New list schemas default to
batch_rows=2048. PassNonefor caller-managed
batching. Existing schemas retain their stored batching behavior. - Arrow/Parquet exports default to 65,536 rows per batch, improving throughput
at the cost of larger temporary allocations and default Parquet row groups.
Imports and persisted list batches retain the 2,048-row default. - Parquet list imports now default to MessagePack, matching Arrow imports.
Uselist_serializer="arrow"or CLI--list-serializer arrowto opt in.
Dense local Arrow-backed list exports keep their Arrow buffers where possible.
Bug fixes
- Fixed
asarray()data corruption for arrays larger than 16 MB with block
partitions that divide chunks but are not C-contiguous (#723, PR #724).
Thanks to @jeandet. - Fixed concurrent NDArray reads sharing mutable decompression state, including
views, and released the GIL while waiting for read/frame locks (#555, PR #713).
Thanks to @Johnny-Kao. - Fixed stale remote index handles, cache cleanup, failed-open resource leaks,
double-closing fsspec sessions, and Windows path/refresh handling. - Hardened HDF5 metadata validation and bounded deflate decoding. Preserved
query strings in dataset URLs and NumPy attributes in metadata round trips. - Fixed Parquet list-cell shape preservation, list predicates, and cached-read
exception handling. Arrow imports preserve nested item nullability. - Preserved metadata-only groups during tree traversal and materialization;
handled root entries in store summaries without looping. - Added tested CTable query and sorting examples (#650, PR #711).
Thanks to @armutlutost.
Compatibility notes
hdf5_indexreplacesrefsinblosc2.open()andRemoteArray. It accepts
a native index dictionary or JSON index path; legacy Kerchunk maps are rejected.
Thehdf5extra now requires only h5py and hdf5plugin.RemoteCTable.save()now saves a reference and returns its path. Use
materialize()for the former full-copy behavior. LocalCTable.save()is
unchanged. Replacing a reference destination requiresoverwrite=True.- Shared cache factories default to a 256 MiB aggregate compressed-payload
budget, matching ordinary opens. Passmax_cache_bytes=Nonefor unlimited
retention. This does not limit peak RAM or total disk use. - Shared caches require lazy access; explicit
lazy=Falseis rejected.
Ordinary table/store disk caches remain exclusive to one owner at a time.