Skip to content

0.14.0

Latest

Choose a tag to compare

@github-actions github-actions released this 22 Sep 16:13
· 1 commit to dev since this release
13aa028

This release marks a major transition for hdmf-zarr, updating the core storage format from Zarr v2 to the new Zarr v3 standard. As part of this transition, hdmf-zarr has adopted the unified Zarr v3 storage convention developed in collaboration with Zindi and LINDI.

Migration Notes (Zarr V2 to V3)

  • Zarr V3 Standard: The primary ZarrIO and NWBZarrIO classes now write and read Zarr v3 format exclusively, using zarr-python v3.
  • Legacy Zarr V2 Support: Legacy Zarr v2 files cannot be read by ZarrIO directly. Instead, use the newly added ZarrV2IO and NWBZarrV2IO backends to read them. Attempting to open a v2 file with ZarrIO will now raise a clear error directing users to the v2 classes.
  • Migration Path: To migrate existing files to the new format, you can use NWBZarrV2IO.export_to_v3(...) or the one-shot NWBZarrV2IO.convert_to_v3(...) static helper to efficiently convert NWB Zarr v2 files to Zarr v3.
  • Forward Compatibility: Files written by this release are Zarr v3 and cannot be opened by hdmf-zarr 0.13 or earlier, which require zarr<3. Upgrade to hdmf-zarr 0.14 in order to read Zarr v3 files written with hdmf-zarr.

Changes

The core functionality has been overhauled to transition to Zarr V3, including structural, encoding, and reading changes.

  • Updated ZarrIO to use Zarr V3: The primary ZarrIO and NWBZarrIO classes now write and read Zarr v3 format exclusively, using zarr-python v3. (see list of contributors and related pull requests below)
  • Adopted New Storage Convention: Adopted the unified Zarr v3 storage convention shared with Zindi/LINDI. Links are stored in a _LINKS attribute on the parent group, dataset types in _DTYPE and _REFERENCE_FIELDS. Scalar datasets, including scalar object references and scalar compound datasets, are stored as zero-dimensional arrays with shape (), matching HDF5IO, and are read back as scalars. Attribute references are formatted as {"_REFERENCE": {"source", "path"}}, while references in datasets are plain target path strings. ZarrReference removes unused object_id and source_object_id keys. A Zarr v3 store generated by Zindi from an HDF5 NWB file can now be read directly with NWBZarrIO. @bendichter #336 #325
  • Aligned ZarrDataIO with the Zarr v3 codec API: The compressor argument is renamed to compressors, the serializer (ArrayBytesCodec) slot is now directly accessible, and filters only applies to ArrayArrayCodec. @h-mayorquin #369
  • Support dataset sharding via ZarrDataIO: Added new shards argument to ZarrDataIO to define custom sharding properties, including support for automatic sharding and guarded parallel iterator writes. @h-mayorquin @alejoe91 #374 #372
  • Requirements Changes: Bumped minimum required Python version to 3.12 and Zarr-Python dependency to >=3.4.0 due to backwards-incompatible changes in Zarr v3 structured data types and bug fixes in Zarr-Python 3.4.0. Also bumped hdmf to >=6.2.0, numpy to >=2.2.3 (earlier versions write bytes in string datasets as their Python representation, e.g., "b'ok'", instead of decoding them), and pynwb to >=4.2.0, and removed the numcodecs<0.16.0 upper bound. @oruebel @rly #373 #382 #402 #404
  • Removed the synchronizer argument: ZarrIO and NWBZarrIO no longer accept a synchronizer argument, and the ZarrIO.synchronizer property is removed. zarr-python v3 does not provide the synchronizer mechanism this wrapped. Passing synchronizer=... raises TypeError.
  • Removed the object_codec_class argument: ZarrIO and NWBZarrIO no longer accept an object_codec_class argument, and the ZarrIO.object_codec_class property is removed. References, cached specs, and compound datasets are serialized as JSON in StringDType arrays, so there is no object codec to select. Passing object_codec_class=... raises TypeError.
  • Removed the extensions argument (breaking): NWBZarrIO and NWBZarrV2IO no longer accept an extensions argument, matching the removal of the same argument from NWBHDF5IO in PyNWB 4.0.0. Load the cached namespaces from the file (the default, load_namespaces=True) or pass a prebuilt manager instead. Passing extensions=... raises TypeError. @rly #399
  • Supported store classes: SUPPORTED_ZARR_STORES covers LocalStore, FsspecStore, and any other zarr v3 Store subclass. Reading a remote store requires fsspec, which pip install hdmf-zarr[full] installs. DirectoryStore, NestedDirectoryStore, and TempStore do not exist in zarr-python v3, so code that constructs one needs to use LocalStore instead.
  • Remote write: Passing storage_options is supported only with mode="r". Writing an NWB Zarr file directly to a remote store is not currently supported and raises ValueError.

Added

Added support for reading and converting existing data using Zarr V2:

  • Zarr V2 Legacy Read Support: Added ZarrV2IO and NWBZarrV2IO read-only backends for interacting with legacy Zarr v2 files. Opening a Zarr v2 file with ZarrIO directly will now raise a clear error message directing users to the appropriate legacy reader. @alejoe91 #349
  • Migration Helpers: Added helpers NWBZarrV2IO.export_to_v3 and NWBZarrV2IO.convert_to_v3 to conveniently convert legacy Zarr v2 files to Zarr v3. @alejoe91 #349

Fixes

Addressed the following bugs that are independent of the migration to Zarr V3:

  • Fixed bug where ZarrIO.generate_dataset_html would raise an error when called with a non-Zarr object. @oruebel #355
  • Fixed bug where a compound dtype field declared with the spec type uint (or short) was written as float64. @ehennestad #365
  • Fixed bug where writing a scalar dataset with a compound dtype, such as ElectrodeGroup.position, raised IndexError. @rly #277
  • Reading or writing a scalar dataset with a compound dtype that has a reference field raises NotImplementedError. This combination is not supported. @rly #391
  • Fixed the S3 streaming tutorial catching every exception from its read, which let the gallery tests pass without reading the file. It now skips the read only when fsspec is not installed or the network is unavailable. @rly #400

Contributors

This release was made possible by the efforts of @bendichter, @alejoe91, @h-mayorquin, @ehennestad, @rly, and @oruebel.

Related Pull Requests

  • #325 : Migrate hdmf-zarr from zarr-python v2 to v3
  • #336 : Adopt unified Zarr v3 convention for hdmf-zarr/Zindi interop
  • #363 : Fix Zarr-to-HDF5 export deadlock
  • #365 : Resolve compound fields declared uint instead of writing them as float64
  • #366 : Review and propose changes to v3 integration
  • #367 : Rebuild a resolved compound row as a record instead of a list
  • #368 : Decode a zarr v2 compound that holds an object field
  • #369 : Align ZarrDataIO with the new zarr version 3 codec API
  • #370 : Raise instead of truncating when a compound string field is narrower than the data
  • #372 : Add support for custom sharding properties
  • #373 : Increase Zarr and Python minimum versions
  • #374 : Support automatic shards and guard parallel iterator writes
  • #382 : Fill CHANGELOG gaps and close doc and test gaps from the v3 migration review
  • #393 : Report a Zarr v2 file in every mode, and improve allow_pickle discoverability and the error for the renamed compressor argument
  • #396 : Name the dataset when a value is not valid UTF-8, and correct the dtype spec values and compound data types in the docs
  • #397 : Write a text dataset in one assignment instead of one per element
  • #399 : Remove the extensions argument from NWBZarrIO and NWBZarrV2IO
  • #400 : Remove code for zarr<3.3 and fix issues found along the way
  • #402 : Require zarr>=3.4.0
  • #404 : Prepare for release of HDMF-Zarr 0.14.0