Repository navigation
This release marks a major transition for hdmf-zarr, updating the core storage format from Zarr v2 to the new Zarr v3 standard. As part of this transition, hdmf-zarr has adopted the unified Zarr v3 storage convention developed in collaboration with Zindi and LINDI.
Migration Notes (Zarr V2 to V3)
- Zarr V3 Standard: The primary
ZarrIOandNWBZarrIOclasses now write and read Zarr v3 format exclusively, usingzarr-pythonv3. - Legacy Zarr V2 Support: Legacy Zarr v2 files cannot be read by
ZarrIOdirectly. Instead, use the newly addedZarrV2IOandNWBZarrV2IObackends to read them. Attempting to open a v2 file withZarrIOwill now raise a clear error directing users to the v2 classes. - Migration Path: To migrate existing files to the new format, you can use
NWBZarrV2IO.export_to_v3(...)or the one-shotNWBZarrV2IO.convert_to_v3(...)static helper to efficiently convert NWB Zarr v2 files to Zarr v3. - Forward Compatibility: Files written by this release are Zarr v3 and cannot be opened by
hdmf-zarr0.13 or earlier, which requirezarr<3. Upgrade tohdmf-zarr0.14 in order to read Zarr v3 files written withhdmf-zarr.
Changes
The core functionality has been overhauled to transition to Zarr V3, including structural, encoding, and reading changes.
- Updated
ZarrIOto use Zarr V3: The primaryZarrIOandNWBZarrIOclasses now write and read Zarr v3 format exclusively, usingzarr-pythonv3. (see list of contributors and related pull requests below) - Adopted New Storage Convention: Adopted the unified Zarr v3 storage convention shared with Zindi/LINDI. Links are stored in a
_LINKSattribute on the parent group, dataset types in_DTYPEand_REFERENCE_FIELDS. Scalar datasets, including scalar object references and scalar compound datasets, are stored as zero-dimensional arrays with shape(), matchingHDF5IO, and are read back as scalars. Attribute references are formatted as{"_REFERENCE": {"source", "path"}}, while references in datasets are plain target path strings.ZarrReferenceremoves unusedobject_idandsource_object_idkeys. A Zarr v3 store generated by Zindi from an HDF5 NWB file can now be read directly withNWBZarrIO. @bendichter #336 #325 - Aligned
ZarrDataIOwith the Zarr v3 codec API: Thecompressorargument is renamed tocompressors, theserializer(ArrayBytesCodec) slot is now directly accessible, andfiltersonly applies toArrayArrayCodec. @h-mayorquin #369 - Support dataset sharding via
ZarrDataIO: Added newshardsargument toZarrDataIOto define custom sharding properties, including support for automatic sharding and guarded parallel iterator writes. @h-mayorquin @alejoe91 #374 #372 - Requirements Changes: Bumped minimum required Python version to 3.12 and Zarr-Python dependency to
>=3.4.0due to backwards-incompatible changes in Zarr v3 structured data types and bug fixes in Zarr-Python 3.4.0. Also bumpedhdmfto>=6.2.0,numpyto>=2.2.3(earlier versions write bytes in string datasets as their Python representation, e.g.,"b'ok'", instead of decoding them), andpynwbto>=4.2.0, and removed thenumcodecs<0.16.0upper bound. @oruebel @rly #373 #382 #402 #404 - Removed the
synchronizerargument:ZarrIOandNWBZarrIOno longer accept asynchronizerargument, and theZarrIO.synchronizerproperty is removed. zarr-python v3 does not provide the synchronizer mechanism this wrapped. Passingsynchronizer=...raisesTypeError. - Removed the
object_codec_classargument:ZarrIOandNWBZarrIOno longer accept anobject_codec_classargument, and theZarrIO.object_codec_classproperty is removed. References, cached specs, and compound datasets are serialized as JSON inStringDTypearrays, so there is no object codec to select. Passingobject_codec_class=...raisesTypeError. - Removed the
extensionsargument (breaking):NWBZarrIOandNWBZarrV2IOno longer accept anextensionsargument, matching the removal of the same argument fromNWBHDF5IOin PyNWB 4.0.0. Load the cached namespaces from the file (the default,load_namespaces=True) or pass a prebuiltmanagerinstead. Passingextensions=...raisesTypeError. @rly #399 - Supported store classes:
SUPPORTED_ZARR_STOREScoversLocalStore,FsspecStore, and any other zarr v3Storesubclass. Reading a remote store requires fsspec, whichpip install hdmf-zarr[full]installs.DirectoryStore,NestedDirectoryStore, andTempStoredo not exist in zarr-python v3, so code that constructs one needs to useLocalStoreinstead. - Remote write: Passing
storage_optionsis supported only withmode="r". Writing an NWB Zarr file directly to a remote store is not currently supported and raisesValueError.
Added
Added support for reading and converting existing data using Zarr V2:
- Zarr V2 Legacy Read Support: Added
ZarrV2IOandNWBZarrV2IOread-only backends for interacting with legacy Zarr v2 files. Opening a Zarr v2 file withZarrIOdirectly will now raise a clear error message directing users to the appropriate legacy reader. @alejoe91 #349 - Migration Helpers: Added helpers
NWBZarrV2IO.export_to_v3andNWBZarrV2IO.convert_to_v3to conveniently convert legacy Zarr v2 files to Zarr v3. @alejoe91 #349
Fixes
Addressed the following bugs that are independent of the migration to Zarr V3:
- Fixed bug where
ZarrIO.generate_dataset_htmlwould raise an error when called with a non-Zarr object. @oruebel #355 - Fixed bug where a compound dtype field declared with the spec type
uint(orshort) was written asfloat64. @ehennestad #365 - Fixed bug where writing a scalar dataset with a compound dtype, such as
ElectrodeGroup.position, raisedIndexError. @rly #277 - Reading or writing a scalar dataset with a compound dtype that has a reference field raises
NotImplementedError. This combination is not supported. @rly #391 - Fixed the S3 streaming tutorial catching every exception from its read, which let the gallery tests pass without reading the file. It now skips the read only when fsspec is not installed or the network is unavailable. @rly #400
Contributors
This release was made possible by the efforts of @bendichter, @alejoe91, @h-mayorquin, @ehennestad, @rly, and @oruebel.
Related Pull Requests
- #325 : Migrate hdmf-zarr from zarr-python v2 to v3
- #336 : Adopt unified Zarr v3 convention for hdmf-zarr/Zindi interop
- #363 : Fix Zarr-to-HDF5 export deadlock
- #365 : Resolve compound fields declared
uintinstead of writing them as float64 - #366 : Review and propose changes to v3 integration
- #367 : Rebuild a resolved compound row as a record instead of a list
- #368 : Decode a zarr v2 compound that holds an object field
- #369 : Align
ZarrDataIOwith the new zarr version 3 codec API - #370 : Raise instead of truncating when a compound string field is narrower than the data
- #372 : Add support for custom sharding properties
- #373 : Increase Zarr and Python minimum versions
- #374 : Support automatic shards and guard parallel iterator writes
- #382 : Fill CHANGELOG gaps and close doc and test gaps from the v3 migration review
- #393 : Report a Zarr v2 file in every mode, and improve
allow_picklediscoverability and the error for the renamedcompressorargument - #396 : Name the dataset when a value is not valid UTF-8, and correct the
dtypespec values and compound data types in the docs - #397 : Write a text dataset in one assignment instead of one per element
- #399 : Remove the
extensionsargument fromNWBZarrIOandNWBZarrV2IO - #400 : Remove code for zarr<3.3 and fix issues found along the way
- #402 : Require zarr>=3.4.0
- #404 : Prepare for release of HDMF-Zarr 0.14.0