Releases: huggingface/datasets
Release list
5.0.1
Bug fixes
- Fix version string in init.py by @qgallouedec in #8244
- fix conda build by @lhoestq in #8250
- Fix JSON loader schema inference for files starting with a UTF-8 BOM (#8241) by @archievi in #8243
- Support hermes traces by @lhoestq in #8255
- Fix batch(by_column=...) crashing after shard/shuffle/split by @pkooij in #8259
- fix traces streaming by @lhoestq in #8277
- support droid agent traces by @cfahlgren1 in #8263
- Fix CI: commit operation equality (hfh 1.20.0) and pytest parametrize collection error by @Wauplin in #8283
- Fix lance auth by @lhoestq in #8301
- Fix symlink-following arbitrary file write in archive extraction by @AAtomical in #8303
- Bug Fix: Resuming Twice Resets the Dataloader by @francesco-bertolotti in #8295
- Bump fsspec and simpler wds compr by @lhoestq in #8337
- fix: validate Arrow IPC record batches by @XciD in #8350
- Make the dataset fingerprint independent of Arrow chunking by @SuryanshSS1011 in #8339
- Fix casting a nullable LargeList to a different inner type by @vineethsaivs in #8346
- docs: replace AutoFeatureExtractor with AutoImageProcessor in image preprocessing docs by @gautamkishore in #8326
- Fix column drop in Arrow path of axis=1 concatenation by @ebarkhordar in #8342
- Raise on length mismatch in batched IterableDataset.map by @sohumt123 in #8332
- Fix require_storage_embed recursing into require_storage_cast by @vineethsaivs in #8349
- Fix path traversal via metadata file_name in folder-based builders by @Kaif10 in #8325
- Support batched=True in Dataset.to_dict by @vineethsaivs in #8333
- Fix hdf5 external files by @lhoestq in #8355
- remove bad require_storage test by @lhoestq in #8357
- ensure fiels are in repo by @lhoestq in #8356
- Keep flat numeric columns with nulls numeric in numpy format by @ebarkhordar in #8352
- Fix bucket dataset card handling and push metadata accounting by @pjh4993 in #8354
- Fix CSV loader dropping on_bad_lines/encoding_errors on pandas 2.0-2.2 by @ebarkhordar in #8358
- Keep integers on the python read path for fixed-shape ArrayXD columns with nulls by @ebarkhordar in #8363
- Rebatch arrow source before formatting in IterableDataset.filter to fix resume data loss by @ebarkhordar in #8360
- Decode Json() columns in Dataset.to_pandas() by @ebarkhordar in #8344
- fix buckets on windows by @lhoestq in #8369
- Fix DatasetDict.push_to_hub leaving removed splits in the dataset card by @pjh4993 in #8367
- Preserve nullable integer columns in to_json/to_csv/to_sql by @ebarkhordar in #8366
Docs
- docs: fix duplicate "to" in IterableDataset push-to-hub example by @DaoyuanLi2816 in #8252
- Clarify dataset creation vs loading workflows in create_dataset tutorial by @zanvari in #8235
New Contributors
- @DaoyuanLi2816 made their first contribution in #8252
- @zanvari made their first contribution in #8235
- @archievi made their first contribution in #8243
- @pkooij made their first contribution in #8259
- @AAtomical made their first contribution in #8303
- @francesco-bertolotti made their first contribution in #8295
- @XciD made their first contribution in #8350
- @SuryanshSS1011 made their first contribution in #8339
- @vineethsaivs made their first contribution in #8346
- @gautamkishore made their first contribution in #8326
- @ebarkhordar made their first contribution in #8342
- @sohumt123 made their first contribution in #8332
- @Kaif10 made their first contribution in #8325
- @pjh4993 made their first contribution in #8354
Full Changelog: 5.0.0...5.0.1
5.0.0
Datasets Features
Agent traces
-
Parse Agent traces messages for SFT using
teichby @lhoestq in #8232- Agent traces from claude_code/pi/codex and others can now be loaded with load_dataset
- Using the
teichlibrary (new optional dependency), traces are parsed tomessagesto enable training on traces using e.g.trl - Load the data:
>>> from datasets import load_dataset >>> ds = load_dataset("lhoestq/agent-traces-example", split="train") >>> ds[0]["messages"] [{'role': 'user', 'content': 'Download a random dataset from Hugging Face, use DuckDB to inspect it, and come back with a short report about it. Be concise and include: dataset name, what files/format you found, row count or rough size if you can determine it,...' ...]
- Train on agent traces:
trl sft --dataset-name lhoestq/agent-traces-example ...
- find all the Agent traces datasets on HF here: https://huggingface.co/datasets?format=format:agent-traces&sort=trending
Next-level shuffling in streaming mode
-
Use multiple input shards for shuffle buffer by @lhoestq in #8194
ds = load_dataset(..., streaming=True) ds = ds.shuffle(seed=42) # or configure local buffer shuffling manually, default is: ds = ds.shuffle(seed=42, buffer_size=1000, max_buffer_input_shards=10)
toy example comparison
from datasets import IterableDataset ds = IterableDataset.from_dict({"i": range(123_456_789)}, num_shards=1024) ds = ds.shuffle(seed=42) print("Cold start ids:") print(list(ds.take(10)["i"])) print("Nominal regime ids:") print(list(ds.skip(10_000).take(10)["i"]))
before👎:
Cold start ids: [6148853, 6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858] Nominal regime ids: [6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858, 6149290]after✨:
Cold start ids: [7836668, 9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871] Nominal regime ids: [9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871, 16758448]Note:
ds.state_dict()andds.load_state_dict()are still supported for this improved shuffling :) enabling dataset checkpointingNote 2: it uses threads to fetch the first examples in parallel from the input shards
Note 3: This is a BREAKING CHANGE: the default shuffling mechanism now uses multiple input shards. You can get the old mechanism by passing
max_buffer_input_shards=1toIterableDataset.shuffle()
New batching features for robotics datasets
-
Add batch(by_column=...) by @lhoestq in #8172
from datasets import Dataset ds = Dataset.from_dict({"episode": [0] * 10 + [1] * 10, "frame": list(range(10)) * 2}) # ds = ds.to_iterable_dataset() ds = ds.batch(by_column="episode") for x in ds: print(x) # {'episode': [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]} # {'episode': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]}
New supported formats
- Add Apache Iceberg format support by @frankliee in #8148
- feat: add TsFile (Apache IoTDB) packaged builder with per-device wide format by @JackieTien97 in #8160
- feat: add 3D mesh support and MeshFolder builder by @Vinay-Umrethe in #8055
- Add
.conll/.conlludataset format loader (CoNLL-2003 / 2000 / U) by @CrypticCortex in #8219
Other improvements and bug fixes
- Pass library_name/version to HfApi in dataset push and delete paths by @davanstrien in #8161
- Fix storage_options lookup for streaming Lance datasets by @ericjaebeom in #8166
- add agent trace prompt, sent_at, count fields by @cfahlgren1 in #8163
- fix: add
num_procargument toDataset.to_sqlby @EricSaikali in #7791 - Support fsspec 2026.4.0 by @lhoestq in #8175
- Fix Parquet streaming hangs at the end of script by @lhoestq in #8176
ClassLabeldocs: Correct value for unknown labels by @l-uuz in #7645- fix parquet reshard by @lhoestq in #8193
- Fix parquet columns arg by @lhoestq in #8210
- update readme by @lhoestq in #8208
- update single seg repos in ci by @lhoestq in #8213
- Fix single lance file form pylance 7.0 by @lhoestq in #8225
- fix(map): fix progress bar exceeding total when load_from_cache_file=False by @Nitin-Rajasekar in #8170
- fix: embed_external_files=True for mesh support by @Vinay-Umrethe in #8224
- Fix iterable skip over full Arrow blocks by @my17th2 in #8236
- Keep None as a real null in Json() columns instead of the string "null" by @adityasingh2400 in #8231
- Support composed splits in streaming datasets by @lanarkite99 in #8220
New Contributors
- @ericjaebeom made their first contribution in #8166
- @EricSaikali made their first contribution in #7791
- @l-uuz made their first contribution in #7645
- @CrypticCortex made their first contribution in #8219
- @frankliee made their first contribution in #8148
- @Vinay-Umrethe made their first contribution in #8055
- @Nitin-Rajasekar made their first contribution in #8170
- @JackieTien97 made their first contribution in #8160
- @my17th2 made their first contribution in #8236
- @adityasingh2400 made their first contribution in #8231
- @lanarkite99 made their first contribution in #8220
Full Changelog: 4.8.5...5.0.0
4.8.5
Main bug fixes
- fix: decode Json() values before calling DataFrame.to_json() (#8116) by @Brianzhengca in #8122
- Fix: decode JSON type before to_list or to_dict is called by @ItsTania in #8137
- Fix batching for table-formatted datasets by @bluehyena in #8126
- Fix iterable map resume state by @Brianzhengca in #8147
- don't embed remote files in download_and_prepare to parquet by @lhoestq in #8150
Other improvements and bug fixes
- Parse agent traces by @lhoestq in #8113
- 🔒 Pin GitHub Actions to commit SHAs by @paulinebm in #8114
- chore: bump doc-builder SHA for PR upload workflow by @rtrompier in #8134
- Remove print statement in JSON processing by @lhoestq in #8136
- Don't include files list DatasetInfo (and remove old stuff) by @lhoestq in #8128
- update ci uer by @lhoestq in #8139
- fix warning in ci by @lhoestq in #8140
- fix mask in embed_storage for remote files by @lhoestq in #8151
- fix original_files missing in ci json test by @lhoestq in #8152
- Fix null in embed storage by @lhoestq in #8154
- Fix base_path in integration tests by @lhoestq in #8155
New Contributors
- @paulinebm made their first contribution in #8114
- @Brianzhengca made their first contribution in #8122
- @bluehyena made their first contribution in #8126
- @rtrompier made their first contribution in #8134
- @ItsTania made their first contribution in #8137
Full Changelog: 4.8.4...4.8.5
4.8.4
4.8.3
What's Changed
- Fix split_dataset_by_node step by @lhoestq in #8081
- Fix docstring of Json.cast_storage by @albertvillanova in #8080
Full Changelog: 4.8.2...4.8.3
4.8.2
4.8.1
What's Changed
- Fix formatted iter arrow double yield by @HaukurPall in #8063
Full Changelog: 4.8.0...4.8.1
4.8.0
Dataset Features
-
Read (and write) from HF Storage Buckets: load raw data, process and save to Dataset Repos by @lhoestq in #8064
from datasets import load_dataset # load raw data from a Storage Bucket on HF ds = load_dataset("buckets/username/data-bucket", data_files=["*.jsonl"]) # or manually, using hf:// paths ds = load_dataset("json", data_files=["hf://buckets/username/data-bucket/*.jsonl"]) # process, filter ds = ds.map(...).filter(...) # publish the AI-ready dataset ds.push_to_hub("username/my-dataset-ready-for-training")
This also fixes multiprocessed push_to_hub on macos that was causing segfault (now it uses spawn instead of fork).
And it bumpsdillandmultiprocessversions to support python 3.14 -
Datasets streaming iterable packaged improvements and fixes by @Michael-RDev in #8068
- added
max_shard_sizeto IterableDataset.push_to_hub (but requires iterating twice to know the full dataset twice - improvements are welcome) - more arrow-native iterable operations for IterableDataset
- better support of glob patterns in archives, e.g.
zip://*.jsonl::hf://datasets/username/dataset-name/data.zip - fixes for to_pandas, videofolder, load_dataset_builder kwargs
- added
What's Changed
- fix reshard_data_sources by @lhoestq in #8061
- Improve error message for invalid data_files pattern format by @kushalkkb in #8060
- fix null filling in missing jsonl columns by @lhoestq in #8069
New Contributors
- @kushalkkb made their first contribution in #8060
- @Michael-RDev made their first contribution in #8068
Full Changelog: 4.7.0...4.8.0
4.7.0
Datasets Features
- Add
Json()type by @lhoestq in #8027- JSON Lines files that contain arbitrary JSON objects like tool calling datasets are now supported. When there is a field or subfield containing mixed types (e.g. mix of str/int/float/dict/list or dictionaries with arbitrary keys), the
Json()type is used to store such data that would normally not be supported in Arrow/Parquet - Use the
Json()type inFeatures()for any dataset, it is supported in any functions that acceptsfeatures=likeload_dataset(),.map(),.cast(),.from_dict(),.from_list() - Use
on_mixed_types="use_json"to automatically set theJson()type on mixed types in.from_dict(),.from_list()and.map()
- JSON Lines files that contain arbitrary JSON objects like tool calling datasets are now supported. When there is a field or subfield containing mixed types (e.g. mix of str/int/float/dict/list or dictionaries with arbitrary keys), the
Examples:
You can use on_mixed_types="use_json" or specify features= with a [Json] type:
>>> ds = Dataset.from_dict({"a": [0, "foo", {"subfield": "bar"}]})
Traceback (most recent call last):
...
File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: Could not convert 'foo' with type str: tried to convert to int64
>>> features = Features({"a": Json()})
>>> ds = Dataset.from_dict({"a": [0, "foo", {"subfield": "bar"}]}, features=features)
>>> ds.features
{'a': Json()}
>>> list(ds["a"])
[0, "foo", {"subfield": "bar"}]This is also useful for lists of dictionaries with arbitrary keys and values, to avoid filling missing fields with None:
>>> ds = Dataset.from_dict({"a": [[{"b": 0}, {"c": 0}]]})
>>> ds.features
{'a': List({'b': Value('int64'), 'c': Value('int64')})}
>>> list(ds["a"])
[[{'b': 0, 'c': None}, {'b': None, 'c': 0}]] # missing fields are filled with None
>>> features = Features({"a": List(Json())})
>>> ds = Dataset.from_dict({"a": [[{"b": 0}, {"c": 0}]]}, features=features)
>>> ds.features
{'a': List(Json())}
>>> list(ds["a"])
[[{'b': 0}, {'c': 0}]] # OKAnother example with tool calling data and the on_mixed_types="use_json" argument (useful to not have to specify features= manually):
>>> messages = [
... {"role": "user", "content": "Turn on the living room lights and play my electronic music playlist."},
... {"role": "assistant", "tool_calls": [
... {"type": "function", "function": {
... "name": "control_light",
... "arguments": {"room": "living room", "state": "on"}
... }},
... {"type": "function", "function": {
... "name": "play_music",
... "arguments": {"playlist": "electronic"} # mixed-type here since keys ["playlist"] and ["room", "state"] are different
... }}]
... },
... {"role": "tool", "name": "control_light", "content": "The lights in the living room are now on."},
... {"role": "tool", "name": "play_music", "content": "The music is now playing."},
... {"role": "assistant", "content": "Done!"}
... ]
>>> ds = Dataset.from_dict({"messages": [messages]}, on_mixed_types="use_json")
>>> ds.features
{'messages': List({'role': Value('string'), 'content': Value('string'), 'tool_calls': List(Json()), 'name': Value('string')})}
>>> ds[0][1]["tool_calls"][0]["function"]["arguments"]
{"room": "living room", "state": "on"}What's Changed
- Fix typos in iterable_dataset.py by @omkar-334 in #8049
- Fix non-deterministic by sorting metadata extensions (#8034) by @Nexround in #8039
- Use num_examples instead of len(self) for iterable_dataset's SplitInfo by @HaukurPall in #8041
- Fix silent data loss in push_to_hub when num_proc > num_shards by @HaukurPall in #8044
- Don't extract bad files by @lhoestq in #8056
- fix(iterable_dataset): preserve features when chaining filter() on typed IterableDataset by @s-zx in #8053
- fix: handle nested null types in feature alignment for multi-proc map by @ain-soph in #8047
- Fix unstable tokenizer fingerprinting (enables map cache reuse) by @KOKOSde in #7982
- Limit dataset listing to first 20 entries in readme by @lhoestq in #8057
New Contributors
- @omkar-334 made their first contribution in #8049
- @Nexround made their first contribution in #8039
- @HaukurPall made their first contribution in #8041
- @s-zx made their first contribution in #8053
- @ain-soph made their first contribution in #8047
- @KOKOSde made their first contribution in #7982
Full Changelog: 4.6.1...4.7.0

