Skip to content

Releases: huggingface/datasets

5.0.1

Choose a tag to compare

@lhoestq lhoestq released this 28 Jul 11:11
921c2a7

Bug fixes

Docs

  • docs: fix duplicate "to" in IterableDataset push-to-hub example by @DaoyuanLi2816 in #8252
  • Clarify dataset creation vs loading workflows in create_dataset tutorial by @zanvari in #8235

New Contributors

Full Changelog: 5.0.0...5.0.1

5.0.0

Choose a tag to compare

@lhoestq lhoestq released this 05 Jun 13:29
68ac1a9

Datasets Features

Agent traces

  • Parse Agent traces messages for SFT using teich by @lhoestq in #8232

    • Agent traces from claude_code/pi/codex and others can now be loaded with load_dataset
    • Using the teich library (new optional dependency), traces are parsed to messages to enable training on traces using e.g. trl
    • Load the data:
    >>> from datasets import load_dataset
    >>> ds = load_dataset("lhoestq/agent-traces-example", split="train")
    >>> ds[0]["messages"]
    [{'role': 'user', 'content': 'Download a random dataset from Hugging Face, use DuckDB to inspect it, and come back with a short report about it. Be concise and include: dataset name, what files/format you found, row count or rough size if you can determine it,...'
     ...]
    • Train on agent traces:
    trl sft --dataset-name lhoestq/agent-traces-example ...

Next-level shuffling in streaming mode

  • Use multiple input shards for shuffle buffer by @lhoestq in #8194

    ds = load_dataset(..., streaming=True)
    ds = ds.shuffle(seed=42)
    # or configure local buffer shuffling manually, default is:
    ds = ds.shuffle(seed=42, buffer_size=1000, max_buffer_input_shards=10)

    before👎:
    image

    after✨:
    image

    toy example comparison

    from datasets import IterableDataset
    
    ds = IterableDataset.from_dict({"i": range(123_456_789)}, num_shards=1024)
    ds = ds.shuffle(seed=42)
    
    print("Cold start ids:")
    print(list(ds.take(10)["i"]))
    print("Nominal regime ids:")
    print(list(ds.skip(10_000).take(10)["i"]))

    before👎:

    Cold start ids:
    [6148853, 6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858]
    Nominal regime ids:
    [6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858, 6149290]
    

    after✨:

    Cold start ids:
    [7836668, 9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871]
    Nominal regime ids:
    [9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871, 16758448]
    

    Note: ds.state_dict() and ds.load_state_dict() are still supported for this improved shuffling :) enabling dataset checkpointing

    Note 2: it uses threads to fetch the first examples in parallel from the input shards

    Note 3: This is a BREAKING CHANGE: the default shuffling mechanism now uses multiple input shards. You can get the old mechanism by passing max_buffer_input_shards=1 to IterableDataset.shuffle()

New batching features for robotics datasets

  • Add batch(by_column=...) by @lhoestq in #8172

    from datasets import Dataset
    
    ds = Dataset.from_dict({"episode": [0] * 10 + [1] * 10, "frame": list(range(10)) * 2})
    # ds = ds.to_iterable_dataset()
    ds = ds.batch(by_column="episode")
    for x in ds:
        print(x)
    # {'episode': [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]}
    # {'episode': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]}

New supported formats

Other improvements and bug fixes

New Contributors

Full Changelog: 4.8.5...5.0.0

4.8.5

Choose a tag to compare

@lhoestq lhoestq released this 27 Apr 15:46
a015b2f

Main bug fixes

Other improvements and bug fixes

New Contributors

Full Changelog: 4.8.4...4.8.5

4.8.4

Choose a tag to compare

@lhoestq lhoestq released this 23 Mar 14:21
a0ce369

What's Changed

  • Support latest torchvision by @lhoestq in #8087
  • fix regression when loading JSON with one file = one object by @lhoestq in #8086

Full Changelog: 4.8.3...4.8.4

4.8.3

Choose a tag to compare

@lhoestq lhoestq released this 19 Mar 17:44
d4942e2

What's Changed

Full Changelog: 4.8.2...4.8.3

4.8.2

Choose a tag to compare

@lhoestq lhoestq released this 17 Mar 01:10
2c46ab1

What's Changed

Full Changelog: 4.8.1...4.8.2

4.8.1

Choose a tag to compare

@lhoestq lhoestq released this 17 Mar 00:13
9352382

What's Changed

Full Changelog: 4.8.0...4.8.1

4.8.0

Choose a tag to compare

@lhoestq lhoestq released this 16 Mar 23:52
c988a5d

Dataset Features

  • Read (and write) from HF Storage Buckets: load raw data, process and save to Dataset Repos by @lhoestq in #8064

    from datasets import load_dataset
    # load raw data from a Storage Bucket on HF
    ds = load_dataset("buckets/username/data-bucket", data_files=["*.jsonl"])
    # or manually, using hf:// paths
    ds = load_dataset("json", data_files=["hf://buckets/username/data-bucket/*.jsonl"])
    # process, filter
    ds = ds.map(...).filter(...)
    # publish the AI-ready dataset
    ds.push_to_hub("username/my-dataset-ready-for-training")

    This also fixes multiprocessed push_to_hub on macos that was causing segfault (now it uses spawn instead of fork).
    And it bumps dill and multiprocess versions to support python 3.14

  • Datasets streaming iterable packaged improvements and fixes by @Michael-RDev in #8068

    • added max_shard_size to IterableDataset.push_to_hub (but requires iterating twice to know the full dataset twice - improvements are welcome)
    • more arrow-native iterable operations for IterableDataset
    • better support of glob patterns in archives, e.g. zip://*.jsonl::hf://datasets/username/dataset-name/data.zip
    • fixes for to_pandas, videofolder, load_dataset_builder kwargs

What's Changed

New Contributors

Full Changelog: 4.7.0...4.8.0

4.7.0

Choose a tag to compare

@lhoestq lhoestq released this 09 Mar 19:09
ac9c452

Datasets Features

  • Add Json() type by @lhoestq in #8027
    • JSON Lines files that contain arbitrary JSON objects like tool calling datasets are now supported. When there is a field or subfield containing mixed types (e.g. mix of str/int/float/dict/list or dictionaries with arbitrary keys), the Json()type is used to store such data that would normally not be supported in Arrow/Parquet
    • Use the Json() type in Features() for any dataset, it is supported in any functions that accepts features=like load_dataset(), .map(), .cast(), .from_dict(), .from_list()
    • Use on_mixed_types="use_json" to automatically set the Json() type on mixed types in .from_dict(), .from_list() and .map()

Examples:

You can use on_mixed_types="use_json" or specify features= with a [Json] type:

>>> ds = Dataset.from_dict({"a": [0, "foo", {"subfield": "bar"}]})
Traceback (most recent call last):
  ...
  File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: Could not convert 'foo' with type str: tried to convert to int64

>>> features = Features({"a": Json()})
>>> ds = Dataset.from_dict({"a": [0, "foo", {"subfield": "bar"}]}, features=features)
>>> ds.features
{'a': Json()}
>>> list(ds["a"])
[0, "foo", {"subfield": "bar"}]

This is also useful for lists of dictionaries with arbitrary keys and values, to avoid filling missing fields with None:

>>> ds = Dataset.from_dict({"a": [[{"b": 0}, {"c": 0}]]})
>>> ds.features
{'a': List({'b': Value('int64'), 'c': Value('int64')})}
>>> list(ds["a"])
[[{'b': 0, 'c': None}, {'b': None, 'c': 0}]]  # missing fields are filled with None

>>> features = Features({"a": List(Json())})
>>> ds = Dataset.from_dict({"a": [[{"b": 0}, {"c": 0}]]}, features=features)
>>> ds.features
{'a': List(Json())}
>>> list(ds["a"])
[[{'b': 0}, {'c': 0}]]  # OK

Another example with tool calling data and the on_mixed_types="use_json" argument (useful to not have to specify features= manually):

>>> messages = [
...     {"role": "user", "content": "Turn on the living room lights and play my electronic music playlist."},
...     {"role": "assistant", "tool_calls": [
...         {"type": "function", "function": {
...             "name": "control_light",
...             "arguments": {"room": "living room", "state": "on"}
...         }},
...         {"type": "function", "function": {
...             "name": "play_music",
...             "arguments": {"playlist": "electronic"}  # mixed-type here since keys ["playlist"] and ["room", "state"] are different
...         }}]
...     },
...     {"role": "tool", "name": "control_light", "content": "The lights in the living room are now on."},
...     {"role": "tool", "name": "play_music", "content": "The music is now playing."},
...     {"role": "assistant", "content": "Done!"}
... ]
>>> ds = Dataset.from_dict({"messages": [messages]}, on_mixed_types="use_json")
>>> ds.features
{'messages': List({'role': Value('string'), 'content': Value('string'), 'tool_calls': List(Json()), 'name': Value('string')})}
>>> ds[0][1]["tool_calls"][0]["function"]["arguments"]
{"room": "living room", "state": "on"}

What's Changed

  • Fix typos in iterable_dataset.py by @omkar-334 in #8049
  • Fix non-deterministic by sorting metadata extensions (#8034) by @Nexround in #8039
  • Use num_examples instead of len(self) for iterable_dataset's SplitInfo by @HaukurPall in #8041
  • Fix silent data loss in push_to_hub when num_proc > num_shards by @HaukurPall in #8044
  • Don't extract bad files by @lhoestq in #8056
  • fix(iterable_dataset): preserve features when chaining filter() on typed IterableDataset by @s-zx in #8053
  • fix: handle nested null types in feature alignment for multi-proc map by @ain-soph in #8047
  • Fix unstable tokenizer fingerprinting (enables map cache reuse) by @KOKOSde in #7982
  • Limit dataset listing to first 20 entries in readme by @lhoestq in #8057

New Contributors

Full Changelog: 4.6.1...4.7.0

4.6.1

Choose a tag to compare

@lhoestq lhoestq released this 27 Feb 23:27
7afef69

Bug fix

Full Changelog: 4.6.0...4.6.1