Skip to content

[v2.2.0] `hf jobs stats` snapshots, download and `HfFileSystem` fixes

Latest

Choose a tag to compare

@github-actions github-actions released this 08 Oct 14:54
· 1 commit to main since this release

A small release this week, mostly bug fixes and quality-of-life improvements. hf jobs stats now returns right away instead of streaming forever, concurrent downloads reliably use hf_xet, and HfFileSystem caching is more predictable. Two behavior changes are listed under Breaking changes.

📊 hf jobs stats prints a snapshot by default

hf jobs stats used to stream metrics until the job finished, so scripts and agents calling it on a running job hung until their own timeout. It now prints the latest sample of each job once and exits, like hf jobs logs and hf spaces logs. Use -f/--follow for the previous live view. Snapshot output also supports --format json|agent|quiet, which outputs the raw API metrics.

On the Python side, HfApi.fetch_job_metrics gets the same follow parameter as fetch_job_logs. It defaults to False and returns only the current sample. Pass follow=True to stream until the job completes.

>>> hf jobs stats 6ac3bd9afbc85ba6823a9ad4
JOB ID                CPU % NUM CPU MEM % MEM USAGE      NET I/O         GPU UTIL % GPU MEM % GPU MEM USAGE
--------------------- ----- ------- ----- -------------- --------------- ---------- --------- -------------
6ac3bd9afbc85ba682... 0%    2.0     0.01% 1.0MB / 16.0GB 0.0bps / 0.0bps N/A        N/A       N/A
Hint: Use `hf jobs stats -f celinah/6ac3bd9afbc85ba6823a9ad4` to follow live stats.

# Stream live stats until the job completes
>>> hf jobs stats -f 6ac3bd9afbc85ba6823a9ad4
>>> from huggingface_hub import fetch_job_metrics
>>> for metrics in fetch_job_metrics(job_id=job.id, follow=True):
...     print(metrics)

Other Jobs fixes in this release: hf jobs logs and hf spaces logs without -f no longer hang while the job keeps printing, hf jobs stats no longer fails when one of the requested jobs has finished, and hf jobs stats --json prints [] when no job is running.

📚 Documentation: Jobs guide

⚡ Concurrent downloads always use hf_xet

In a fresh process, the first concurrent hf_hub_download calls could read a placeholder version for hf_xet while another thread was still looking it up. Those calls treated hf_xet as missing and downloaded through the plain HTTP path. This mostly affected parallel downloads such as snapshot_download, where some of the files skipped Xet. The version lookup is now cached only once it is resolved, so every download sees the real value.

  • Fix race in _get_version that makes concurrent downloads skip hf_xet by @BramVanroy in #5102

📂 HfFileSystem cache fixes

Three fixes to how HfFileSystem caches listings and missing paths:

  • ls() on a newly created empty bucket returns an empty list instead of raising FileNotFoundError.
  • Directory listings are now replaced in the cache instead of appended to. Before, ls() after find() or ls(refresh=True) could return duplicate entries, refresh=True kept deleted files, and listing a single bucket file cached it as the whole folder's listing.
  • exists(path, refresh=True) and invalidate_cache(path) now clear cached "not found" results. Before, a repo, branch or bucket created in the same process stayed missing until a full invalidate_cache().
>>> fs.exists(repo_id)
False
>>> api.create_repo(repo_id)
>>> fs.exists(repo_id, refresh=True)  # was False before this release
True

💔 Breaking changes

  • hf jobs stats and HfApi.fetch_job_metrics return a single snapshot by default instead of streaming until the job completes. Use -f/--follow or follow=True to keep the previous behavior. See the section above (#5090).

  • load_torch_model now raises FileNotFoundError instead of ValueError when the checkpoint path does not exist, and model_index_to_eval_results raises ValueError on an empty list (#5063).

  • [Bugfix] Fix wrong exception type and unbound variable guard by @vector15-05 in #5063

🖥️ CLI

  • [CLI] Don't glob-expand arguments on Windows by @hanouticelina in #5117. On Windows, quoted patterns such as --exclude ".venv/**" or --include "*.json" were expanded into local file names before reaching the command. hf download <repo> --include "*.json" in a folder containing a .json file silently downloaded nothing.
  • [CLI] Read --env-file / --secrets-file as UTF-8 and strip the BOM by @hanouticelina in #5066
  • [CLI] Upload existing paths with glob characters as-is in hf upload by @Ganesh2325 in #5059

🔧 Other QoL improvements

  • [Download] Warn when allow/ignore patterns match no files in snapshot_download by @Ad1th in #5067
  • [Cache] Speed up delete_revisions with a set lookup by @hanouticelina in #5118. Planning hf cache prune / hf cache rm no longer scales quadratically with the number of files per revision.
  • [Setup] Install hf-xet on Windows ARM64 by @angt in #5074

🐛 Bug and typo fixes

📖 Documentation

  • [Docs] Fix HTTP migration and async session references by @Yuilona in #5079
  • [Docs] Fix from_pretrained integration example by @Ayushdevo in #5105

🏗️ Internal

  • Post-release: bump version to 2.2.0.dev0 by @huggingface-hub-bot[bot] in #5054
  • [CI] Split tests into api / transfer / inference suites by @Wauplin in #5042
  • [Tests] Always clean up staging repos and buckets created by tests by @Wauplin in #5057
  • [CI] Reduce flaky tests (jedi cache, read-after-write, fewer reruns) by @Wauplin in #5058
  • Bump doc-builder workflow pin by @paulinebm in #5087
  • Bump doc-builder main docs workflow pin by @paulinebm in #5093
  • [Tests] Wait for bucket listings after writes by @hanouticelina in #5095