v0.2.5
Summary
Adds cache_scope to @huggingface_hub to control where Hugging Face caching mechanism, enabling sharing across runs, users, and flows while reducing redundant downloads and avoiding throttling. The cache key scheme for non-checkpoint scopes now considers revision, allow_patterns, and ignore_patterns to prevent collisions. See commit 467b430.
What’s new
- New decorator option to
@huggingface_hub:cache_scope- Values:
checkpoint(default),flow,global - Controls the base path used for HF caches written via
@huggingface_hub flowwill cache all repos under the same flow regardless of namespaceglobalwill cache all repos under a fixed path regardless of flow/namespace.- For
checkpoint, cache key format is unchanged for backward compatibility - For
flowandglobal, cache keys also include material parameters:revision,allow_patterns,ignore_patterns
- Values:
Usage
from metaflow import huggingface_hub
# Default (per-step scope): unchanged behavior
@huggingface_hub
@step
def default_scope(self):
from metaflow import current
ref = current.huggingface_hub.snapshot_download(
repo_id="mistralai/Mistral-7B-Instruct-v0.1",
allow_patterns=["*.safetensors", "*.json"],
)
# Share cache within a flow even across namespaces (even @project/branch)
@huggingface_hub(cache_scope="flow")
@step
def flow_scope(self):
from metaflow import current
ref = current.huggingface_hub.snapshot_download(
repo_id="mistralai/Mistral-7B-Instruct-v0.1",
allow_patterns=["*.safetensors", "*.json"],
)
# Model cached accessed across flows (Accessible to any/all flow executions within metaflow environment)
@huggingface_hub(load=["mistralai/Mistral-7B-Instruct-v0.1"], cache_scope="global")
@step
def load_into_path(self):
from metaflow import current
model_path = current.huggingface_hub.loaded["mistralai/Mistral-7B-Instruct-v0.1"]