Skip to content

2.9.0: Latest release

Latest

Choose a tag to compare

@IsaacYangSLA IsaacYangSLA released this 04 Sep 20:53
· 67 commits to main since this release
16dbf27

2.9.0 Release Contributors (PR Count Order)

Total PRs counted in this release: 459


🎉 Welcome First-Time Contributors!


Feature Highlights

NVIDIA FLARE 2.9.0 focuses on agent-assisted federated development, a Python-first research API, a new HPC job launcher, and hardening large-model training and the internal transport for production deployments. It also ships Kubernetes/OpenShift staging support, a Hugging Face Client API, and Client API/Recipe consolidation.

  • Agent Skills: two categories of bundled skills for agent-assisted federated development, validated by pre-merge security scans including prompt-injection and untrusted-input eval coverage.
    • Conversion skills generate a reviewable federated job from an existing project or dataset — PyTorch, PyTorch Lightning, and Hugging Face Trainer training conversion, plus a federated-statistics skill that builds a FedStatsRecipe job directly from a dataset.
    • Auto-FL optimization is an agent-directed campaign that tunes an existing job within its declared training budget — NVFLARE owns the deterministic campaign import, execution, policy boundaries, and provenance, while a coding agent proposes hypothesis-driven, budget- and schema-bounded candidates.
  • Collaboration API (Technical Preview): a Python-first way to express custom federated algorithms — decorate the functions a server or client publishes, write coordination logic in ordinary Python, and package/simulate/submit with CollabRecipe. Every Collab call is authorized against the caller's authenticated CellNet origin before dispatch.
  • Slurm job launcher: a new HPC execution target alongside process, Docker, and Kubernetes — a long-lived NVFLARE parent submits each client or server job as a Slurm batch job, with Apptainer, Pyxis/Enroot, and bare-Python execution backends, GPU-aware worker setup, and a shared-file worker channel for clusters without direct node connectivity.
  • Large-model training hardening: reliable streaming (bounded chunk retries, progress-aware liveness instead of fixed timeouts), better sender/receiver flow-control synchronization, and tensor disk offload extended from FedAvg to Scaffold, FedOpt, and Swarm keep peak aggregator memory flatter as models and client counts grow. FedAvg now works out of the box for federated LLM training, with no configuration tuning required — validated up to a 72-billion-parameter model, with larger models likely to work as well though not yet tested.
  • Security hardening: CellNet messages move to signed AES-256-GCM envelopes with verified sender signatures, internal CellNet TCP links default to mutual TLS across Docker, Slurm, Kubernetes, and Network Attach deployments, admin sessions fail closed on unverifiable tokens, cross-client authentication routes through the server trust boundary, and require_signed_jobs is now also enforced client-side.
  • Kubernetes and OpenShift deployment: nvflare deploy k8s stage/unstage stage a prepared kit as Kubernetes ConfigMaps and Secrets through a generated Helm chart, with OpenShift (--kubectl oc) and multicloud examples.
  • Hugging Face Client API: federate an existing Trainer or TRL SFTTrainer through flare.patch(trainer), with FLARE owning round exchange, global-weight loading, local-budget enforcement, rank-0 communication, checkpoint continuity, and metric reporting.
  • Unified Client API execution paths: ClientAPIExecutor consolidates trainer-process ownership behind in_process, external_process, and a new attach mode for independently managed trainers, replacing the previous InProcessClientAPIExecutor / ClientAPILauncherExecutor stacks.
  • Deprecation: the deprecated FL HUB feature and its runtime, documentation, and test surface are removed.

See the full 2.9.0 release note: https://nvflare.readthedocs.io/en/2.9.0/release_notes/flare_290.html


What's Changed

Full Changelog: 2.8.0...2.9.0