Repository navigation
2.9.0 Release Contributors (PR Count Order)
Total PRs counted in this release: 459
- Peter Cnudde — 104 PRs
- Yuan-Ting Hsieh (謝沅廷) — 89 PRs
- Chester Chen — 73 PRs
- Holger Roth — 59 PRs
- Zhihong Zhang — 42 PRs
- Isaac Yang — 34 PRs
- nvkevlu — 25 PRs
- Ziyue Xu — 10 PRs
- Yuanyuan Chen — 4 PRs
- nvshaxie — 4 PRs
- Oleksandr_Sanin — 4 PRs
- Minh Vu — 4 PRs
- Zare2001 — 1 PR
- Roshni Bhowmik — 1 PR
- mrityunjay2627 — 1 PR
- Kevin Ta — 1 PR
- Hashim Khan — 1 PR
- Douwe van der Wal — 1 PR
- Chih-Chieh — 1 PR
🎉 Welcome First-Time Contributors!
- Oleksandr_Sanin — 4 PRs
- Minh Vu — 4 PRs
- Roshni Bhowmik — 1 PR
- mrityunjay2627 — 1 PR
- Hashim Khan — 1 PR
- Chih-Chieh — 1 PR
Feature Highlights
NVIDIA FLARE 2.9.0 focuses on agent-assisted federated development, a Python-first research API, a new HPC job launcher, and hardening large-model training and the internal transport for production deployments. It also ships Kubernetes/OpenShift staging support, a Hugging Face Client API, and Client API/Recipe consolidation.
- Agent Skills: two categories of bundled skills for agent-assisted federated development, validated by pre-merge security scans including prompt-injection and untrusted-input eval coverage.
- Conversion skills generate a reviewable federated job from an existing project or dataset — PyTorch, PyTorch Lightning, and Hugging Face Trainer training conversion, plus a federated-statistics skill that builds a
FedStatsRecipejob directly from a dataset. - Auto-FL optimization is an agent-directed campaign that tunes an existing job within its declared training budget — NVFLARE owns the deterministic campaign import, execution, policy boundaries, and provenance, while a coding agent proposes hypothesis-driven, budget- and schema-bounded candidates.
- Conversion skills generate a reviewable federated job from an existing project or dataset — PyTorch, PyTorch Lightning, and Hugging Face Trainer training conversion, plus a federated-statistics skill that builds a
- Collaboration API (Technical Preview): a Python-first way to express custom federated algorithms — decorate the functions a server or client publishes, write coordination logic in ordinary Python, and package/simulate/submit with
CollabRecipe. Every Collab call is authorized against the caller's authenticated CellNet origin before dispatch. - Slurm job launcher: a new HPC execution target alongside process, Docker, and Kubernetes — a long-lived NVFLARE parent submits each client or server job as a Slurm batch job, with Apptainer, Pyxis/Enroot, and bare-Python execution backends, GPU-aware worker setup, and a shared-file worker channel for clusters without direct node connectivity.
- Large-model training hardening: reliable streaming (bounded chunk retries, progress-aware liveness instead of fixed timeouts), better sender/receiver flow-control synchronization, and tensor disk offload extended from FedAvg to Scaffold, FedOpt, and Swarm keep peak aggregator memory flatter as models and client counts grow. FedAvg now works out of the box for federated LLM training, with no configuration tuning required — validated up to a 72-billion-parameter model, with larger models likely to work as well though not yet tested.
- Security hardening: CellNet messages move to signed AES-256-GCM envelopes with verified sender signatures, internal CellNet TCP links default to mutual TLS across Docker, Slurm, Kubernetes, and Network Attach deployments, admin sessions fail closed on unverifiable tokens, cross-client authentication routes through the server trust boundary, and
require_signed_jobsis now also enforced client-side. - Kubernetes and OpenShift deployment:
nvflare deploy k8s stage/unstagestage a prepared kit as Kubernetes ConfigMaps and Secrets through a generated Helm chart, with OpenShift (--kubectl oc) and multicloud examples. - Hugging Face Client API: federate an existing
Traineror TRLSFTTrainerthroughflare.patch(trainer), with FLARE owning round exchange, global-weight loading, local-budget enforcement, rank-0 communication, checkpoint continuity, and metric reporting. - Unified Client API execution paths:
ClientAPIExecutorconsolidates trainer-process ownership behindin_process,external_process, and a newattachmode for independently managed trainers, replacing the previousInProcessClientAPIExecutor/ClientAPILauncherExecutorstacks. - Deprecation: the deprecated FL HUB feature and its runtime, documentation, and test surface are removed.
See the full 2.9.0 release note: https://nvflare.readthedocs.io/en/2.9.0/release_notes/flare_290.html
What's Changed
- Added Agent Skills for PyTorch/Lightning/Hugging Face Trainer conversion and federated statistics, with lint, eval, packaging, and security hardening, by @chesterxgchen in #4837, #4967, #5114, #5139, #5170, #5176, #5181, #5213, #5225, #5231
- Added the Auto-FL Agent Skill's optimization campaign, including native metric-direction support, budget/schema admission guards, and hardened CI contracts, by @holgerroth in #5169, #5182, #5186, #5187, #5188, #5200, #5211, #5229
- Added the Collaboration API (technical preview) with CellNet-authenticated calls, POC/simulation fixes, and Hello Collab / async CIFAR-10 / SplitNN / federated LLM SFT examples by @nvidianz, @ZiyueXu77, and @holgerroth in #4960, #4975, #4978, #4988, #5003, #5005, #5034, #5036, #5050, #5079, #5183, #5185, #5223, #5238
- Added the Slurm job launcher with Apptainer/Pyxis/Enroot backends, GPU-aware worker setup, shared-file worker channel, and internal CellNet mTLS by @pcnudde in #4986, #4996, #4997, #5019, #5037, #5038, #5049, #5071, #5094, #5104, #5144, #5236, #5237
- Hardened large-model streaming reliability and throughput (progress-aware liveness, flow-control synchronization, blob size limits) and extended tensor disk offload to Scaffold, FedOpt, and Swarm by @nvidianz, @pcnudde, and @YuanTingHsieh in #4736, #4768, #4794, #4799, #4844, #4861, #4965, #4995, #5015, #5046, #5068, #5103, #5165, #5190
- Hardened CellNet message authentication (AES-256-GCM), enabled internal mTLS by default for Docker/Kubernetes/Slurm, and routed cross-client authentication through the server by @IsaacYangSLA, @pcnudde, and @YuanTingHsieh in #4879, #4932, #4933, #4956, #5079, #5096, #5107, #5131, #5142, #5144, #5147, #5158, #5159
- Added Kubernetes ConfigMap/Secret staging (
nvflare deploy k8s stage/unstage), OpenShift support, and multicloud examples by @IsaacYangSLA and @pcnudde in #4679, #4760, #4770, #4782, #4804, #4871, #4877, #4911, #5047, #5076, #5077, #5078 - Added the Hugging Face Client API and
flare.patch(trainer)support, plus theexternal_processexecution mode and Attach mode for externally owned trainers by @chesterxgchen and @YuanTingHsieh in #4906, #4948, #4967, #4994, #5027, #5113, #5152, #5179, #5213 - Unified Client API execution paths behind
ClientAPIExecutor(in_process/external_process/attach) and deprecated legacy training APIs by @holgerroth in #4641 - Removed deprecated FL HUB support and migrated tutorials off the deprecated job CLI by @YuanTingHsieh and @nvkevlu in #4971, #5156
- Added automatic FedProx and SCAFFOLD support for PyTorch Lightning clients and fixed Lightning best-model metric reporting and DDP result sending by @holgerroth and @chesterxgchen in #4838, #4895, #5040, #5145
- Documented Slurm/Kubernetes dual-EKU re-provisioning, Slurm shared-file worker transport, and Kubernetes restaging namespace behavior by @nvidianz and @IsaacYangSLA in #5214, #5222, #5237
Full Changelog: 2.8.0...2.9.0