Skip to content

v1.0.0

Latest

Choose a tag to compare

@github-actions github-actions released this 05 Oct 23:47
· 14 commits to main since this release
Immutable release. Only release title and notes can be modified.
v1.0.0
82bccef

AICR's first stable release puts the v1 stability contract into effect, completes the apiVersion migration, extends component upgrade safety to live clusters and generated bundles, and adds validation that proves GPU access and fabric capability rather than assuming it.

Highlights

Community - AICR launched at KubeCon EU in March this year. Today, little over 6 months later, the AICR project reached its first stable release. In that short time over 85 people from more than 20 organizations have contributed, filing >1K issues and submitting >1.5K PRs. Every bug report, question, review, and fix helped decide what v1.0.0 commits to. Thank you!

v1 Stability Contract - The aicr CLI, the REST API, the Go SDK (pkg/client/v1), and the bundle layout and artifact schemas are now frozen. A breaking change on any of them waits for the next major release and is announced first under the deprecation policy.

Component Upgrade Safety - aicr upgrade-check --from cluster compares against what is actually installed, and warns about live objects an upgrade could destroy (NVIDIA/aicr#2840). Upgrade checks now also catch a change to a component's chart, source, path, type, manifests, or object names, and a bundle carries the upgrade steps for its own deployer whenever a pinned version needs them (NVIDIA/aicr#2986).

Validation Hardening

  • slinky-slurm-gpu-access proves that a Slurm job given one GPU can use exactly that GPU, and that a job without one cannot open any (NVIDIA/aicr#3057)
  • GKE checks now verify that the GPUDirect-TCPXO fabric is usable (every a3 node exposes all 8 GPU NICs), not just that its Networks exist, and that bundle-installer pools have driver auto-install turned off
  • Wider conformance coverage: robust-controller and secure-accelerator-access on the OKE, LKE, and RTX PRO 6000 bases, IMEX for recipe-supplied NCCL runtimes, and a fail-closed OKE device-plugin check

Bundler and Deployer Reliability - The Helm readiness gate now blocks the deploy and re-runs on upgrade, and its Jobs schedule on system nodes instead of hanging Pending on tainted clusters. Bundle scripts take the cluster connection from KUBE_CONTEXT and KUBECONFIG. A missing component dependency is reported as missing, not as a false cycle, before any deployer writes output.

Recipes and Coverage

  • GB200 GKE Kubeflow gains torch-distributed-rdma, a GPUDirect-RDMA training runtime a TrainJob selects by name (NVIDIA/aicr#2859)
  • GPU Operator moves to v26.7.1, dropping the ComputeDomain CRD workaround, with the OpenShift OLM channel on v26.7
  • NVSentinel moves to v1.25.0, with the metadata-collector back on for VR200 and enabled on the k0s H200 leaf; Grove moves to v0.1.0-alpha.13
  • k8s-aibom moves to v1.5.1, and aicr recipe --runtime-inventory enabled can now add it to any stock GKE recipe
  • VR200 recipes ship one rebooting NodeWright CR instead of two, replacing the standalone rdma-netns-exclusive component
  • EKS and AKS gain service-level Ubuntu overlays, matching GKE
  • A --data overlay can now add values to an embedded profile without forking the overlay that declares it (NVIDIA/aicr#2985)

Thanks to @ArangoGutierrez, @atif1996, @ayuskauskas, @EronWright, @ezhang3333, @framsouza, @lalitadithya, @lockwobr, @mikecook, @mohityadav8, @njhensley, @rorajani, @SatyamPandey-07, @srao-nv, @varmesh, @yuanchen8911, and @mchmarny.

Changelog

New Features

Bug Fixes

Other Tasks