Repository navigation
NUMA-aware allocation: opt-in design and failure modes #11
Replies: 2 comments
|
On Kube 1.36+, with the KEP-4816, The user can still request for devices with NUMA attribute match as a preferred constraint. This is potentially expressible today as : firstAvailable + matchAttribute on composite/numaNode. With composite driver: |
consumesCounters + TAS/Gang Scheduling IntegrationBeyond the consumesCounters (KEP-5075, beta K8s 1.36)The The composite driver can leverage this for consumable capacity (CPUs, memory, bandwidth) in two modes: Passthrough mode: Composite device's Locked-in mode: Each composite device gets its own CounterSet with pre-divided capacity (e.g., quarter-partition = 32 of 128 cores). Consumed entirely on allocation. Why this matters for TASKueue TAS currently has no DRA awareness — documented limitation: "DRA resources are not accounted for in TAS capacity calculations." However, with The composite driver's per-node ResourceSlices are the right primitive for this integration:
Publishing Intra-node bin packingThis connects to the NUMA bin packing discussion — if composite devices carry both topology attributes (
References: |
Uh oh!
There was an error while loading. Please reload this page.
Context
Removed automatic NUMA MatchAttribute constraints from the webhook (commit 624fbee). Webhook now generates plain DeviceRequests — scheduler picks devices freely. NUMA affinity is not enforced unless explicitly requested.
Problem with automatic NUMA constraints
With
"4/4"(4 pairs, all same NUMA), the webhook generated:If no single NUMA zone had 4 free pairs (but 4 existed across 2 zones), the scheduler hard-failed. The constraint is not best-effort — it is a hard requirement. Pod stays Pending even though devices exist.
The old webhook avoided this by scanning ResourceSlices and falling back to cross-NUMA when a single zone was insufficient. We explicitly do not scan ResourceSlices.
Options for NUMA support (when we add it back)
Option 1: Explicit annotation format
none(default): no constraints, scheduler picks freelypreferred: add constraints, but use separate claims so partial failure is tolerablerequired: hard MatchAttribute — fail if not satisfiableOption 2: NUMA-ordered ResourceSlices (driver-side)
Publish one ResourceSlice per NUMA zone. Allocator uses first-fit → naturally packs same-NUMA when capacity allows. No constraints needed. Heuristic, not guaranteed.
Option 3: Per-NUMA DeviceClasses + extended resources
Separate DeviceClasses per NUMA zone with
extendedResourceName. User explicitly requests from a NUMA zone. Guaranteed but requires user to know topology.Option 4: Webhook scans ResourceSlices
Add ResourceSlice awareness back to the webhook to determine feasible NUMA grouping before generating constraints. Adds coupling and complexity — moves toward the old webhook model.
Current position
Do not manage NUMA unless explicitly asked for. Simple count-based allocation covers the common case. NUMA optimization is an advanced feature that should be opt-in and fail-safe.
Questions
numa-policyannotation, or should NUMA be a DeviceClass concern?All reactions