Skip to content

feat(serving): division actuation — daemon consumes DivisionPolicy (#2 of the split) - #2133

Open
joelteply wants to merge 31 commits into
canaryfrom
feat/division-actuation
Open

feat(serving): division actuation — daemon consumes DivisionPolicy (#2 of the split)#2133
joelteply wants to merge 31 commits into
canaryfrom
feat/division-actuation

Conversation

@joelteply

Copy link
Copy Markdown
Contributor

Stacked on #2109 (this PR targets k3-serving-canary so the split ships as one lineage).

My half of the resident/cache division contract (her DivisionPolicy fe624ea is the brain, her --resident-only quantize tool #40 is the manifest producer):

  • Tier discovery: *.resident.json manifests from the device-fit cache dir (tolerant key-scan; a manifest whose GGUF vanished is dropped). Native as-shipped resident is always tier 0.
  • Warm-started bandit: feasible_divisionsDivisionBandit::warm_start with the measured K3 coverage curve; the chosen resident-tier label is published as a fourth axis on the SAME governed plan file.
  • One budget authority: the live device budget stays Resource lifecycle: load on demand, unload after idle #305's board-derived axis — this slice never publishes a budget. A relaunch that adopts a smaller resident frees VRAM the board then sees; Resource lifecycle: load on demand, unload after idle #305 grows into it organically.
  • Two-speed: publishing resident_tier never triggers a relaunch; it actuates when a relaunch happens for its own reasons.
  • Honest reward: trace-tail token-watermark delta per tick = measured decode tok/s, credited to the tier the spawn ACTUALLY loaded (tracked at reconcile), never the bandit's unlaunched choice. Noise deltas (<64 tokens) rejected; watermark resets re-seed silently.
  • Also resolves the platform_io/fs_portable duplicate from the canary merge to the one canonical fs_portable.

Validated: division_actuation 4 tests + serving_daemon 30 tests + expert-pager-policy 23 tests green with --features metal,accelerate.

Follow-ups (named, not in this slice): spawn-time tier selection unified with the bandit choice (needs resolver #35), per-model coverage curves, fork-side resident_tier consumption at relaunch.

🤖 Generated with Claude Code

https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo

joelteply and others added 30 commits August 1, 2026 03:49
…nto canary

Advances the vendored llama.cpp fork 30 commits (clean FF over canary's stale
66594cc3f): container-serve resident-override (LLAMA_RESIDENT_OVERRIDE), the
rung-2 ResidencyCache plan-file consumer, the score-hint/generation-bias
actuator, PagerCaptureEvent emit, fit-device --reserve-gb. Makes canary USE the
K3 misfit-serving stack (measured 0.33 tok/s WASTE-parity on a 32GB card).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…M5's replay

Live GGML_MOE_TRACE_FILE slice (12B records: u64 tkey + u32 e) + the reverse
table so BanditPlanController recovers (layer,expert): tkey=FNV-1a of
blk.{layer}.ffn_{gate,up,down}_exps.weight, e=within-layer expert idx, expert
identity=(layer,e) deduped across the 3 matrices. RUN-1 static-pin datum input.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
The actual std-only Rust prototypes written against live K3 traces this
session: trace_replay (recency beats LFU 3-4x), predictor (offline
learned-decay +5pts held-out), online_predictor (bandit 49.8 vs 47.8
best-fixed on non-stationary), self_optimize (joint speed×quality). These
are the faithful-port source for the learned policy behind TierPolicy
(continuum-core expert_tier_policy.rs, #276). Numbers are properties of
these exact constants + reward math — reproduce before improving.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
The major GPU speedup: promote hot experts to persistent VRAM so decode's hot
path is GPU-native (zero fetch, zero copy). 3 increments (copy-skip -> VRAM hot
cache -> pipeline), the 32GB rate-distortion constraint (imatrix-enabled resident
shrink frees VRAM for the hot set), measured per-increment via k3-bench.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…DeviceUploadFetcher

Mechanism is the existing (buft,fetcher)-generic ResidencyCache; a VRAM cache =
same class + device buft + host->device fetcher. 3 small parameterized pieces
(DeviceUploadFetcher, instantiate w/ GGML_MOE_VRAM_CACHE_GB, seam hook). Stats
only tune params -> zero mechanism rework.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…rnor seam)

Diagnoses the hardcoded-cache overcommit that collapsed K3 fetch bandwidth
(40GB pinned + mmap = 95.9GB on 63GB -> pagefile thrash -> 205 MB/s -> 0.027
tok/s) and lays out the clean architecture: governor owns the residency budget
net of the model's mmap footprint, plan-file is the one wire, ResidencyCache is
pure mechanism. Governor-interface sections marked [M5 OWNS] for her to edit.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…f-mmap is explicit arithmetic; plan_file.budget_bytes is the lease wire)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo
…path

Records the BigMama measurements feeding M5's #287 derivation (non-cache ~56GB,
per-token working set 5.5GB, governed budget ~6GB, fetch recovers to 2.5GB/s at
fit), the now-complete C++ cache mechanism (enable-from-plan, grow, shrink), and
the three-piece graduated path to serving/load kimi-k3 (catalog row + serving-lane
MoE launch + #287) replacing the rigged .bat.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
start-server.sh now provisions the full Windows CUDA build env before the cargo
builds (the cargo/nvcc path had none, unlike the vcvars-wrapped llama cmake):
import MSVC via vswhere->VS2022-14.4x + a .bat env dump (cl.exe for nvcc), pin
CMAKE to the manifest install, force CMAKE_GENERATOR=Ninja (the VS18-2026 auto-
pick is undefined in cmake 3.30), add the Windows SDK bin (mt.exe/rc.exe), select
a complete CUDA toolkit + CUDA_PATH (a provisioning split left cuda-env with 0
import libs vs cuda-13.2's 12), and RUSTFLAGS -L for pocket-tts (which emits no
link-search) while re-carrying +crt-static so the /MT GPU stack still links.

Portability: expert_container.rs + commands/capacity.rs used Unix-only
std::os::unix::fs::FileExt::read_exact_at. Add crate::platform_io::pread_exact
(unix read_exact_at / windows seek_read loop) - one place for positioned reads.

Build validated (npm start exit 0, continuum-core lib clean). A separate runtime
hot-loop on the #2088 core at startup is tracked apart from this build fix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
Pure calc the governor uses to fit a streaming-MoE's RESIDENT (non-expert)
tier to a device VRAM budget and reconcile it with the expert tier on ONE
budget — fixing the double-count where the expert pager was handed the full
VRAM ceiling while resident silently ate most of it.

Partition (in order): compute reserve -> resident (Native | device-fit
Override | Unfittable) -> sufficient-context KV -> everything left =
hot-expert VRAM budget (maximized: more on-GPU experts, fewer streams).
Context is derived + clamped, never hand-picked. Artifact resolver injected
(no hardcoded paths). Standalone-validated 7/7; M5 wires it into the daemon
spawn path + launch (ServingTarget.resident_override) per the K3 sprint split.

Refs #29 #31 #36. Arch-confirmed on real K3 UD-IQ2 (93 blk/896 exp/top-16).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
Wire foundation for the governor's device_fit plan: ServingTarget carries
resident_override: Option<PathBuf>, and the launcher exports it as
LLAMA_RESIDENT_OVERRIDE so llama.cpp sources the precision-shrunk RESIDENT
(non-expert) tensors from the device-fit GGUF (all offloaded to GPU) while the
primary streams experts. All builders updated; defaults None (resident serves
as-shipped, no behavior change) until compute_resident_override + the
resolve-or-generate resolver (#35) land next. In-crate validated.

Refs #29 #36.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
M5's per-layer KV accessors (n_head_kv_il + n_embd_head_{k,v}_il, continuum #238)
+ graph reconciliation. The K3 engine now builds against these — enables the
device_fit resident-override serve + honest per-layer K3 KV sizing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
# Conflicts:
#	tools/scripts/start-server.sh
The governor now DECIDES the resident source per serve: compute_resident_override
derives resident_bytes (weights - expert_bytes_total) vs the governed VRAM ceiling
via capacity::device_fit, and sets ServingTarget.resident_override. A dense/small
model fits native (None); a >VRAM-resident MoE (K3) resolves a cached device-fit
override that fits, else Unfittable → route to grid / generate (#35), glass-boxed.

resolve_device_fit_override (model_registry::artifacts): looks up a per-user
device-fit cache convention (<storage_root>/device-fit/<id>/) + a resident-bytes
sidecar; returns the override only when its resident fits the usable budget. No
hardcoded paths; generation/HF discovery is #35. The resident-fit decision turns
only on resident_bytes vs budget — per-layer KV (#2107 ModelCapabilities) drives
the context/expert split elsewhere, so KV is not consulted here.

Refs #29 #35 #36. In-crate validated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…naged like VRAM/RAM

Joel: 'like vram and memory, this contention has to be managed between cold
storage and nvme.' Design: NVMe is a governed HOT-SERVING tier (a ResourcePool,
same TrackedDir + evict_at_least machinery as CargoTargetPool), whose eviction =
MIGRATE frozen/duplicate artifacts to the Cold drive, not a manual rm. Serving
asks ensure_hot_resident(model); composes with device_fit's Unfittable one tier
down (VRAM). Corrects the DriveRole bug: Cold (HDD) is FROZEN storage, never the
per-token streaming tier (HDD = unservable). Dissolves today's K3 container disk
fight: the C: IQ2 is a verified duplicate of the D: copy -> governor migrates it
off NVMe -> container fits, no human deletes anything.

Refs #12 #36. Design for M5's system_resources lane.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
… for the storage tier

The gate the NVMe serving-tier eviction (#302) consults before dropping a frozen
GGUF: is an IDENTICAL twin already on cold storage? is_structural_twin (pure) =
same shard count + per-shard name + size, zero-byte shards never match. scan_shards
+ find_cold_twin are the thin fs layer. Never drop an NVMe artifact without a
VERIFIED cold twin (dropping 662GB on a path guess is the failure this guards).
Standalone-validated 5/5. Composes with device_fit + M5's NvmeServingTierPool.

Refs #12 #36. Design: STORAGE-SERVING-TIER-GOVERNOR.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…rve wired

Both halves of the DirContainerFetcher wire (BigMama fetcher + moe_pick_fetcher
branch 175ac9d6a; M5 caller-side encode + record_bytes reader c6469d5). Serving now
reads the aligned per-layer container (GGML_MOE_CONTAINER) instead of the scattered
raw GGUF — the honest ~2.6GB/s path. Retires the built-not-wired ContainerFetcher.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
# Conflicts:
#	core/continuum-core/src/modules/serving_daemon.rs
…coverage measurement

Merge fix: vision_sidecar's ServingTarget was missing resident_override (added by
#29). Plus a measurement test that drains the real K3 routed-access fixture through
the #282 predictive instrument and prints repeat_recall / predicted_delta /
schedulable_coverage — the go/no-go for the LiveUploadPager predictive pipeline
(H2D/token = (1 - coverage) x ~11GB). Prints, never asserts (real routing sample).
NOTE: can't run on windows-msvc (pre-existing cargo-test Unix-socket block, ipc/mod.rs);
runs on M5's Mac.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…, --once, --budget-slots)

The moe-pager-driver gains an offline replay mode so any completed
GGML_MOE_TRACE_FILE can be scored on any box (windows-msvc clean by
crate constraint), not just tailed live next to a serve:

- --synth-layers N: synthesize the tkey->layer map from layer count
  alone (TkeyTable::for_layers, the same zero-config seam MoeTraceTail
  owns) instead of requiring an operator tkey-to-layer-matrix.json.
- --once: exit when the trace stops growing (EOF) and print a SUMMARY
  line with mean DECODE-token serving hit = warm schedulable coverage.
- --budget-slots N: override the predictor residency budget (default
  auto = first token x1.5) to measure the coverage-vs-free-VRAM curve
  (the device-fit tradeoff).

Measured on BigMama run2.trace (302 warm decode tokens): bandit
coverage 13.8% @250 slots -> 51.3% @2000 -> 65.7% @4024, beating naive
last-N recency by +7-9pts at matched VRAM.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…ver for K3

VDD offline measurement (cooccur-ceiling bin) on the real warm serve
trace (run2.trace, 122 held-out decode tokens, 11102 layer-steps):

  cross_layer_cooccur_hit           0.159   (adjacent-layer noisy-OR)
  recency_same_layer_hit            0.403   (last token, same layer)
  structure beyond recency         -0.244
  cooccur_recall_on_recency_misses  0.112   (11824/106058)

Adjacent-layer co-occurrence predicts <half what plain recency does, and
recovers only 11% of the experts recency misses (~base rate). K3 expert
routing has no exploitable cross-layer structure — the CrossLayerExpert-
Predictor prefetch lever is not worth wiring (saves the ggml pass-id
capture slice). Recency-family residency (the bandit EMA curve) is THE
signal; the only lever that lifts K3 is freeing VRAM (device-fit shrink)
so residency coverage can reach the measured 51%.

Caveat: adjacent-layer, one workload trace. Wider-predecessor noisy-OR
would regress toward the frequency baseline (which underperforms recency),
so a large lift is unlikely — but not measured here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…ache half (#23)

Pins the fork at the DeviceUploadFetcher wiring (my half of the LiveUpload-
Pager H2D-kill). Off unless GGML_MOE_VRAM_CACHE_GB / plan device_budget_bytes
enables it; host serving path byte-for-byte unchanged. M5's expert-loop D2D
half lands next on the same seam.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
… + tier manifest (#40)

Enables the device-fit division: produce a small resident override per
precision tier + a (tier_label, resident_bytes) sidecar the governor reads
to co-optimize the VRAM split.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
#3)

The second control rung above the pager's DecayBandit. The pager decides
WHICH experts stay resident (reward=hit-rate, cheap, online). This decides
HOW TO DIVIDE the card — resident (non-expert) weights vs expert cache —
to MAXIMIZE tok/s. That reward (actual tok/s) is EXPENSIVE (a serve), so
naive online RL flails; the fix is SIM-WARM-START: predict tok/s per
division OFFLINE from the measured coverage curve, then a slow bandit
refines each arm from real measured tok/s.

- CoverageModel: piecewise-linear coverage(slots) over MEASURED points
  (k3_measured() = the trace-replay curve); saturates, never extrapolates up.
- predict_tok_s: coverage -> (1-coverage)*experts/token*expert_bytes H2D ->
  t_token -> tok/s. Higher coverage -> less H2D -> faster (the load-bearing
  property, tested).
- feasible_divisions: tier catalog (from --resident-only manifests) x
  HardwareBudget -> cache budget/slots per tier; drops VRAM-overflow tiers.
- DivisionBandit: warm_start from the predictor; observe(tier, measured_tok_s)
  overrides the prior on first serve then EMAs — the expensive reward spent
  only on the arm actually run.

Policy lives here (windows-clean, 4 tests pass); serving_daemon actuates it
(M5's #2: discover manifests, feed catalog+budget+live tok/s, apply the
chosen {resident_tier, device_budget_bytes} to the plan). Fractal control
law: pager (experts<->hit-rate) -> this (VRAM split<->tok/s) -> grid.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…tion

# Conflicts:
#	core/continuum-core/src/capacity/expert_container.rs
#	core/continuum-core/src/commands/capacity.rs
#2 of the resident/cache split)

The serving_daemon half of the contract (BigMama's DivisionPolicy fe624ea is the
brain): discover the --resident-only tier manifests (#40) from the device-fit cache
dir, warm-start a DivisionBandit over the feasible divisions, and publish the chosen
resident-tier label as a fourth axis on the SAME governed plan file.

Division of authority (compression): the LIVE device budget stays #305's board-derived
axis — this never publishes a budget. It owns only the TIER choice; when a relaunch
adopts a smaller resident, the board frees VRAM and #305's budget grows into it
organically. Two-speed: publishing resident_tier never triggers a relaunch.

Reward: the fork's trace-tail token watermark delta per tick = measured decode tok/s,
credited to the tier the spawn ACTUALLY loaded (served_resident tracked at reconcile),
never the bandit's latest unlaunched choice. Sub-64-token deltas rejected as noise;
watermark resets (relaunch) re-seed silently. First measurement replaces the offline
prior outright per the policy contract.

Pure parts (manifest parse, discovery, shape derivation, reward loop) live in
capacity/division_actuation.rs with 4 tests incl. the served-arm attribution and
no-churn publish invariants. Probes: serving.division / serving.division_reward.

Also resolves the platform_io/fs_portable duplicate from the canary merge to the ONE
canonical fs_portable (compression rule).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo
Base automatically changed from k3-serving-canary to canary August 3, 2026 18:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant