Implement context-aware single-command CLI - #31
Conversation
This change completely reimagines the Continuum CLI experience by creating a single-command entry point that intelligently detects context and performs the appropriate action: - Detects and analyzes the current environment (git repo, config files, etc.) - Automatically initializes when no configuration exists - Suggests next steps based on what's missing - Provides direct integration with AI assistants via --ask option - Maintains backward compatibility with command-style usage - Creates a frictionless, intent-based developer experience 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
|
This PR exceeds the recommended size of 1000 lines. Please make sure you are NOT addressing multiple issues with one PR. Note this PR might be rejected due to its size. |
There was a problem hiding this comment.
Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.
Comments suppressed due to low confidence (3)
packages/cli/src/ask.js:14
- [nitpick] Consider renaming askAssistant to more explicitly indicate its purpose (for example, sendTaskToAssistant) to improve clarity.
export async function askAssistant(prompt, options = {}) {
packages/cli/src/ask.js:153
- The prompt variable is directly interpolated into the command string without escaping, which could lead to shell injection vulnerabilities when integrated into production.
return `echo "This would run: claude ask \"${prompt}\""`;
packages/cli/src/ask.js:162
- The prompt variable is directly interpolated into the command string without proper escaping, posing a potential security risk for command injection.
return `echo "This would run: openai api chat \"${prompt}\""`;
…egration - Implement context detection to auto-determine appropriate actions - Add environment and repository awareness to configurations - Enhance assistant integration with improved prompt generation - Update README with compatibility grid for AI models - Add new keywords to package.json for broader AI model coverage 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
|
This PR exceeds the recommended size of 1000 lines. Please make sure you are NOT addressing multiple issues with one PR. Note this PR might be rejected due to its size. |
- Add GPT.json symlink created by our own CLI - Add .clauderc for Claude Code integration - Add .gptrc for GPT integration - Dogfood our own tool to manage configurations 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
|
This PR exceeds the recommended size of 1000 lines. Please make sure you are NOT addressing multiple issues with one PR. Note this PR might be rejected due to its size. |
- Add CLAUDE.md symlink created by our CLI tool - Complete our dogfooding by using all available integrations - Ensure proper routing through our configuration files 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
|
This PR exceeds the recommended size of 1000 lines. Please make sure you are NOT addressing multiple issues with one PR. Note this PR might be rejected due to its size. |
Implement context-aware single-command CLI
…g dataset (#1691) The rooms→training-data bridge of the coordination↔learning flywheel: a persona's recorded room turns become the data that trains its LoRA genome. Coordination (rooms) IS the training data for learning (genome) — and a better genome collaborates better, producing richer turns. This is the first concrete brick of that loop. `dataset/from-turns` reads the recorder's per-turn JSON captures (default ~/.continuum/fixtures/persona-respond), converts each SPOKE turn (system prompt + user message → the persona's spoken response) into a chat SFT example in the SAME {messages:[{role,content}]} format the CSV importer emits, and flows them through the SAME split/write/manifest path (split_and_write). The JSONL is unsloth's canonical training input. Reuses dataset.rs end-to-end (no parallel allocator). Source = recorder turns, NOT engrams — engrams are curated recall memory, not paired SFT turns; only `spoke` turns are training pairs (silent/errored/malformed are skipped). Params: turnsDir, name, splitRatio, includeSystem, includeHistory (role-attributed by sender), personaId/roomId filters, outputDir. Tests (extend the existing dataset test mod): spoke turns → 3-message SFT example through the real write path; a turnsDir with no spoke turns errors loudly rather than writing an empty dataset. Roadmap: #30 (this), #31 (unsloth inference adapter), #32 (close the loop: ForgeRecipe→unsloth train→LoRA→genome page-in). See memory coordination-learning-flywheel. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…g.env template (#1692) Lean into the simplification: continuum reaches ALL models through unsloth (local llama.cpp + cloud providers, via the provider keys configured inside unsloth Studio). So continuum needs ONE credential, not a per-provider list. config.env.example now leads with a MODEL GATEWAY — UNSLOTH section: - UNSLOTH_API_KEY (the single credential continuum needs) - UNSLOTH_BASE_URL (default http://127.0.0.1:8888/v1) …and demotes the per-provider ANTHROPIC_API_KEY/OPENAI_API_KEY/etc. block to "legacy / optional — prefer the unsloth gateway". New-user setup is "open unsloth Studio, configure keys there, paste the one key here". Both readers already auto-load *_API_KEY (Rust secrets.rs + the cached config_env.rs owner). Mechanism slices tracked separately: install autowire that opens Studio + captures the key via config_env::upsert (#33), the UnslothInferenceAdapter gateway (#31), and the config.env single-owner consolidation that deprecates the TS SecretManager write path (#34). See memory unsloth-universal-model-gateway. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…1693) unsloth is a local OpenAI-compatible server (like docker-model-runner): it serves local models (llama.cpp) AND fans out to cloud providers via the keys the user configures in unsloth Studio. So continuum reaches every model through ONE provider with ONE credential (UNSLOTH_API_KEY) — no new adapter, just a catalog entry + registration on the existing OpenAICompatibleAdapter. - model_registry/catalog.rs: add the `unsloth` provider (base_url http://127.0.0.1:8888/v1, auth Bearer, api_key_env UNSLOTH_API_KEY). Dynamic catalog — the live model list comes from /v1/models (empty model_prefixes, runtime discovery), exactly like the DMR precedent. - modules/ai_provider.rs: register the unsloth adapter when UNSLOTH_API_KEY is configured (key-gated, same pattern as the cloud providers). Additive priority for now; making the gateway the *preferred* route is a separate routing-policy change. Live-validated end-to-end against a running unsloth Studio (tests/ unsloth_gateway_live.rs, skip-if-not-configured/unreachable): builds the SAME adapter the core builds, reads the key from ~/.continuum/config.env, authenticates + fetches the live /v1/models catalog. Proves real key + real bearer auth + real transport to the running gateway — the wiring unit tests (which mock dispatch) can't cover. ✓ unsloth gateway live: provider_id=unsloth, /v1/models returned 0 model(s) (0 = no model loaded in Studio yet; init succeeding = auth + transport OK) Part of the unsloth universal-model-gateway direction (#31). Follow-ups: make it the preferred route + UNSLOTH_BASE_URL runtime override (#20), install autowire (#33), close the genome loop (#32). See memory unsloth-universal-model-gateway. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
… first local head-to-head (#1919) Nous Hermes-3-Llama-3.1-8B registered as a servable local model (Llama-3.1 arch, ChatML, ~5GB Q4_K_M) so we can run the SAME coder gym through our OWN system for an honest local-vs-local head-to-head. First result (HumanEval-Rust, 40 tasks, single-pass, valid test_grade — real compile+run, 0 inference errors): OUR local (Qwen2.5-Coder-14B): 88% Hermes-3-Llama-3.1-8B: 42% Honest read: much of the gap is MODEL-FIT — Hermes-3 is a general model, we serve a code specialist (and 14B vs 8B). That's not a trick; it IS the thesis (pick/tune the best-fit local model for the ask — the thing cloud can't do). It is NOT "our loop made an equal model win" — the purer system claim is to run Hermes THROUGH our agentic loop + PX and show the lift, which the unsloth /v1 adapter (#31) makes a one-seam swap. Local beats Hermes on the coder task; the deeper agentic-tool-use comparison is the follow-up. Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Pure calc the governor uses to fit a streaming-MoE's RESIDENT (non-expert) tier to a device VRAM budget and reconcile it with the expert tier on ONE budget — fixing the double-count where the expert pager was handed the full VRAM ceiling while resident silently ate most of it. Partition (in order): compute reserve -> resident (Native | device-fit Override | Unfittable) -> sufficient-context KV -> everything left = hot-expert VRAM budget (maximized: more on-GPU experts, fewer streams). Context is derived + clamped, never hand-picked. Artifact resolver injected (no hardcoded paths). Standalone-validated 7/7; M5 wires it into the daemon spawn path + launch (ServingTarget.resident_override) per the K3 sprint split. Refs #29 #31 #36. Arch-confirmed on real K3 UD-IQ2 (93 blk/896 exp/top-16). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
#2109) * chore(k3): bump llama.cpp submodule to eee635ba2 — K3 serving stack onto canary Advances the vendored llama.cpp fork 30 commits (clean FF over canary's stale 66594cc3f): container-serve resident-override (LLAMA_RESIDENT_OVERRIDE), the rung-2 ResidencyCache plan-file consumer, the score-hint/generation-bias actuator, PagerCaptureEvent emit, fit-device --reserve-gb. Makes canary USE the K3 misfit-serving stack (measured 0.33 tok/s WASTE-parity on a 32GB card). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(pager): RUN-1 K3 trace fixture + tkey->(layer,matrix) table for M5's replay Live GGML_MOE_TRACE_FILE slice (12B records: u64 tkey + u32 e) + the reverse table so BanditPlanController recovers (layer,expert): tkey=FNV-1a of blk.{layer}.ffn_{gate,up,down}_exps.weight, e=within-layer expert idx, expert identity=(layer,e) deduped across the 3 matrices. RUN-1 static-pin datum input. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(pager): reference RL-policy prototypes for M5's TierPolicy port The actual std-only Rust prototypes written against live K3 traces this session: trace_replay (recency beats LFU 3-4x), predictor (offline learned-decay +5pts held-out), online_predictor (bandit 49.8 vs 47.8 best-fixed on non-stationary), self_optimize (joint speed×quality). These are the faithful-port source for the learned policy behind TierPolicy (continuum-core expert_tier_policy.rs, #276). Numbers are properties of these exact constants + reward math — reproduce before improving. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(k3): GPU-resident hot experts design (task #23 'trend to full GPU') The major GPU speedup: promote hot experts to persistent VRAM so decode's hot path is GPU-native (zero fetch, zero copy). 3 increments (copy-skip -> VRAM hot cache -> pipeline), the 32GB rate-distortion constraint (imatrix-enabled resident shrink frees VRAM for the hot set), measured per-increment via k3-bench. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(k3): flag the input_cpy-persistence question gating increment 1 vs 2 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(k3): modular rework-proof impl for #23 — reuse ResidencyCache + DeviceUploadFetcher Mechanism is the existing (buft,fetcher)-generic ResidencyCache; a VRAM cache = same class + device buft + host->device fetcher. 3 small parameterized pieces (DeviceUploadFetcher, instantiate w/ GGML_MOE_VRAM_CACHE_GB, seam hook). Stats only tune params -> zero mechanism rework. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(arch): MoE serving on a governed budget (draft; M5 owns the governor seam) Diagnoses the hardcoded-cache overcommit that collapsed K3 fetch bandwidth (40GB pinned + mmap = 95.9GB on 63GB -> pagefile thrash -> 205 MB/s -> 0.027 tok/s) and lays out the clean architecture: governor owns the residency budget net of the model's mmap footprint, plan-file is the one wire, ResidencyCache is pure mechanism. Governor-interface sections marked [M5 OWNS] for her to edit. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(arch): answer the [M5 OWNS] governor-budget seam in place (net-of-mmap is explicit arithmetic; plan_file.budget_bytes is the lease wire) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo * docs(arch): measured governed-budget inputs + graduated serving/load path Records the BigMama measurements feeding M5's #287 derivation (non-cache ~56GB, per-token working set 5.5GB, governed budget ~6GB, fetch recovers to 2.5GB/s at fit), the now-complete C++ cache mechanism (enable-from-plan, grow, shrink), and the three-piece graduated path to serving/load kimi-k3 (catalog row + serving-lane MoE launch + #287) replacing the rigged .bat. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * windows: make continuum-core build + link on windows-msvc (first time) start-server.sh now provisions the full Windows CUDA build env before the cargo builds (the cargo/nvcc path had none, unlike the vcvars-wrapped llama cmake): import MSVC via vswhere->VS2022-14.4x + a .bat env dump (cl.exe for nvcc), pin CMAKE to the manifest install, force CMAKE_GENERATOR=Ninja (the VS18-2026 auto- pick is undefined in cmake 3.30), add the Windows SDK bin (mt.exe/rc.exe), select a complete CUDA toolkit + CUDA_PATH (a provisioning split left cuda-env with 0 import libs vs cuda-13.2's 12), and RUSTFLAGS -L for pocket-tts (which emits no link-search) while re-carrying +crt-static so the /MT GPU stack still links. Portability: expert_container.rs + commands/capacity.rs used Unix-only std::os::unix::fs::FileExt::read_exact_at. Add crate::platform_io::pread_exact (unix read_exact_at / windows seek_read loop) - one place for positioned reads. Build validated (npm start exit 0, continuum-core lib clean). A separate runtime hot-loop on the #2088 core at startup is tracked apart from this build fix. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(capacity): device_fit VRAM-partition calc for the governor Pure calc the governor uses to fit a streaming-MoE's RESIDENT (non-expert) tier to a device VRAM budget and reconcile it with the expert tier on ONE budget — fixing the double-count where the expert pager was handed the full VRAM ceiling while resident silently ate most of it. Partition (in order): compute reserve -> resident (Native | device-fit Override | Unfittable) -> sufficient-context KV -> everything left = hot-expert VRAM budget (maximized: more on-GPU experts, fewer streams). Context is derived + clamped, never hand-picked. Artifact resolver injected (no hardcoded paths). Standalone-validated 7/7; M5 wires it into the daemon spawn path + launch (ServingTarget.resident_override) per the K3 sprint split. Refs #29 #31 #36. Arch-confirmed on real K3 UD-IQ2 (93 blk/896 exp/top-16). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(serving): resident-override plumbing on ServingTarget + launcher Wire foundation for the governor's device_fit plan: ServingTarget carries resident_override: Option<PathBuf>, and the launcher exports it as LLAMA_RESIDENT_OVERRIDE so llama.cpp sources the precision-shrunk RESIDENT (non-expert) tensors from the device-fit GGUF (all offloaded to GPU) while the primary streams experts. All builders updated; defaults None (resident serves as-shipped, no behavior change) until compute_resident_override + the resolve-or-generate resolver (#35) land next. In-crate validated. Refs #29 #36. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp fork to k3-adopt e3ce51df5 M5's per-layer KV accessors (n_head_kv_il + n_embd_head_{k,v}_il, continuum #238) + graph reconciliation. The K3 engine now builds against these — enables the device_fit resident-override serve + honest per-layer K3 KV sizing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(serving): compute_resident_override — wire device_fit into the plan The governor now DECIDES the resident source per serve: compute_resident_override derives resident_bytes (weights - expert_bytes_total) vs the governed VRAM ceiling via capacity::device_fit, and sets ServingTarget.resident_override. A dense/small model fits native (None); a >VRAM-resident MoE (K3) resolves a cached device-fit override that fits, else Unfittable → route to grid / generate (#35), glass-boxed. resolve_device_fit_override (model_registry::artifacts): looks up a per-user device-fit cache convention (<storage_root>/device-fit/<id>/) + a resident-bytes sidecar; returns the override only when its resident fits the usable budget. No hardcoded paths; generation/HF discovery is #35. The resident-fit decision turns only on resident_bytes vs budget — per-layer KV (#2107 ModelCapabilities) drives the context/expert split elsewhere, so KV is not consulted here. Refs #29 #35 #36. In-crate validated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(arch): storage serving-tier governor — NVMe<->cold contention managed like VRAM/RAM Joel: 'like vram and memory, this contention has to be managed between cold storage and nvme.' Design: NVMe is a governed HOT-SERVING tier (a ResourcePool, same TrackedDir + evict_at_least machinery as CargoTargetPool), whose eviction = MIGRATE frozen/duplicate artifacts to the Cold drive, not a manual rm. Serving asks ensure_hot_resident(model); composes with device_fit's Unfittable one tier down (VRAM). Corrects the DriveRole bug: Cold (HDD) is FROZEN storage, never the per-token streaming tier (HDD = unservable). Dissolves today's K3 container disk fight: the C: IQ2 is a verified duplicate of the D: copy -> governor migrates it off NVMe -> container fits, no human deletes anything. Refs #12 #36. Design for M5's system_resources lane. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(capacity): verified cold-twin detection — safe-to-drop primitive for the storage tier The gate the NVMe serving-tier eviction (#302) consults before dropping a frozen GGUF: is an IDENTICAL twin already on cold storage? is_structural_twin (pure) = same shard count + per-shard name + size, zero-byte shards never match. scan_shards + find_cold_twin are the thin fs layer. Never drop an NVMe artifact without a VERIFIED cold twin (dropping 662GB on a path guess is the failure this guards). Standalone-validated 5/5. Composes with device_fit + M5's NvmeServingTierPool. Refs #12 #36. Design: STORAGE-SERVING-TIER-GOVERNOR.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp fork to k3-adopt c6469d5 — container-serve wired Both halves of the DirContainerFetcher wire (BigMama fetcher + moe_pick_fetcher branch 175ac9d6a; M5 caller-side encode + record_bytes reader c6469d5). Serving now reads the aligned per-layer container (GGML_MOE_CONTAINER) instead of the scattered raw GGUF — the honest ~2.6GB/s path. Retires the built-not-wired ContainerFetcher. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * fix(serving): resident_override on vision_sidecar ServingTarget + K3 coverage measurement Merge fix: vision_sidecar's ServingTarget was missing resident_override (added by #29). Plus a measurement test that drains the real K3 routed-access fixture through the #282 predictive instrument and prints repeat_recall / predicted_delta / schedulable_coverage — the go/no-go for the LiveUploadPager predictive pipeline (H2D/token = (1 - coverage) x ~11GB). Prints, never asserts (real routing sample). NOTE: can't run on windows-msvc (pre-existing cargo-test Unix-socket block, ipc/mod.rs); runs on M5's Mac. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(pager-driver): offline warm-coverage measurement (--synth-layers, --once, --budget-slots) The moe-pager-driver gains an offline replay mode so any completed GGML_MOE_TRACE_FILE can be scored on any box (windows-msvc clean by crate constraint), not just tailed live next to a serve: - --synth-layers N: synthesize the tkey->layer map from layer count alone (TkeyTable::for_layers, the same zero-config seam MoeTraceTail owns) instead of requiring an operator tkey-to-layer-matrix.json. - --once: exit when the trace stops growing (EOF) and print a SUMMARY line with mean DECODE-token serving hit = warm schedulable coverage. - --budget-slots N: override the predictor residency budget (default auto = first token x1.5) to measure the coverage-vs-free-VRAM curve (the device-fit tradeoff). Measured on BigMama run2.trace (302 warm decode tokens): bandit coverage 13.8% @250 slots -> 51.3% @2000 -> 65.7% @4024, beating naive last-N recency by +7-9pts at matched VRAM. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(pager): measure cross-layer prefetch predictor ceiling — DEAD lever for K3 VDD offline measurement (cooccur-ceiling bin) on the real warm serve trace (run2.trace, 122 held-out decode tokens, 11102 layer-steps): cross_layer_cooccur_hit 0.159 (adjacent-layer noisy-OR) recency_same_layer_hit 0.403 (last token, same layer) structure beyond recency -0.244 cooccur_recall_on_recency_misses 0.112 (11824/106058) Adjacent-layer co-occurrence predicts <half what plain recency does, and recovers only 11% of the experts recency misses (~base rate). K3 expert routing has no exploitable cross-layer structure — the CrossLayerExpert- Predictor prefetch lever is not worth wiring (saves the ggml pass-id capture slice). Recency-family residency (the bandit EMA curve) is THE signal; the only lever that lifts K3 is freeing VRAM (device-fit shrink) so residency coverage can reach the measured 51%. Caveat: adjacent-layer, one workload trace. Wider-predecessor noisy-OR would regress toward the frequency baseline (which underperforms recency), so a large lift is unlikely — but not measured here. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp to de29843e0 — device-resident expert cache half (#23) Pins the fork at the DeviceUploadFetcher wiring (my half of the LiveUpload- Pager H2D-kill). Off unless GGML_MOE_VRAM_CACHE_GB / plan device_budget_bytes enables it; host serving path byte-for-byte unchanged. M5's expert-loop D2D half lands next on the same seam. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp to 0fbe4e27a — quantize --resident-only + tier manifest (#40) Enables the device-fit division: produce a small resident override per precision tier + a (tier_label, resident_bytes) sidecar the governor reads to co-optimize the VRAM split. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(pager): DivisionPolicy — the governor's VRAM-division RL brain (#2/#3) The second control rung above the pager's DecayBandit. The pager decides WHICH experts stay resident (reward=hit-rate, cheap, online). This decides HOW TO DIVIDE the card — resident (non-expert) weights vs expert cache — to MAXIMIZE tok/s. That reward (actual tok/s) is EXPENSIVE (a serve), so naive online RL flails; the fix is SIM-WARM-START: predict tok/s per division OFFLINE from the measured coverage curve, then a slow bandit refines each arm from real measured tok/s. - CoverageModel: piecewise-linear coverage(slots) over MEASURED points (k3_measured() = the trace-replay curve); saturates, never extrapolates up. - predict_tok_s: coverage -> (1-coverage)*experts/token*expert_bytes H2D -> t_token -> tok/s. Higher coverage -> less H2D -> faster (the load-bearing property, tested). - feasible_divisions: tier catalog (from --resident-only manifests) x HardwareBudget -> cache budget/slots per tier; drops VRAM-overflow tiers. - DivisionBandit: warm_start from the predictor; observe(tier, measured_tok_s) overrides the prior on first serve then EMAs — the expensive reward spent only on the arm actually run. Policy lives here (windows-clean, 4 tests pass); serving_daemon actuates it (M5's #2: discover manifests, feed catalog+budget+live tok/s, apply the chosen {resident_tier, device_budget_bytes} to the plan). Fractal control law: pager (experts<->hit-rate) -> this (VRAM split<->tok/s) -> grid. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp to 2a32025dd — device-cache un-crashable (clamp to free VRAM + null-buffer guard, #23) Testing convicted the segfault as VRAM oversubscription (K3 33GB resident + env cache on a 32GB card, cudaMalloc lazy-VMM deferred fault). Fix: clamp device budget to measured free VRAM (mine) + M5's D2D null-buffer guard. Device cache now disables safely where there's no room (K3) and works where there is (V4-Flash); can't crash from any budget source. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * test(division): bandit learns residency saturation from measured V4-Flash curve Feeds the real BigMama RTX 5090 --n-cpu-moe sweep (DeepSeek-V4-Flash UD-IQ2_M) into DivisionBandit: 0 resident=1.39, 8 resident=1.69, 14 resident=1.68 tok/s. Asserts the bandit converges on the SATURATION KNEE (8 layers), not max residency — 8->14 layers buys nothing at +11GB VRAM. Encodes the measured finding that the governor must learn 'minimal static residency + max device cache', the freed VRAM belonging to the recency cache (#43), not to over-pinned static layers. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * test(division): bandit finds non-monotonic device-cache budget optimum Measured V4-Flash device-cache coverage curve (5090, GGML_MOE_VRAM_CACHE_GB sweep): 6GB/992slots=1.80, 12GB/1985=3.10, 22GB/3630=2.96 tok/s, all 100% hit. tok/s is NON-MONOTONIC in budget: undersized churns, 12GB is the plateau knee, 22GB is no better (100% hit but O(slots) reserve_slot eviction scan). predict_tok_s's monotonic prior would pick 22GB; only the measured reward lands on 12GB — which frees ~20GB of a 32GB card for co-resident lanes. Pins the invariant that the governor must not oversize the cache and starve other models. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp to fa7e0d8e9 — #43 device-cache fix + async restore + enum fix Includes the prefetch host_visible guard (THE #43 crash fix, validated 3.05 tok/s V4-Flash device cache on the 5090), M5's async cpy_tensor_async restore, and the moe-pack quant-enum fix. A fresh parent build now includes the un-crashable device cache instead of the pre-fix pin. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Summary
This PR completely reimagines the Continuum CLI experience by creating a single-command entry point that intelligently detects context and performs the appropriate action:
The Philosophy Shift
This redesign represents a fundamental shift in how we think about the CLI:
Usage Examples
Implementation Notes
This implementation is a proof of concept that demonstrates the core ideas. It includes:
Testing
I've tested this implementation in a variety of scenarios:
🤖 Generated with Claude Code