Build(deps-dev): Bump @typescript-eslint/eslint-plugin from 8.29.1 to 8.31.0 - #40
Closed
dependabot[bot] wants to merge 1 commit into
Closed
Conversation
Bumps [@typescript-eslint/eslint-plugin](https://github.com/typescript-eslint/typescript-eslint/tree/HEAD/packages/eslint-plugin) from 8.29.1 to 8.31.0. - [Release notes](https://github.com/typescript-eslint/typescript-eslint/releases) - [Changelog](https://github.com/typescript-eslint/typescript-eslint/blob/main/packages/eslint-plugin/CHANGELOG.md) - [Commits](https://github.com/typescript-eslint/typescript-eslint/commits/v8.31.0/packages/eslint-plugin) --- updated-dependencies: - dependency-name: "@typescript-eslint/eslint-plugin" dependency-version: 8.31.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com>
Contributor
Author
|
Superseded by #43. |
dependabot
Bot
deleted the
dependabot/npm_and_yarn/typescript-eslint/eslint-plugin-8.31.0
branch
April 28, 2025 22:20
joelteply
added a commit
that referenced
this pull request
Jun 21, 2026
§3.5 — probed the running unsloth Studio (:8888). Two HTTP surfaces: /v1
(OpenAI-compatible serving) + /api/* (Training & Model Management backend,
openapi at /openapi.json). The complete delegation surface, every capability
endpoint-backed:
- model lifecycle: POST /api/inference/load (+ /load-progress), /unload,
GET /api/inference/status|models, /v1/models
- capability checks: /api/models/check-embedding/{name}, /check-vision/{name}
- hub: POST /api/hub/download (HF pull), GET /api/hub/cached-models|cached-gguf
- serving: POST /v1/chat/completions, POST /v1/embeddings
- foundry handoff: POST /api/export/load-checkpoint, /api/export/export/gguf
- training data: /api/datasets/*, /api/data-recipe/*; health: GET /api/health
Implications: auto-load (#24) is clean HTTP (status→load→poll→health-probe),
not a CLI subprocess; embeddings (#40) verify via check-embedding then
/v1/embeddings; the foundry handoff is endpoint-backed.
CRITICAL live state: unsloth is running but EMPTY (/v1/models -> [],
/v1/embeddings -> "No GGUF model loaded"). Nothing serves until auto-load
runs — the prerequisite for peers-online and for testing #40. Local GGUFs
exist to load against.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
joelteply
added a commit
that referenced
this pull request
Jun 21, 2026
…in verified (#1711) Corrects the shallow earlier read. unsloth multimodal is NOT separate endpoints — vision-in + audio-in (STT) ride as OpenAI-style multimodal `content` parts (image_url, input_audio) on /v1/chat/completions; speech-out is /v1/audio/generate (a ChatCompletionRequest); per-model capability gated by load-response flags is_vision/is_audio/has_audio_input/is_diffusion. So the HTTP adapter handles ALL modalities through ONE chat seam carrying multimodal content + the right loaded model — multimodal is non-negotiable and fully covered. Avatar video calls = unsloth multimodal (see/hear/speak) ⊕ continuum live layer (LiveKitAgentManager + Bevy renderer, in main.rs). Live-state updated: VERIFIED a manual POST /api/inference/load of a local GGUF loads instantly and /v1/chat replies — the persona reasoning chain works end-to-end through unsloth on this box. (Embeddings need --embeddings or a dedicated embed model — the #40 detail.) Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
joelteply
added a commit
that referenced
this pull request
Jun 21, 2026
…dapter (#40) (#1714) Routes continuum's neural recall embeddings through the unsloth gateway's OpenAI-compatible /v1/embeddings — the first half of #40 ("embeddings → unsloth"). The cognition NeuralEmbeddingProvider already calls adapter.create_embedding(); until now the OpenAICompatibleAdapter inherited the trait's default Err ("does not support embeddings"), so every neural embed silently degraded to the lexical fallback. This implements it. - create_embedding(): POSTs /v1/embeddings (same base_url / Bearer-auth / error-surfacing shape as generate_text). Degrades to Err — never panics — on an unreachable endpoint or non-embedding model, so recall falls back to the lexical embedder rather than crashing. - supports_embeddings now true for "unsloth" as well as "openai" (the gateway exposes OpenAI-compatible /v1/embeddings). - The model is required (no silent default among chat models, [[no-fallbacks-ever]]) — it IS the embedding-space identity the cache keys on. Pure helpers, TDD'd apart from the HTTP I/O: - build_embedding_body() — single → string, batch → array - parse_embedding_response() — orders vectors by the response `index` field (the spec does not guarantee input order; mis-ordering silently misaligns every vector with its source text), errors on missing data instead of fabricating a vector - parse_embedding_usage() — usage is observability, defaults to 0, never fails 5 new unit tests cover all of the above. This does NOT yet swap the live default off fastembed or delete the ort/fastembed embedder — that's the follow-up "trim" half of #40, which requires migrating the legacy sync memory/ consumers onto this async path. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
joelteply
added a commit
that referenced
this pull request
Jun 21, 2026
…#1715) The live persona recall path (build_workspace_cycle → RecallFaculty) hard-coded the lexical bootstrap embedder. Now it uses neural embeddings (via unsloth's /v1/embeddings, landed last commit) when the embed model actually serves, and falls back to lexical otherwise — so personas get semantic recall, not just word-overlap, the moment an embed model is available. resolve_recall_embedder(adapter): - Prefers NeuralEmbeddingProvider (semantic) but PROBES it once with a one-shot embed. A usable probe (non-empty, non-zero, all-finite) → neural; an empty / zero / NaN probe (model not loaded, endpoint error) → lexical fallback. This is the no-signal guard: without it, an adapter that advertises embedding support but has no embed model loaded would embed every memory into a zero vector and recall would return nothing — strictly worse than lexical. - The choice is PROCESS-STABLE (decided once at spawn, never per-embed): a query and the stored vectors must live in the SAME embedding space, so neural and lexical are never mixed per call (cosine across spaces is meaningless). - Result is wrapped in the content-addressed CachingEmbeddingProvider — each message embedded ONCE and shared across every persona (the latency win; the embed model slug is the cache/space key). Further perf is a later pass. - Always returns a working embedder — never errors/panics. A box with no embed model still gets real lexical relevance ("solve for public users"). Wiring: - PersonaBrainConfig gains `embedder: Option<Arc<dyn EmbeddingProvider>>`; build_workspace_cycle uses it, defaulting to lexical+cache when None (harnesses). - The live supervisor spawn sets it via resolve_recall_embedder(adapter). - CANONICAL_EMBED_MODEL = "qwen3-embedding-0.6b" (dim 1024), overridable via UNSLOTH_EMBED_MODEL — one embedding space across the grid. 7 new tests: the probe gate (usable vs no-signal), and the resolver picking neural when the model serves / falling back to lexical on empty-probe / using lexical when the adapter lacks embedding support. Scope: this does NOT delete fastembed — that path (the separate IPC Hippocampus / PersonaMemoryManager, with real blast radius into ORM vector + search) is the follow-up trim, deferred per "get it working first, optimize once healthy." Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
joelteply
added a commit
that referenced
this pull request
Aug 3, 2026
…f of the resident/cache split seam (#2129) The split policy (replaces resident-first-scraps-last): serving_daemon will pick the (resident_tier, cache_bytes) pair maximizing predicted tok/s from the measured coverage curve, writing both into the ONE plan file. This lands the wire field + optional-contract tests now so BigMama's --resident-only quantize tool (#40 enabler) and my policy build against the same name in parallel; absent never serializes, prior documents parse unchanged. Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
joelteply
added a commit
that referenced
this pull request
Aug 3, 2026
… + tier manifest (#40) Enables the device-fit division: produce a small resident override per precision tier + a (tier_label, resident_bytes) sidecar the governor reads to co-optimize the VRAM split. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
joelteply
added a commit
that referenced
this pull request
Aug 3, 2026
#2109) * chore(k3): bump llama.cpp submodule to eee635ba2 — K3 serving stack onto canary Advances the vendored llama.cpp fork 30 commits (clean FF over canary's stale 66594cc3f): container-serve resident-override (LLAMA_RESIDENT_OVERRIDE), the rung-2 ResidencyCache plan-file consumer, the score-hint/generation-bias actuator, PagerCaptureEvent emit, fit-device --reserve-gb. Makes canary USE the K3 misfit-serving stack (measured 0.33 tok/s WASTE-parity on a 32GB card). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(pager): RUN-1 K3 trace fixture + tkey->(layer,matrix) table for M5's replay Live GGML_MOE_TRACE_FILE slice (12B records: u64 tkey + u32 e) + the reverse table so BanditPlanController recovers (layer,expert): tkey=FNV-1a of blk.{layer}.ffn_{gate,up,down}_exps.weight, e=within-layer expert idx, expert identity=(layer,e) deduped across the 3 matrices. RUN-1 static-pin datum input. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(pager): reference RL-policy prototypes for M5's TierPolicy port The actual std-only Rust prototypes written against live K3 traces this session: trace_replay (recency beats LFU 3-4x), predictor (offline learned-decay +5pts held-out), online_predictor (bandit 49.8 vs 47.8 best-fixed on non-stationary), self_optimize (joint speed×quality). These are the faithful-port source for the learned policy behind TierPolicy (continuum-core expert_tier_policy.rs, #276). Numbers are properties of these exact constants + reward math — reproduce before improving. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(k3): GPU-resident hot experts design (task #23 'trend to full GPU') The major GPU speedup: promote hot experts to persistent VRAM so decode's hot path is GPU-native (zero fetch, zero copy). 3 increments (copy-skip -> VRAM hot cache -> pipeline), the 32GB rate-distortion constraint (imatrix-enabled resident shrink frees VRAM for the hot set), measured per-increment via k3-bench. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(k3): flag the input_cpy-persistence question gating increment 1 vs 2 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(k3): modular rework-proof impl for #23 — reuse ResidencyCache + DeviceUploadFetcher Mechanism is the existing (buft,fetcher)-generic ResidencyCache; a VRAM cache = same class + device buft + host->device fetcher. 3 small parameterized pieces (DeviceUploadFetcher, instantiate w/ GGML_MOE_VRAM_CACHE_GB, seam hook). Stats only tune params -> zero mechanism rework. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(arch): MoE serving on a governed budget (draft; M5 owns the governor seam) Diagnoses the hardcoded-cache overcommit that collapsed K3 fetch bandwidth (40GB pinned + mmap = 95.9GB on 63GB -> pagefile thrash -> 205 MB/s -> 0.027 tok/s) and lays out the clean architecture: governor owns the residency budget net of the model's mmap footprint, plan-file is the one wire, ResidencyCache is pure mechanism. Governor-interface sections marked [M5 OWNS] for her to edit. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(arch): answer the [M5 OWNS] governor-budget seam in place (net-of-mmap is explicit arithmetic; plan_file.budget_bytes is the lease wire) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo * docs(arch): measured governed-budget inputs + graduated serving/load path Records the BigMama measurements feeding M5's #287 derivation (non-cache ~56GB, per-token working set 5.5GB, governed budget ~6GB, fetch recovers to 2.5GB/s at fit), the now-complete C++ cache mechanism (enable-from-plan, grow, shrink), and the three-piece graduated path to serving/load kimi-k3 (catalog row + serving-lane MoE launch + #287) replacing the rigged .bat. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * windows: make continuum-core build + link on windows-msvc (first time) start-server.sh now provisions the full Windows CUDA build env before the cargo builds (the cargo/nvcc path had none, unlike the vcvars-wrapped llama cmake): import MSVC via vswhere->VS2022-14.4x + a .bat env dump (cl.exe for nvcc), pin CMAKE to the manifest install, force CMAKE_GENERATOR=Ninja (the VS18-2026 auto- pick is undefined in cmake 3.30), add the Windows SDK bin (mt.exe/rc.exe), select a complete CUDA toolkit + CUDA_PATH (a provisioning split left cuda-env with 0 import libs vs cuda-13.2's 12), and RUSTFLAGS -L for pocket-tts (which emits no link-search) while re-carrying +crt-static so the /MT GPU stack still links. Portability: expert_container.rs + commands/capacity.rs used Unix-only std::os::unix::fs::FileExt::read_exact_at. Add crate::platform_io::pread_exact (unix read_exact_at / windows seek_read loop) - one place for positioned reads. Build validated (npm start exit 0, continuum-core lib clean). A separate runtime hot-loop on the #2088 core at startup is tracked apart from this build fix. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(capacity): device_fit VRAM-partition calc for the governor Pure calc the governor uses to fit a streaming-MoE's RESIDENT (non-expert) tier to a device VRAM budget and reconcile it with the expert tier on ONE budget — fixing the double-count where the expert pager was handed the full VRAM ceiling while resident silently ate most of it. Partition (in order): compute reserve -> resident (Native | device-fit Override | Unfittable) -> sufficient-context KV -> everything left = hot-expert VRAM budget (maximized: more on-GPU experts, fewer streams). Context is derived + clamped, never hand-picked. Artifact resolver injected (no hardcoded paths). Standalone-validated 7/7; M5 wires it into the daemon spawn path + launch (ServingTarget.resident_override) per the K3 sprint split. Refs #29 #31 #36. Arch-confirmed on real K3 UD-IQ2 (93 blk/896 exp/top-16). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(serving): resident-override plumbing on ServingTarget + launcher Wire foundation for the governor's device_fit plan: ServingTarget carries resident_override: Option<PathBuf>, and the launcher exports it as LLAMA_RESIDENT_OVERRIDE so llama.cpp sources the precision-shrunk RESIDENT (non-expert) tensors from the device-fit GGUF (all offloaded to GPU) while the primary streams experts. All builders updated; defaults None (resident serves as-shipped, no behavior change) until compute_resident_override + the resolve-or-generate resolver (#35) land next. In-crate validated. Refs #29 #36. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp fork to k3-adopt e3ce51df5 M5's per-layer KV accessors (n_head_kv_il + n_embd_head_{k,v}_il, continuum #238) + graph reconciliation. The K3 engine now builds against these — enables the device_fit resident-override serve + honest per-layer K3 KV sizing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(serving): compute_resident_override — wire device_fit into the plan The governor now DECIDES the resident source per serve: compute_resident_override derives resident_bytes (weights - expert_bytes_total) vs the governed VRAM ceiling via capacity::device_fit, and sets ServingTarget.resident_override. A dense/small model fits native (None); a >VRAM-resident MoE (K3) resolves a cached device-fit override that fits, else Unfittable → route to grid / generate (#35), glass-boxed. resolve_device_fit_override (model_registry::artifacts): looks up a per-user device-fit cache convention (<storage_root>/device-fit/<id>/) + a resident-bytes sidecar; returns the override only when its resident fits the usable budget. No hardcoded paths; generation/HF discovery is #35. The resident-fit decision turns only on resident_bytes vs budget — per-layer KV (#2107 ModelCapabilities) drives the context/expert split elsewhere, so KV is not consulted here. Refs #29 #35 #36. In-crate validated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * docs(arch): storage serving-tier governor — NVMe<->cold contention managed like VRAM/RAM Joel: 'like vram and memory, this contention has to be managed between cold storage and nvme.' Design: NVMe is a governed HOT-SERVING tier (a ResourcePool, same TrackedDir + evict_at_least machinery as CargoTargetPool), whose eviction = MIGRATE frozen/duplicate artifacts to the Cold drive, not a manual rm. Serving asks ensure_hot_resident(model); composes with device_fit's Unfittable one tier down (VRAM). Corrects the DriveRole bug: Cold (HDD) is FROZEN storage, never the per-token streaming tier (HDD = unservable). Dissolves today's K3 container disk fight: the C: IQ2 is a verified duplicate of the D: copy -> governor migrates it off NVMe -> container fits, no human deletes anything. Refs #12 #36. Design for M5's system_resources lane. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(capacity): verified cold-twin detection — safe-to-drop primitive for the storage tier The gate the NVMe serving-tier eviction (#302) consults before dropping a frozen GGUF: is an IDENTICAL twin already on cold storage? is_structural_twin (pure) = same shard count + per-shard name + size, zero-byte shards never match. scan_shards + find_cold_twin are the thin fs layer. Never drop an NVMe artifact without a VERIFIED cold twin (dropping 662GB on a path guess is the failure this guards). Standalone-validated 5/5. Composes with device_fit + M5's NvmeServingTierPool. Refs #12 #36. Design: STORAGE-SERVING-TIER-GOVERNOR.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp fork to k3-adopt c6469d5 — container-serve wired Both halves of the DirContainerFetcher wire (BigMama fetcher + moe_pick_fetcher branch 175ac9d6a; M5 caller-side encode + record_bytes reader c6469d5). Serving now reads the aligned per-layer container (GGML_MOE_CONTAINER) instead of the scattered raw GGUF — the honest ~2.6GB/s path. Retires the built-not-wired ContainerFetcher. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * fix(serving): resident_override on vision_sidecar ServingTarget + K3 coverage measurement Merge fix: vision_sidecar's ServingTarget was missing resident_override (added by #29). Plus a measurement test that drains the real K3 routed-access fixture through the #282 predictive instrument and prints repeat_recall / predicted_delta / schedulable_coverage — the go/no-go for the LiveUploadPager predictive pipeline (H2D/token = (1 - coverage) x ~11GB). Prints, never asserts (real routing sample). NOTE: can't run on windows-msvc (pre-existing cargo-test Unix-socket block, ipc/mod.rs); runs on M5's Mac. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(pager-driver): offline warm-coverage measurement (--synth-layers, --once, --budget-slots) The moe-pager-driver gains an offline replay mode so any completed GGML_MOE_TRACE_FILE can be scored on any box (windows-msvc clean by crate constraint), not just tailed live next to a serve: - --synth-layers N: synthesize the tkey->layer map from layer count alone (TkeyTable::for_layers, the same zero-config seam MoeTraceTail owns) instead of requiring an operator tkey-to-layer-matrix.json. - --once: exit when the trace stops growing (EOF) and print a SUMMARY line with mean DECODE-token serving hit = warm schedulable coverage. - --budget-slots N: override the predictor residency budget (default auto = first token x1.5) to measure the coverage-vs-free-VRAM curve (the device-fit tradeoff). Measured on BigMama run2.trace (302 warm decode tokens): bandit coverage 13.8% @250 slots -> 51.3% @2000 -> 65.7% @4024, beating naive last-N recency by +7-9pts at matched VRAM. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(pager): measure cross-layer prefetch predictor ceiling — DEAD lever for K3 VDD offline measurement (cooccur-ceiling bin) on the real warm serve trace (run2.trace, 122 held-out decode tokens, 11102 layer-steps): cross_layer_cooccur_hit 0.159 (adjacent-layer noisy-OR) recency_same_layer_hit 0.403 (last token, same layer) structure beyond recency -0.244 cooccur_recall_on_recency_misses 0.112 (11824/106058) Adjacent-layer co-occurrence predicts <half what plain recency does, and recovers only 11% of the experts recency misses (~base rate). K3 expert routing has no exploitable cross-layer structure — the CrossLayerExpert- Predictor prefetch lever is not worth wiring (saves the ggml pass-id capture slice). Recency-family residency (the bandit EMA curve) is THE signal; the only lever that lifts K3 is freeing VRAM (device-fit shrink) so residency coverage can reach the measured 51%. Caveat: adjacent-layer, one workload trace. Wider-predecessor noisy-OR would regress toward the frequency baseline (which underperforms recency), so a large lift is unlikely — but not measured here. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp to de29843e0 — device-resident expert cache half (#23) Pins the fork at the DeviceUploadFetcher wiring (my half of the LiveUpload- Pager H2D-kill). Off unless GGML_MOE_VRAM_CACHE_GB / plan device_budget_bytes enables it; host serving path byte-for-byte unchanged. M5's expert-loop D2D half lands next on the same seam. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp to 0fbe4e27a — quantize --resident-only + tier manifest (#40) Enables the device-fit division: produce a small resident override per precision tier + a (tier_label, resident_bytes) sidecar the governor reads to co-optimize the VRAM split. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * feat(pager): DivisionPolicy — the governor's VRAM-division RL brain (#2/#3) The second control rung above the pager's DecayBandit. The pager decides WHICH experts stay resident (reward=hit-rate, cheap, online). This decides HOW TO DIVIDE the card — resident (non-expert) weights vs expert cache — to MAXIMIZE tok/s. That reward (actual tok/s) is EXPENSIVE (a serve), so naive online RL flails; the fix is SIM-WARM-START: predict tok/s per division OFFLINE from the measured coverage curve, then a slow bandit refines each arm from real measured tok/s. - CoverageModel: piecewise-linear coverage(slots) over MEASURED points (k3_measured() = the trace-replay curve); saturates, never extrapolates up. - predict_tok_s: coverage -> (1-coverage)*experts/token*expert_bytes H2D -> t_token -> tok/s. Higher coverage -> less H2D -> faster (the load-bearing property, tested). - feasible_divisions: tier catalog (from --resident-only manifests) x HardwareBudget -> cache budget/slots per tier; drops VRAM-overflow tiers. - DivisionBandit: warm_start from the predictor; observe(tier, measured_tok_s) overrides the prior on first serve then EMAs — the expensive reward spent only on the arm actually run. Policy lives here (windows-clean, 4 tests pass); serving_daemon actuates it (M5's #2: discover manifests, feed catalog+budget+live tok/s, apply the chosen {resident_tier, device_budget_bytes} to the plan). Fractal control law: pager (experts<->hit-rate) -> this (VRAM split<->tok/s) -> grid. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp to 2a32025dd — device-cache un-crashable (clamp to free VRAM + null-buffer guard, #23) Testing convicted the segfault as VRAM oversubscription (K3 33GB resident + env cache on a 32GB card, cudaMalloc lazy-VMM deferred fault). Fix: clamp device budget to measured free VRAM (mine) + M5's D2D null-buffer guard. Device cache now disables safely where there's no room (K3) and works where there is (V4-Flash); can't crash from any budget source. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * test(division): bandit learns residency saturation from measured V4-Flash curve Feeds the real BigMama RTX 5090 --n-cpu-moe sweep (DeepSeek-V4-Flash UD-IQ2_M) into DivisionBandit: 0 resident=1.39, 8 resident=1.69, 14 resident=1.68 tok/s. Asserts the bandit converges on the SATURATION KNEE (8 layers), not max residency — 8->14 layers buys nothing at +11GB VRAM. Encodes the measured finding that the governor must learn 'minimal static residency + max device cache', the freed VRAM belonging to the recency cache (#43), not to over-pinned static layers. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * test(division): bandit finds non-monotonic device-cache budget optimum Measured V4-Flash device-cache coverage curve (5090, GGML_MOE_VRAM_CACHE_GB sweep): 6GB/992slots=1.80, 12GB/1985=3.10, 22GB/3630=2.96 tok/s, all 100% hit. tok/s is NON-MONOTONIC in budget: undersized churns, 12GB is the plateau knee, 22GB is no better (100% hit but O(slots) reserve_slot eviction scan). predict_tok_s's monotonic prior would pick 22GB; only the measured reward lands on 12GB — which frees ~20GB of a 32GB card for co-resident lanes. Pins the invariant that the governor must not oversize the cache and starve other models. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc * chore(vendor): bump llama.cpp to fa7e0d8e9 — #43 device-cache fix + async restore + enum fix Includes the prefetch host_visible guard (THE #43 crash fix, validated 3.05 tok/s V4-Flash device cache on the 5090), M5's async cpy_tensor_async restore, and the moe-pack quant-enum fix. A fresh parent build now includes the un-crashable device cache instead of the pre-fix pin. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bumps @typescript-eslint/eslint-plugin from 8.29.1 to 8.31.0.
Release notes
Sourced from
@typescript-eslint/eslint-plugin's releases.Changelog
Sourced from
@typescript-eslint/eslint-plugin's changelog.Commits
2cc7656chore(release): publish 8.31.080bd7a5feat(eslint-plugin): [no-unnecessary-type-assertion] add option to ignore str...1a3ab0dchore(eslint-plugin): migrate to vitest (#10579)9531492chore(release): publish 8.30.1152def7fix(eslint-plugin): fix mistake with eslintrc config generation (#11072)b3688bechore(release): publish 8.30.03ccd79cfeat(eslint-plugin): [no-explicit-any] suggest to replace keyof any with Prop...128d95bfix(eslint-plugin): [promise-function-async] use a different error message fo...69e2f6cfeat: support stringly-typed extends (#10973)Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting
@dependabot rebase.Dependabot commands and options
You can trigger Dependabot actions by commenting on this PR:
@dependabot rebasewill rebase this PR@dependabot recreatewill recreate this PR, overwriting any edits that have been made to it@dependabot mergewill merge this PR after your CI passes on it@dependabot squash and mergewill squash and merge this PR after your CI passes on it@dependabot cancel mergewill cancel a previously requested merge and block automerging@dependabot reopenwill reopen this PR if it is closed@dependabot closewill close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually@dependabot show <dependency name> ignore conditionswill show all of the ignore conditions of the specified dependency@dependabot ignore this major versionwill close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this minor versionwill close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this dependencywill close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)