Skip to content

v0.76.0-rc8

Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 29 Aug 09:46
· 113 commits to main since this release

What's Changed

  • Fix crates.io publishing for new workspace crates, and make dev builds identify themselves by @michaelneale in #1223
  • Split Skippy by functional boundary by @i386 in #1194
  • feat: pass reasoning effort through Skippy templates by @i386 in #1211
  • chore: update pinned llama.cpp revision by @i386 in #1216
  • fix: IPC idle timeout for external-relay plugins by @MahdiHedhli in #1191
  • Publish Rust crate API docs with the website by @i386 in #1230
  • feat(mesh): expose model-bound tokenizer capability by @i386 in #1227
  • fix: repair the Fly console image build by @michaelneale in #1241
  • chore(llama): regenerate patch queue for latest upstream by @i386 in #1232
  • feat(catalog): add Muse Glimmer 30B to the model catalog by @michaelneale in #1240
  • feat(bench): run SWE-Gym through Harbor by @i386 in #1239
  • fix: Skippy Nemotron MTP loading by @ndizazzo in #1245
  • fix: complete remaining Skippy MTP runtime callers by @i386 in #1247
  • chore(deps): pin iroh to 1.0.3 explicitly by @michaelneale in #1257
  • feat(logging): add durable lifecycle logging foundation by @ndizazzo in #1174
  • feat(logging): expose audited local log APIs and runtime telemetry by @ndizazzo in #1175
  • feat(console): add typed logs ledger and maintenance UI by @ndizazzo in #1176
  • feat(openai): integrate exact lifecycle and trusted sessions by @ndizazzo in #1258
  • feat: make plugin-only Serve mode nodes discoverable/reachable by other mesh peers by @MahdiHedhli in #1192
  • Fix install one-liner 404 on NVIDIA hosts without a CUDA toolkit by @michaelneale in #1267
  • ci: reshape PR and main CI around composable slices, prepping for Depot workers by @ndizazzo in #1244
  • ci: distinguish skipped bootstrap routes by @ndizazzo in #1272
  • ci: fail fast on topic lane projections by @ndizazzo in #1274
  • ci: flatten paginated lane checks by @ndizazzo in #1276
  • Retry transient model route misses in console chat by @i386 in #1277
  • fix(ci): make CUDA release checks hermetic by @ndizazzo in #1270
  • fix: CI smoke matrix routing by @ndizazzo in #1284
  • fix(moa): restore single-model degrade for model=mesh on host ingress by @michaelneale in #1291
  • ci: split planning and platform workflow graphs by @ndizazzo in #1282
  • fix(ci): keep main planning exhaustive by @ndizazzo in #1295
  • fix(cache): certify Muse-Glimmer family for the resident-KV prefix cache by @michaelneale in #1294
  • Fix split serving: stages above layer 0 fail to load by @michaelneale in #1290
  • fix(logging): harden lifecycle capture and console recovery by @ndizazzo in #1268
  • chore: update mesh-llm-ui console dependencies by @ndizazzo in #1286
  • fix: move model lifecycle controls off peer-reachable ingress by @ndizazzo in #1279
  • fix: name a local GGUF with --model instead of trying to resolve it by @michaelneale in #1269
  • fix: cross-platform product bundle tree hashes by @michaelneale in #1271
  • Keep degraded MoA requests on the resolved pipeline model by @michaelneale in #1292
  • ci: prepare protected Depot PR runners by @ndizazzo in #1302
  • ci: add bounded Depot PR canary gate by @ndizazzo in #1306
  • ci: isolate native caches on Depot PR runners by @ndizazzo in #1310
  • ci: keep Depot cache namespace inert by @ndizazzo in #1317
  • CI: pin Depot audit to merged policy by @ndizazzo in #1319
  • Durable KV prefix cache: agent prefixes survive eviction, restart, and cold nodes by @michaelneale in #1228
  • feat(topology): route Qwen3.8 identities to the qwen35 recurrent family by @michaelneale in #1283
  • Make the optional GitHub star step reliable in setup by @michaelneale in #1278
  • Fix/windows skippy package large gguf by @wangwenjunfromlanzhou in #1307
  • CI: identify failing Depot canary endpoint by @ndizazzo in #1321
  • CI: classify rejected Depot cache endpoint by @ndizazzo in #1323
  • CI: add protected Depot authority sentinel by @ndizazzo in #1324
  • CI: prove Depot sentinel cache writes by @ndizazzo in #1326
  • Docs: record unsafe Depot PR authority by @ndizazzo in #1328
  • Docs: define Depot PR cache isolation contract by @ndizazzo in #1329
  • Allow bounded Depot caching for approved PRs by @ndizazzo in #1333
  • Guard approved Depot PR routing by @ndizazzo in #1334
  • ci: repin Depot PR isolation audit by @ndizazzo in #1336
  • ci: fingerprint Depot macOS toolchains by @ndizazzo in #1337
  • chore: sync source versions during GitHub releases by @ndizazzo in #1330
  • docs: record Depot PR rollout evidence by @ndizazzo in #1341
  • fix(skippy): restore recurrent shared prefixes by @michaelneale in #1342
  • feat(catalog): recommend Qwen3.8 27B Q4_K_M by @michaelneale in #1343
  • fix(moa): preserve caller tool schemas by @michaelneale in #1344
  • fix(split): let Iroh own transport path selection by @michaelneale in #1345
  • Consolidate model=auto and model=mesh into one automatic-routing directive by @michaelneale in #1309
  • Return a visible error when streaming prompts exceed context by @michaelneale in #1352
  • Treat max_tokens as a ceiling, not a context reservation by @michaelneale in #1354
  • Stop aborting healthy long prefills at five minutes by @michaelneale in #1355
  • Return IDs for OpenAI tool calls on local-model-only serving by @michaelneale in #1318
  • Cap remotely advertised model lists to protect the console from a hostile peer by @michaelneale in #1357
  • Reject non-protobuf content types on OTLP/HTTP ingestion by @michaelneale in #1356
  • Add mode-aware management health endpoint by @i386 in #1351
  • fix(skippy): serialize native proposal lifecycle by @i386 in #1242
  • CI recognizes Apple provider paths by @i386 in #1363
  • fix(ui): remove dead ungated ErrorBoundary that leaks stack traces (#39) by @michaelneale in #1361
  • Reuse recurrent KV across growing chat turns by @i386 in #1253
  • Fix blank responses from thinking models by @michaelneale in #1365
  • perf(skippy): coalesce native KV page transfers by @michaelneale in #1368
  • task: refine logging console UX and live delivery by @ndizazzo in #1339
  • Stream tool calls as they are generated instead of after the turn ends by @michaelneale in #1371
  • fix(skippy): stop a stalled SSE consumer pinning a generation worker by @ndizazzo in #1367
  • fix(mtp): make native MTP actually work on single-node serving by @michaelneale in #1366
  • fix: keep --log-format json parseable and block new raw console prints in CI by @ndizazzo in #1376
  • fix(xtask): regenerate console-print ratchet drifted 2 lines by #1376 by @ndizazzo in #1379
  • Wire Playwright into CI and fix the 9 pre-existing /logs failures by @ndizazzo in #1377
  • fix(ci): select static-abi in pr-draft when its matrix consumers are planned by @ndizazzo in #1383
  • test(ui): tolerate hosted-runner frame cadence by @ndizazzo in #1386
  • perf(kv): keep exact-state recording off inference path by @michaelneale in #1389
  • ci: cancel sibling PR lanes after failure by @ndizazzo in #1388
  • fix(tui): make the dashboard immune to stray console output by @ndizazzo in #1382
  • ci: bound compiler and local build caches by @ndizazzo in #1390
  • Revert "ci: bound compiler and local build caches (#1390)" by @ndizazzo in #1394
  • ci: prebuilt runner images for smoke/SDK/web jobs by @ndizazzo in #1380
  • ci: bound compiler and local build caches by @ndizazzo in #1395
  • fix(moa): prevent reasoning-only model poisoning by @michaelneale in #1391
  • docs(skippy): remove stale disk cache defaults by @michaelneale in #1398
  • Stop persisting KV cache state to disk by @michaelneale in #1399
  • Fix release builds failing in the containerized smoke job by @michaelneale in #1401
  • fix(ci): point pnpm at the runner image's baked store instead of the Actions cache by @ndizazzo in #1400
  • refactor: complete evidence-backed codebase audit waves by @ndizazzo in #1396
  • ci(quality): install just before the CI-contracts python tests run by @ndizazzo in #1407
  • fix: close the four rc6 release-validation defects by @ndizazzo in #1405
  • Advance llama.cpp and simplify the Skippy patch queue by @i386 in #1402
  • skippy: address #1402 review follow-up (dead comparison + Inkling test coverage) by @i386 in #1412
  • fix(api): stop leaking plugin endpoint credentials over the management API by @michaelneale in #1364
  • fix(build): link libvendor-hash.a and test Rust on llama.cpp pin bumps by @michaelneale in #1419
  • Rescue terminal Buzz replies in MoA by @michaelneale in #1417
  • Compact oversized chat context for provider windows by @i386 in #1280
  • Unblock release dispatches after the config schema module split by @michaelneale in #1426
  • fix(packaging): retry transient HF artifact uploads by @michaelneale in #1424
  • Fix chat-chain pipelined splits: verify-history accounting, early-stop replies, session leaks by @danielwinterw in #1408
  • ci: skip build slices for draft pull requests by @ndizazzo in #1415
  • fix(mesh): close the iroh endpoint on shutdown instead of dropping it by @ndizazzo in #1431
  • fix: redact plugin provider health details by @ndizazzo in #1432
  • fix(release): compare release bases, not raw tag SHAs, in provenance check by @ndizazzo in #1430
  • fix: drain hosted split stages on shutdown by @ndizazzo in #1433
  • fix: bound downstream endpoint resolution by @ndizazzo in #1438
  • feat: iteration-level scheduler for concurrent staged serving (#1416) by @i386 in #1420
  • task: make operational logs durable, attributable, and queryable by @ndizazzo in #1440
  • task: add lifecycle timelines and recovery to the logs console by @ndizazzo in #1441
  • feat(skippy): replace flat prefix caches with unified radix by @i386 in #1429
  • perf(skippy): add scheduler lab and cache-aware radix scheduling by @i386 in #1447
  • test(skippy): pin scheduler workload fixtures by @i386 in #1450
  • feat(skippy): make resident KV admission capacity aware by @i386 in #1452
  • ci(llama): tiered family battery + agent-repaired canary on the family-certify runner by @i386 in #1436
  • Run the llama.cpp canary daily and accept split GGUF metadata by @i386 in #1457
  • fix(skippy): restore small-context admission and state handoff by @i386 in #1459
  • fix(proxy): cancel upstream before response headers by @ndizazzo in #1448
  • ci: assign imported Just files to Rust ownership by @ndizazzo in #1464
  • refactor(just): split root Justfile into semantic imports by @ndizazzo in #1458
  • fix(ci): keep change planning below expression limit by @ndizazzo in #1468
  • Send remote mesh inference directly into the serving pipeline by @michaelneale in #1463
  • Clarify end-to-end encryption for mesh traffic by @michaelneale in #1470
  • Fix llama canary family certification failures by @i386 in #1471
  • skippy-quantize: compose-mtp — splice an MTP draft into a sharded target GGUF by @i386 in #1439
  • Route repeated prompts to workers with verified radix cache reuse by @i386 in #1449
  • fix(skippy): use raw f32 activation transport by @i386 in #1482
  • ci(llama): wire the agent repair loop into the upstream canary by @i386 in #1481
  • ci(llama): heartbeat, arm64 build guard, and push-permission diagnostics for the canary repair loop by @i386 in #1489
  • ci(llama): preflight repair-token push permission before any repair work by @i386 in #1493
  • ci(llama): probe actual token write capability in the repair preflight by @i386 in #1496
  • ci(llama): prune stale /tmp worktrees and fix the preflight probe delete by @i386 in #1499
  • ci(llama): auto-approve agent tool permissions in the repair sandbox by @i386 in #1502
  • fix(llama): rebase patch queue onto upstream b19cbe925b by @i386 in #1504
  • ci(llama): ensure the repair PR exists before applying the agent's body by @i386 in #1506
  • fix(llama): rebase patch queue onto upstream b19cbe925b by @i386 in #1505
  • fix(ci): validate CUDA runtime toolkit labels by @ndizazzo in #1465
  • fix(ui): render LaTeX in active chat responses by @ndizazzo in #1466
  • ci(llama): review certified canary repairs with a fresh agent turn by @i386 in #1507
  • skippy: support staged Qwen3.8 Flash-Next inference by @michaelneale in #1509
  • Certify Qwen3.8 Flash-Next in the llama family battery by @i386 in #1513
  • fix(skippy): address Qwen4 split framing review findings by @i386 in #1516
  • fix(release): use dotted workspace version form in skippy-scheduler by @michaelneale in #1518
  • ci: regenerate console-print ratchet after #1516 by @i386 in #1525

New Contributors

Full Changelog: v0.75.1...v0.76.0-rc8