Skip to content

Releases: i-mrDed/weight-streaming

v0.15.0 — security hardening + auth + honest frontend

Choose a tag to compare

@mrDedchai mrDedchai released this 10 Aug 13:54

v0.15.0 — security hardening + auth + honest frontend

Updated 2026-08-13 — tag moved to current main (the release that goes
public). This build adds auto-tiering, the honest bench harness, hub
recommendations and 30+ research experiments on top of the original
security release. Changelog: CHANGELOG.md.

This release came out of a 4-platform public code review. The headline is
security: five verified High-severity bugs are fixed, and the API is
now optionally protected by a token. It also ships the console inside the
wheel (previously pip install produced a server with no web UI).

🔐 Security hardening (OpenCode W1–W5, verified)

  • W1 deadlock fix — model eviction no longer runs while holding the
    load lock (server froze permanently at max_loaded_models).
  • W2 CORS hardening — API is loopback-only by default (WS_CORS_ORIGINS
    to extend); any-website drive-by was possible before.
  • W3 MCP RCE — MCP server command must now be a bare allowlisted
    runner (npx, python, uvx, …) instead of arbitrary shell.
  • W4 mmap leak — failed GGUF metadata parse no longer leaks the

    100 GB mapping + fd.

  • W5 path traversalassistant_id/issue_id validated
    ^[A-Za-z0-9_.-]+$ at the store level (Windows %5C escape closed).
  • 11 regression tests, test-first against the pre-fix code.

🔐 API auth (B1) + frontend tests (B2) + thinking-marker fix (B3)

  • WS_API_TOKEN — Bearer-token auth for /v1/* (console stores + attaches
    automatically). 20 frontend tests (thinks 7 + api 13). Thinking/answer
    markers now only split at line boundaries.

⚡ Auto-tiering — route requests to the right model

  • User-configurable fast/quality tiers, per-tier max_tokens budget
    (fast 2048 / quality 8192 — kills the temp-0 repetition loop), live
    stats card, history export, /v1/tiering/route returning the tier
    budget, and /v1/models/load now accepts extra_args (flags were
    silently dropped before).

⚡ Daily driver upgraded — Gemma 4 12B QAT+MTP (EXP-022) + 26B (EXP-019)

  • Thai gate 9/9 + tonal 6/6 PERFECT — first model on this rig to pass the
    tonal discriminator at usable speed. MTP works on Gemma (+20%) unlike DS
    V4's embedded MTP (EXP-015). Qwen3.6 IQ2_M demoted.

🌐 Hub — "Proven on this rig" recommendations

  • /v1/hub/recommended with localisation + in-app evidence; hardened
    against bad recommendations.

⚡ Bench harness — honest-measurement core

  • weight_stream/bench/ — clean-room methodology, cmdline verify,
    cold/warm runs, --sweep-threads, Gemma 4 channel-think support,
    docs/BENCHMARKING.md official methodology.

🧪 Research — 30 experiments + a full paper (this release's payload)

  • Physics-calibrated simulator (EXP-025), streaming buffer abstraction
    (EXP-026), real MoE compute/I/O ratio (EXP-027), Phase 4 evaluation
    metrics + real Qwen benchmark (EXP-028), K3 >RAM simulation (EXP-029),
    expert offloading on fits-VRAM (EXP-030), 104 GB model on 64 GB RAM
    measured honestly (EXP-012).
  • Full draft paper: research/paper/paper.md, auto fact-checked
    (scripts/factcheck_paper.py — 33/33 PASS).

🧪 Full suite

  • 471 passed / 7 skipped Python tests · mypy clean (60 files) ·
    20 frontend tests · CI green on Windows Python 3.11–3.13.

Install / run — see README. Requires llama.cpp's
llama-server (or point the API at your own llama-server).

v0.14.0 — Assistants · MCP · GPU backend · DS V4 Flash (104 GB) measured

Choose a tag to compare

@mrDedchai mrDedchai released this 10 Aug 09:25

🤖 P7 — Assistants, MCP, GPU backend & tool calling

  • LlamaServerBackend (GPU): offload via -ngl/--n-cpu-moe, reasoning mode, date injection, subprocess page-fault telemetry (Stats ตัวจริงสำหรับ GPU path)
  • Assistants: CRUD API + console page; assistant references guard hub delete/clear
  • MCP host: จัดการ stdio/SSE MCP servers + list/call tools
  • Tool calling protocol (tools/tool_calls)
  • GPU load options: gpu_layers/kv_cache_type (ModelLoadRequest + Settings) + quant advisor

📡 EXP-012 — DeepSeek-V4-Flash 0731 (104 GB) measured honestly

  • 1.48–1.89 tok/s บน i9-9900KF + RTX 3060 12 GB + 64 GB RAM — disk-bound (36–77k faults/token ≈ 150–300 MB disk/token)
  • Full download + measure harness; hub รองรับ sharded/subdir/Xet + resumable .part + GGUF structural gate
  • Qwen3.6-35B-A3B IQ1_M: 75.9 tok/s (GPU-bound) เป็น anchor

🔬 EXP-009…EXP-013

  • KV-q8 no-op, spec-decode dead end, clean-room gate, IQ1_M 72–78 tok/s + Thai tonal quality eval, kimi-k3-in-c deep-research

🏠 Repo & CI

  • Project ย้ายมาที่ repo root (สะอาด — ไม่มีงานอื่นปน) + GitHub Actions CI เขียว (Python Windows + frontend)
  • Packaging fixes (deps ที่ import แต่ไม่เคยประกาศ) + flake fixes

รายละเอียดเต็ม: CHANGELOG.md