Skip to content

v0.14.0 — Assistants · MCP · GPU backend · DS V4 Flash (104 GB) measured

Choose a tag to compare

@mrDedchai mrDedchai released this 10 Aug 09:25
· 130 commits to main since this release

🤖 P7 — Assistants, MCP, GPU backend & tool calling

  • LlamaServerBackend (GPU): offload via -ngl/--n-cpu-moe, reasoning mode, date injection, subprocess page-fault telemetry (Stats ตัวจริงสำหรับ GPU path)
  • Assistants: CRUD API + console page; assistant references guard hub delete/clear
  • MCP host: จัดการ stdio/SSE MCP servers + list/call tools
  • Tool calling protocol (tools/tool_calls)
  • GPU load options: gpu_layers/kv_cache_type (ModelLoadRequest + Settings) + quant advisor

📡 EXP-012 — DeepSeek-V4-Flash 0731 (104 GB) measured honestly

  • 1.48–1.89 tok/s บน i9-9900KF + RTX 3060 12 GB + 64 GB RAM — disk-bound (36–77k faults/token ≈ 150–300 MB disk/token)
  • Full download + measure harness; hub รองรับ sharded/subdir/Xet + resumable .part + GGUF structural gate
  • Qwen3.6-35B-A3B IQ1_M: 75.9 tok/s (GPU-bound) เป็น anchor

🔬 EXP-009…EXP-013

  • KV-q8 no-op, spec-decode dead end, clean-room gate, IQ1_M 72–78 tok/s + Thai tonal quality eval, kimi-k3-in-c deep-research

🏠 Repo & CI

  • Project ย้ายมาที่ repo root (สะอาด — ไม่มีงานอื่นปน) + GitHub Actions CI เขียว (Python Windows + frontend)
  • Packaging fixes (deps ที่ import แต่ไม่เคยประกาศ) + flake fixes

รายละเอียดเต็ม: CHANGELOG.md