v0.14.0 — Assistants · MCP · GPU backend · DS V4 Flash (104 GB) measured
🤖 P7 — Assistants, MCP, GPU backend & tool calling
- LlamaServerBackend (GPU): offload via
-ngl/--n-cpu-moe, reasoning mode, date injection, subprocess page-fault telemetry (Stats ตัวจริงสำหรับ GPU path) - Assistants: CRUD API + console page; assistant references guard hub delete/clear
- MCP host: จัดการ stdio/SSE MCP servers + list/call tools
- Tool calling protocol (
tools/tool_calls) - GPU load options:
gpu_layers/kv_cache_type(ModelLoadRequest + Settings) + quant advisor
📡 EXP-012 — DeepSeek-V4-Flash 0731 (104 GB) measured honestly
- 1.48–1.89 tok/s บน i9-9900KF + RTX 3060 12 GB + 64 GB RAM — disk-bound (36–77k faults/token ≈ 150–300 MB disk/token)
- Full download + measure harness; hub รองรับ sharded/subdir/Xet + resumable
.part+ GGUF structural gate - Qwen3.6-35B-A3B IQ1_M: 75.9 tok/s (GPU-bound) เป็น anchor
🔬 EXP-009…EXP-013
- KV-q8 no-op, spec-decode dead end, clean-room gate, IQ1_M 72–78 tok/s + Thai tonal quality eval, kimi-k3-in-c deep-research
🏠 Repo & CI
- Project ย้ายมาที่ repo root (สะอาด — ไม่มีงานอื่นปน) + GitHub Actions CI เขียว (Python Windows + frontend)
- Packaging fixes (deps ที่ import แต่ไม่เคยประกาศ) + flake fixes
รายละเอียดเต็ม: CHANGELOG.md