Skip to content

History / Home

Revisions

  • Add GLM-4.7-Flash llama.cpp guide; refresh the 4-GPU layout New guide covers the 31B MoE reasoning model on 2x RTX 3080: why the vLLM AWQ route dies mid-inference, why --kv-cache-dtype fp8 cannot work on Ampere, how MLA makes 64K context far cheaper than ordinary attention, and the provider/backend concurrency mismatch that got the model auto-disabled. The Qwen guide's box layout was stale: GPUs 0-1 now run GLM, not Cydonia.

    @flyworker flyworker committed Sep 15, 2026
  • docs: benchmark Qwen3.6-35B-A3B FreeToken 1xRTX3080 vs llama.cpp 2xRTX3080 (32K-64K), with chart

    flyworker committed Aug 26, 2026
  • docs: list Qwen3.8-27B llama.cpp and FreeToken guides on Home

    flyworker committed Aug 26, 2026
  • docs: add guide for serving Cydonia 24B v4.3 (AWQ) on 4x RTX 3080

    @flyworker flyworker committed Jul 24, 2026
  • Add wiki structure: Home, Quick Start, Guides, CLIProxyAPI guide

    flyworker committed Jun 25, 2026
  • Initial Home page

    Charles Cao committed Jun 25, 2026