Skip to content

History

Revisions

  • Add GLM-4.7-Flash llama.cpp guide; refresh the 4-GPU layout New guide covers the 31B MoE reasoning model on 2x RTX 3080: why the vLLM AWQ route dies mid-inference, why --kv-cache-dtype fp8 cannot work on Ampere, how MLA makes 64K context far cheaper than ordinary attention, and the provider/backend concurrency mismatch that got the model auto-disabled. The Qwen guide's box layout was stale: GPUs 0-1 now run GLM, not Cydonia.

    @flyworker flyworker committed Sep 15, 2026
    f04d46e
  • Qwen3.8-27B guide: two slots, n-gram drafting, correct declared window The running config moved on from what this page documented: - --parallel 2, not 1. One slot plus a long request starved the health check into a timeout, which deregistered the model, which stopped the traffic, which made it healthy again — a loop. Documented as a gotcha. - Declared context is one slot (32768), not the -c total (65536). No single request can ever have the total, and the server 400s anything over a slot. - --spec-type ngram-mod, with the measured numbers: 31.7 -> 198.9 tok/s reproducing a document already in context, chat unchanged. Notes that -lcd is accepted but inert on this build. - Prefill measured at ~1,750 tok/s on a 30,021-token prompt (17 s), so the long-prompt cost is generation, not prefill. Also drops a link into the private tracker and adds the shell-quoting trap that restart-loops the container.

    flyworker committed Sep 6, 2026
    819c1a9
  • docs: benchmark Qwen3.6-35B-A3B FreeToken 1xRTX3080 vs llama.cpp 2xRTX3080 (32K-64K), with chart

    flyworker committed Aug 26, 2026
    5cfe35a
  • docs: list Qwen3.8-27B llama.cpp and FreeToken guides on Home

    flyworker committed Aug 26, 2026
    d8b764c
  • docs: add FreeToken guide (Qwen3.8-27B / DeepSeek-V4-Flash limits, Qwen3.6-35B-A3B on one RTX 3080)

    flyworker committed Aug 26, 2026
    fe114a3
  • docs: add Qwen3.8-27B llama.cpp 2x RTX 3080 guide, Cydonia 2-GPU variant

    flyworker committed Aug 26, 2026
    42f2d84
  • docs: add guide for serving Cydonia 24B v4.3 (AWQ) on 4x RTX 3080

    @flyworker flyworker committed Jul 24, 2026
    0667ba1
  • docs: point CLIProxyAPI to upstream router-for-me, drop stale ZK/UBI/account content, add reliability section

    charles committed Jul 12, 2026
    3e2543b
  • docs: add local_model note and docker group crash to CLIProxyAPI guide

    flyworker committed Jun 27, 2026
    3cf61e7
  • Add Configuration, Inference Mode, WebSocket Protocol, Troubleshooting pages

    flyworker committed Jun 25, 2026
    69d2e22
  • Add wiki structure: Home, Quick Start, Guides, CLIProxyAPI guide

    flyworker committed Jun 25, 2026
    84a232b
  • Initial Home page

    Charles Cao committed Jun 25, 2026
    c642549