Add GLM-4.7-Flash llama.cpp guide; refresh the 4-GPU layout New guide covers the 31B MoE reasoning model on 2x RTX 3080: why the vLLM AWQ route dies mid-inference, why --kv-cache-dtype fp8 cannot work on Ampere, how MLA makes 64K context far cheaper than ordinary attention, and the provider/backend concurrency mismatch that got the model auto-disabled. The Qwen guide's box layout was stale: GPUs 0-1 now run GLM, not Cydonia.
docs: benchmark Qwen3.6-35B-A3B FreeToken 1xRTX3080 vs llama.cpp 2xRTX3080 (32K-64K), with chart
docs: add FreeToken guide (Qwen3.8-27B / DeepSeek-V4-Flash limits, Qwen3.6-35B-A3B on one RTX 3080)
docs: add Qwen3.8-27B llama.cpp 2x RTX 3080 guide, Cydonia 2-GPU variant
docs: add guide for serving Cydonia 24B v4.3 (AWQ) on 4x RTX 3080
Add wiki structure: Home, Quick Start, Guides, CLIProxyAPI guide