-
Notifications
You must be signed in to change notification settings - Fork 2
Guides
Step-by-step guides for specific computing-provider use cases.
Turn a ChatGPT account into a Swan Chain computing provider — no GPU, no model weights, no code changes. Just two config files.
Use when: you want frontier GPT-5.x models available in the marketplace for development, benchmarking, or low-volume production without GPU infrastructure.
Serve a 24B model on four consumer 10 GB GPUs with vLLM tensor parallelism, AWQ 4-bit weights, and an fp8 KV cache — including the full 131K context window.
Use when: you have modest multi-GPU consumer hardware and want to serve a large model that won't fit in fp16, with the real VRAM/context/concurrency tuning worked out.
Serve a 27B dense vision-language model on two consumer 10 GB GPUs with llama.cpp and Unsloth's dynamic 4-bit GGUF — 64K context, image input, thinking off by default — and share the box with a vLLM model on the other GPUs. Includes the KV-cache and CPU-starvation traps that silently kill performance, and a FreeToken note for MoE models.
Use when: the model you want has no vLLM-compatible 4-bit checkpoint that fits your VRAM, or you need to split a multi-GPU box between two engines.
Serve a 31B MoE reasoning model on two consumer 10 GB GPUs at ~120 tok/s with a 64K context. Covers why the vLLM AWQ route dies mid-inference, why --kv-cache-dtype fp8 cannot work on Ampere, how MLA makes long context far cheaper than it looks, and the concurrency mismatch that got the model auto-disabled off the marketplace.
Use when: serving a reasoning or MoE model on modest GPUs, or when a vLLM checkpoint loads fine and then crashes under real traffic.
Why FreeToken can't serve Qwen3.8-27B (or DeepSeek-V4-Flash) on a 10 GB card, and the MoE model it does serve at ~40 tok/s on a single RTX 3080 — with the CUDA-devel container setup, ft serve flags, provider models.json, and measured numbers.
Use when: you want to run a MoE model larger than your VRAM by keeping experts in host RAM, or need to know up front whether a checkpoint's non-expert weights will fit your GPU.
Same model, same 32K–64K prompts: prefill, decode, TTFT, end-to-end throughput, VRAM, host RAM/CPU, KV capacity, cold start and cost per million tokens, with a chart. Two cards with everything in VRAM is ~2.4× faster; FreeToken does it on one card for ~15 % more GPU-time per token.
Use when: deciding whether to spend one GPU or two on a MoE model.