Skip to content
github-actions[bot] edited this page Sep 30, 2026 · 2 revisions

Speed

Latency is the reason a yes/no model exists: a decision that takes 26 ms can sit inside a request, a webhook or a loop. These pages explain where the time goes in a language model, why a model that generates nothing is fast, and how to measure it without fooling yourself. The measured numbers come from one laptop, an Intel Core Ultra 7 255H with 16 threads.

Guides

Measurements

Comparisons

Speed

Probability and thresholds

Question design

Use cases

Evaluation

Agents and routing

Integrations

Local and private AI

llama.cpp and GGUF

Clone this wiki locally