Releases: JediKinght9527/precision-bench
Releases · JediKinght9527/precision-bench
Release list
v0.3.0 — first public release
First public release. Precision Bench is a local dashboard for testing third-party LLM relay endpoints: it drives real traffic at a channel, measures latency from the SSE stream, normalizes the provider's usage fields, and compares runs against a stored baseline.
What it does
- Load testing — closed-loop, open-loop with coordinated-omission correction, and duration profiles. TTFT / TPOT / ITL / E2E percentiles to P99.9, goodput, throughput, and cost split across input / output / cache-read / cache-write.
- Prompt-cache verification — usage fields normalized on every run (OpenAI, Anthropic, DeepSeek, Kimi, Qwen, GLM, Gemini), plus an active fixed-prefix probe that compares hit vs miss TTFT and reports effective / suspected fake / not reported.
- Degradation detection — eight scored dimensions (math, knowledge, Chinese, instruction following, JSON, code, long-context needle, self-consistency) with per-channel baseline and output-fingerprint comparison.
Numbers from the maintainer's own runs
112 runs / 1826 recorded samples against real channels, 11 cache probes, 8 degradation runs. Representative results (median E2E P95 across runs):
| Channel | Reqs | Success | E2E P95 | Output tok/s | Cache hit |
|---|---|---|---|---|---|
stealth/space-bunny-alpha (OpenRouter) |
100 | 100% | 6.13 s | 31.8 | 89.8% |
kimi-k3 |
43 | 93% | 17.9 s | 33.3 | 49.0% |
z-ai/glm-5.3 |
100 | 31% | 10.6 s | 9.8 | not reported |
The z-ai/glm-5.3 row is why the tool exists: a channel can answer 100 requests and still fail two thirds of them under concurrency, and it reports no cache usage at all.
Running it
./run.sh # dashboard on http://127.0.0.1:8787
docker compose up -d # or DockerSingle process only — live run state lives in memory. Do not bind HOST=0.0.0.0 to the public internet.
Known limitations
- The dashboard UI is Chinese-only; the English README is the primary entry point for documentation.
- Cache metrics are reported as
n/awhen the upstream does not report cache usage. They are never converted to 0%. - Degradation dimensions use a self-built curated subset (GSM8K/MMLU/C-Eval/IFEval-style), not the official datasets, so scores are not directly comparable to published leaderboards. lm-eval tasks are available for scoreable comparisons.
- The active cache probe is serial by design; long prefixes take proportionally longer.
Verification
- 88 automated tests, 63.9% statement coverage (gate 55%), CI green on the release commit.
- Docker image builds in CI on every push.