Skip to content

Releases: JediKinght9527/precision-bench

v0.3.0 — first public release

Choose a tag to compare

@JediKinght9527 JediKinght9527 released this 26 Sep 02:08

First public release. Precision Bench is a local dashboard for testing third-party LLM relay endpoints: it drives real traffic at a channel, measures latency from the SSE stream, normalizes the provider's usage fields, and compares runs against a stored baseline.

What it does

  • Load testing — closed-loop, open-loop with coordinated-omission correction, and duration profiles. TTFT / TPOT / ITL / E2E percentiles to P99.9, goodput, throughput, and cost split across input / output / cache-read / cache-write.
  • Prompt-cache verification — usage fields normalized on every run (OpenAI, Anthropic, DeepSeek, Kimi, Qwen, GLM, Gemini), plus an active fixed-prefix probe that compares hit vs miss TTFT and reports effective / suspected fake / not reported.
  • Degradation detection — eight scored dimensions (math, knowledge, Chinese, instruction following, JSON, code, long-context needle, self-consistency) with per-channel baseline and output-fingerprint comparison.

Numbers from the maintainer's own runs

112 runs / 1826 recorded samples against real channels, 11 cache probes, 8 degradation runs. Representative results (median E2E P95 across runs):

Channel Reqs Success E2E P95 Output tok/s Cache hit
stealth/space-bunny-alpha (OpenRouter) 100 100% 6.13 s 31.8 89.8%
kimi-k3 43 93% 17.9 s 33.3 49.0%
z-ai/glm-5.3 100 31% 10.6 s 9.8 not reported

The z-ai/glm-5.3 row is why the tool exists: a channel can answer 100 requests and still fail two thirds of them under concurrency, and it reports no cache usage at all.

Running it

./run.sh                 # dashboard on http://127.0.0.1:8787
docker compose up -d     # or Docker

Single process only — live run state lives in memory. Do not bind HOST=0.0.0.0 to the public internet.

Known limitations

  • The dashboard UI is Chinese-only; the English README is the primary entry point for documentation.
  • Cache metrics are reported as n/a when the upstream does not report cache usage. They are never converted to 0%.
  • Degradation dimensions use a self-built curated subset (GSM8K/MMLU/C-Eval/IFEval-style), not the official datasets, so scores are not directly comparable to published leaderboards. lm-eval tasks are available for scoreable comparisons.
  • The active cache probe is serial by design; long prefixes take proportionally longer.

Verification

  • 88 automated tests, 63.9% statement coverage (gate 55%), CI green on the release commit.
  • Docker image builds in CI on every push.