Skip to content

local-enough 0.1.0

Latest

Choose a tag to compare

@B0yko B0yko released this 29 Sep 00:30
· 10 commits to main since this release

First release: measure whether a back-office AI task can run on hardware you own, cost it against cloud models, and route each task to the cheapest model that meets your quality bar.

uvx --from git+https://github.com/B0yko/local-enough@v0.1.0 local-enough --help
docker pull ghcr.io/b0yko/local-enough:0.1.0

What's in it

  • bench: five bundled tasks (CRM extraction, BANKING77 intent classification, PII redaction, meeting summaries, vendor-record matching) with gold labels and calib/test splits, run against local MLX models, OpenAI-compatible endpoints and cloud models, plus three non-LLM baselines. Bootstrap CIs, p50/p95 latency, throughput at concurrency 1 and 4, a 20-minute soak, peak memory, and a cost ledger with a hard budget cap.
  • report: per-task accuracy-vs-cost frontiers, break-even volume with a capacity check (dedicated and shared-machine scenarios), a deterministic per-task verdict and a filled-in ADR, as HTML and Markdown.
  • judge: a summarisation judge calibrated on construction-labelled summaries, with Rogan–Gladen bias correction.
  • route: an OpenAI-compatible router that serves each task alias with the cheapest candidate that met the bar on calib, escalates on deterministic gate failures, keeps data_must_stay_local tasks off cloud endpoints, and can be simulated offline or replayed live.
  • Bring your own task via task.yaml; init, datasets, models pull, memory, power-probe; router image for amd64 and arm64.

Reference run (Mac Studio M4 Max 128 GB; x-ai/grok-4.7, openai/gpt-5.4-nano, deepseek/deepseek-v3.2, qwen/qwen3-235b-a22b-2507; Qwen3-4B and Qwen2.5-1.5B in MLX 4-bit): local met the quality bar on 2 of 5 tasks (classification through the TF-IDF baseline, entity matching through the 1.5B model on calib), PII redaction cannot stay local at a 0.95 F2 bar (best local 0.922), and the gated router costs 76.9% less on the mixed workload than sending everything to the frontier model. Every number in the README is regenerated from the committed run in CI.

Total API spend to build and verify this release: $5.99. Local model downloads: 3.16 GB.