bench: LM Studio / Ollama latency+tok/s comparator as a native Tool #4804
MauricioPerera
started this conversation in
Show Your Plugins!
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Repo: https://github.com/MauricioPerera/bench
Third plugin in this batch (after kdd-gates and gh-discussions):
bench_runcompares latency and tokens/sec on the identical prompt across LM Studio (local, OpenAI-compatible/v1/chat/completions) and Ollama (local or:cloudmodels, native/api/chat). Purefetch(), no subprocess involved.Two real bugs found while verifying it live, might save someone else the trace:
dsh-toolsvalidates a tool's return value as "lossless JSON" and rejectsundefinedas a property value —JSON.stringifysilently drops those, but this validator throws"value is not lossless JSON"instead. Usenullfor a field you don't have, notundefined.:cloudmodels (tested withgpt-oss:20b-cloud) don't return theeval_duration/prompt_eval_durationbreakdown that local models do — onlytotal_duration. Worth a fallback if you're computing tok/s from it.Runs targets sequentially (parallel local calls would contend for the same GPU/CPU and skew the timing). MIT licensed, tagged
dsh-plugin.All reactions