Repository navigation
tiiny-bench 0.1.10
Everything in this release came out of one thing: a sweep of all 52
installed models that came back with eleven measurements empty. Reading
what the device actually said in those failures turned four of them into
real numbers and four more into a diagnosis.
Numbers that did not exist before
- Music generation is measured. Foundation-1 runs at 0.29x real time,
about seventy seconds of wall clock for twenty seconds of audio. It also
ignores the duration you ask for and returns twenty seconds regardless,
which the record now shows because it keeps both the ask and the result. - Speech for Supertone-3, 2.10x real time.
- Both 0.6B embedding models, around 22 vectors a second at 1024 wide.
Why they were missing
- Music came back as a JSON envelope with the audio base64 inside it. The
benchmark only understood a raw file or a job to poll, so it recorded a
failure for a model that had just produced audio. - Two music models disagree about their own request shape: one requires a
duration, the other refuses it. One refuses the prompt and then says a
prompt is required. - Supertone's refusal listed the ten speakers it does have. Nobody read it.
- Two embedding models answered 502 because the device lists a model as
running before its runtime accepts connections.
What changed
Refusals are now read rather than discarded, everywhere. The benchmark
parses the field a 400 objects to, the speakers a 500 enumerates, and the
shape a refusal asks for instead, and asks again properly. It never
deletes its own request to get past a validator. A 502 straight after a
load is retried once. A test that cannot measure returns the reason
instead of a null.
Five tests could report a failure and record nothing: prefill, sustained,
concurrency, reasoning and image. That is thirty of the fifty-two models.
All five capture now, and a test enforces the rule so a new one cannot
skip it.
Two things the report was hiding
- Thirty-two models showed a name, a size and no number at all, because
only the four chat tests were ever drawn. Every class has its headline
figure now, with the detail that makes it mean something. - Sweeps driven from the browser recorded no provenance envelope, so a
three-hour run could not say which OS, Python or commit produced it,
while a one-model run from the terminal could. A test now fails when the
two writers disagree.
Also new: round-trip timing to the device, because every measured number
includes the trip to it and nothing recorded what that cost.
Not measured, and why
All four page-reading models are installable and callable by nothing on
firmware 0.1.34: the OCR route answers 404 or "model not loaded", and
chat answers "does not support chat". Two text-to-speech models do not
implement the voice mode the route defaults to. One music model wants a
session created before you generate into it.
204 tests pass, up from 183.