Repository navigation
TiinyBench 0.1.7
Four more kinds of model can be measured, and the run page will offer them.
Four new test types
Twenty-nine models are installed across nine classes and the suite had tests
for five of them, so nine models wore a type with nothing behind it.
| Class | Test | The figure |
|---|---|---|
| ASR | asr |
seconds of audio transcribed per second of wall clock |
| Image-to-Text | ocr |
seconds per page, and whether the digits came back |
| Music Generation | music |
seconds of audio produced per second of wall clock |
| Text Reranking | rerank |
query-document pairs scored per second |
The audio clip and the page of digits are generated in code rather than
committed: a fixture you cannot diff is one you cannot trust when a number
moves. Both are byte for byte identical every run, and the tests decode them
back and read the digits off the pixels.
OCR takes whichever of two paths answers and records which, because a box whose
OCR model is a vision language model has no gateway route and has to be read
through chat completions, and those two numbers are not comparable. Music
handles both a blocking call and a polled job for the same reason.
Transcription accuracy is deliberately not scored: it needs a clip of known
speech, which cannot be generated without a text-to-speech model loaded, and a
benchmark that depends on a second model is measuring two things. Reranking
does get a free sense check, because one of its four passages actually answers
the query.
Three lists that had drifted are now one each
CHAT_TYPES was a second hand-written copy of the class list and had gone
stale, which is how four classes came to have no test at all. It is a view of
SUITES now.
The leaderboard named the three result blocks it knew how to read, so any class
added later scored blank on the report. It derives them from the registry now.
The run page kept a third copy, and the narrowest: it greyed out every model
that was not a chat model, which was already wrong for the image, speech and
embedding tests. The suites travel with the catalogue now, so the page has no
list of its own, and it plans per model: the ticked tests narrowed to the ones
that mean something for that class, and the whole class suite when that leaves
nothing. A Text-to-Image model picked beside a chat model runs its own test
rather than refusing "thinking". The estimate counts tests that will really
run rather than models times ticked boxes.
The version is one number again
v0.1.5 and v0.1.6 were both tagged and both shipped with VERSION = "0.1.4",
so the app printed 0.1.4 everywhere and stamped it into every result file it
wrote. Fixed, with a test so it cannot drift again.
Two things learned about the device
/openapi.json stopped being per-service in TiinyOS 1.0.0. The AI Gateway
moved behind port 80 and that path is served by the front proxy rather than
proxied on, so every vhost spelling returns the same 44-path device management
spec. Confirm a route without running it by asking with GET: a POST-only route
that exists answers 405, one that does not answers 404.
A route whose class of model is not loaded answers 503 with an envelope nothing
else on the device uses:
{"error": {"message": "No suitable model is currently running.",
"type": "service_unavailable"}}
Code that reads that as a transport failure reports a missing model as a broken
box. The fake device reproduces it and four tests hold the behaviour.
Also
A troubleshooting entry for a stale cached gateway port, which survives the
move to port 80 and which Detect again will not clear because it prefers what
it already has. Restart the app. The README's table of what the suite measures
listed four tests; it lists all eleven.
131 tests, offline, no device needed.