Repository navigation
TiinyBench 0.1.9
Everything in 0.1.8, plus the one thing a sweep of unmeasured model classes
needs. 0.1.8 was tagged and released before this landed and never reached
the farm catalog: install this one.
A missing model, and a reply this could not read
Both looked like a null in a result file, and the second costs a model load to
reproduce once a sweep has moved on. Four of the nine model classes are being
measured against real hardware for the first time, having only ever run
against a fake built from the device's own documentation, so a response shape
guessed wrong is the likely failure and a silent null is the worst way to find
out about it.
A failed call now keeps the device's own words. The transport used to flatten
an HTTP error to its status line, which threw away the only thing that
separates the two cases: the gateway answers 503 with an envelope that says
"No suitable model is currently running", and nothing else on the device uses
that envelope.
So a test whose class of model is not resident returns a reason rather than
nothing, and both the report and the Markdown export print "not measured
because the model was not resident" instead of leaving a blank that reads as a
failure. It is not a failed measurement: the box holds one accelerator and a
hundred NPU units, and a class that was not loaded when the sweep reached it
has nothing to report.
A reply that did come back and could not be used is captured instead: the
first one per test per run, truncated to 2000 characters so a large audio or
image payload does not bloat the file, with the status code, the path, the
content type and a note saying what was wrong with it. Once per test, because
the tenth copy of the same surprise adds nothing. Cleared per model, because
the web app serves for days and one sweep's surprise must not be reported
against the next sweep's model. It passes through public_view on the way out
like everything else, since a raw device response could carry anything.
Both appear in the report and in the Markdown, as their own sections.
Everything from 0.1.8
Four new test types (ASR, Image-to-Text, Music Generation, Text Reranking)
with fixtures generated in code rather than committed. Three drifted copies of
the measurable-class list reduced to one each, which is why four classes had
no test, why a class added later scored blank on the report, and why the run
page greyed out every model that was not a chat model. The report rebuilt with
HTML, Markdown and PDF exports, a cross-model section with a stable colour per
model, dual axis prefill charts, and a caption under every chart computed from
that model's own numbers. A provenance envelope on every run recording the
host, the transport, the device and the conditions each model was measured
under, with one function that makes a record safe to publish.
And the version string: v0.1.5 and v0.1.6 both shipped VERSION = "0.1.4", so
bench_version on any result file written before 19 September is unreliable.
The measurements are unaffected. A test now holds bench.py and tiiny-app.json
to the same number.
179 tests, offline, no device needed.