Skip to content

Releases: Titanium-Devops/tiiny-bench

tiiny-bench 0.1.11

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 24 Sep 17:28

TiinyBench 0.1.11

key() gains one more fallback, after every existing source (env var, farm
config, saved config, TiinyOS scrape) comes up empty: it asks for the box's
password and calls the device's own account API to get a working key
directly, no TiinyOS desktop app needed. Only runs at an interactive
terminal, never in a cron job.

Also carries the fixes since 0.1.10:

  • Two numbers on a public page that could not be true
  • Results survive an update now
  • A rate divided by two tokens is not a rate
  • A 404 that meant four different things
  • The alternation was two backends, not a fussy validator
  • A predicate that could raise instead of answering

tiiny-bench 0.1.10

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 20 Sep 00:14

Everything in this release came out of one thing: a sweep of all 52
installed models that came back with eleven measurements empty. Reading
what the device actually said in those failures turned four of them into
real numbers and four more into a diagnosis.

Numbers that did not exist before

  • Music generation is measured. Foundation-1 runs at 0.29x real time,
    about seventy seconds of wall clock for twenty seconds of audio. It also
    ignores the duration you ask for and returns twenty seconds regardless,
    which the record now shows because it keeps both the ask and the result.
  • Speech for Supertone-3, 2.10x real time.
  • Both 0.6B embedding models, around 22 vectors a second at 1024 wide.

Why they were missing

  • Music came back as a JSON envelope with the audio base64 inside it. The
    benchmark only understood a raw file or a job to poll, so it recorded a
    failure for a model that had just produced audio.
  • Two music models disagree about their own request shape: one requires a
    duration, the other refuses it. One refuses the prompt and then says a
    prompt is required.
  • Supertone's refusal listed the ten speakers it does have. Nobody read it.
  • Two embedding models answered 502 because the device lists a model as
    running before its runtime accepts connections.

What changed

Refusals are now read rather than discarded, everywhere. The benchmark
parses the field a 400 objects to, the speakers a 500 enumerates, and the
shape a refusal asks for instead, and asks again properly. It never
deletes its own request to get past a validator. A 502 straight after a
load is retried once. A test that cannot measure returns the reason
instead of a null.

Five tests could report a failure and record nothing: prefill, sustained,
concurrency, reasoning and image. That is thirty of the fifty-two models.
All five capture now, and a test enforces the rule so a new one cannot
skip it.

Two things the report was hiding

  • Thirty-two models showed a name, a size and no number at all, because
    only the four chat tests were ever drawn. Every class has its headline
    figure now, with the detail that makes it mean something.
  • Sweeps driven from the browser recorded no provenance envelope, so a
    three-hour run could not say which OS, Python or commit produced it,
    while a one-model run from the terminal could. A test now fails when the
    two writers disagree.

Also new: round-trip timing to the device, because every measured number
includes the trip to it and nothing recorded what that cost.

Not measured, and why

All four page-reading models are installable and callable by nothing on
firmware 0.1.34: the OCR route answers 404 or "model not loaded", and
chat answers "does not support chat". Two text-to-speech models do not
implement the voice mode the route defaults to. One music model wants a
session created before you generate into it.

204 tests pass, up from 183.

TiinyBench 0.1.9

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 19 Sep 16:51

Everything in 0.1.8, plus the one thing a sweep of unmeasured model classes
needs. 0.1.8 was tagged and released before this landed and never reached
the farm catalog: install this one.

A missing model, and a reply this could not read

Both looked like a null in a result file, and the second costs a model load to
reproduce once a sweep has moved on. Four of the nine model classes are being
measured against real hardware for the first time, having only ever run
against a fake built from the device's own documentation, so a response shape
guessed wrong is the likely failure and a silent null is the worst way to find
out about it.

A failed call now keeps the device's own words. The transport used to flatten
an HTTP error to its status line, which threw away the only thing that
separates the two cases: the gateway answers 503 with an envelope that says
"No suitable model is currently running", and nothing else on the device uses
that envelope.

So a test whose class of model is not resident returns a reason rather than
nothing, and both the report and the Markdown export print "not measured
because the model was not resident" instead of leaving a blank that reads as a
failure. It is not a failed measurement: the box holds one accelerator and a
hundred NPU units, and a class that was not loaded when the sweep reached it
has nothing to report.

A reply that did come back and could not be used is captured instead: the
first one per test per run, truncated to 2000 characters so a large audio or
image payload does not bloat the file, with the status code, the path, the
content type and a note saying what was wrong with it. Once per test, because
the tenth copy of the same surprise adds nothing. Cleared per model, because
the web app serves for days and one sweep's surprise must not be reported
against the next sweep's model. It passes through public_view on the way out
like everything else, since a raw device response could carry anything.

Both appear in the report and in the Markdown, as their own sections.

Everything from 0.1.8

Four new test types (ASR, Image-to-Text, Music Generation, Text Reranking)
with fixtures generated in code rather than committed. Three drifted copies of
the measurable-class list reduced to one each, which is why four classes had
no test, why a class added later scored blank on the report, and why the run
page greyed out every model that was not a chat model. The report rebuilt with
HTML, Markdown and PDF exports, a cross-model section with a stable colour per
model, dual axis prefill charts, and a caption under every chart computed from
that model's own numbers. A provenance envelope on every run recording the
host, the transport, the device and the conditions each model was measured
under, with one function that makes a record safe to publish.

And the version string: v0.1.5 and v0.1.6 both shipped VERSION = "0.1.4", so
bench_version on any result file written before 19 September is unreliable.
The measurements are unaffected. A test now holds bench.py and tiiny-app.json
to the same number.

179 tests, offline, no device needed.

TiinyBench 0.1.8

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 19 Sep 16:46

Four more kinds of model can be measured, the report is rebuilt, and every run
now records what makes it comparable to somebody else's box.

Every run carries its provenance

A run taken without this is a run that cannot be compared tomorrow. Each file
now has a provenance block with a schema number on it.

Per run: the host that drove the benchmark (os, release, version, arch,
python, logical CPUs, physical RAM, machine model), the transport (plane,
gateway port, vhost if used, and the round trip measured before any work
starts), the device (name, model, serial, TiinyOS, service version, NPU
unit budget), and the tool (bench version, suite version, envelope schema,
git commit). A benchmark driven from a Mac over a cable and one driven from a
Windows box over Wi-Fi are not the same measurement, and until now nothing in
the file said which.

Per model rather than per run, because a sweep loads and unloads as it goes
and the conditions the fifth model met are not the ones the first met: the
resident set with each model's NPU units at the moment that model's tests
begin, units used and free, whether another process held the device lock, the
sampling parameters actually sent, which tests ran and which were skipped, and
the served context window. The lock is read and never taken, because a
benchmark that fought for it would change the thing it measures.

Temperature and power are recorded as null with a reason rather than omitted.
A missing key reads as an oversight and a zero reads as a measurement, and
neither is true of something nobody could measure. The machine model on macOS
is the same: reading it means starting a program, which this app may not do.

public_view is the one function that makes a record shareable, so an upload
path and a report cannot disagree about what is safe. It hashes the serial so
runs group by box without naming one, drops the device address, the vhost and
the owner's chosen name for the box, drops any string that looks like a path
under a home directory, and adds an account attribution. Every measured number
survives untouched: it makes a record publishable, not smaller.

All additive. Files written before any of this still load, render, export and
publish.

Four new test types

Twenty-nine models are installed across nine classes and the suite had tests
for five of them.

Class Test The figure
ASR asr seconds of audio transcribed per second of wall clock
Image-to-Text ocr seconds per page, and whether the digits came back
Music Generation music seconds of audio produced per second of wall clock
Text Reranking rerank query-document pairs scored per second

Their fixtures are generated in code, not committed: a WAV whose header states
its duration exactly, because the real-time factor divides by it, and a PNG of
seven-segment digits, because a font table is three hundred lines nobody can
check by eye. Both are byte for byte identical every run.

Three copies of one list, now one each

The set of measurable model classes was written down three times and all three
had drifted. CHAT_TYPES was a second hand-written copy, which is the actual
reason four classes had no test. The report's leaderboard named the three
result blocks it knew how to read, so a class added later scored blank. And
the run page kept the narrowest copy of all, greying out every model that was
not a chat model, which was already wrong for the image, speech and embedding
tests that already existed.

All three read the registry now. The run page also plans per model, so a
Text-to-Image model picked beside a chat model runs its own test rather than
refusing "thinking", and the time estimate counts tests that will really run.

The report, rebuilt

Exports: HTML hands over the existing self-contained file rather than
generating a second one that could disagree with it; Markdown renders server
side off the same loader; PDF is the browser's own print through a stylesheet,
because a PDF library is a dependency this app will not take.

The print path was tested rather than assumed, by emulating print media and
then producing the PDF. That testing caught three real defects that would
otherwise have shipped: a stranded near-empty page, a grey block where an
empty grid cell printed, and a bordered section stretched down a mostly empty
sheet.

Per model: a dual axis prefill chart with time to first token solid on the
left and decode rate dashed on the right and every point labelled on both, the
same points as a table, concurrency as grouped bars with wall time under each
level, reasoning as two horizontal bars carrying wall time and token count, a
device telemetry strip, and a caption under every chart computed from that
model's own numbers. A box whose aggregate is flat gets a caption that says it
is queueing, and the words come out of the ratio rather than out of a template.

A cross-model section with a stable colour per model taken from a hash of the
id, so adding or dropping a model never repaints the others: sustained decode,
throughput per NPU unit, every prefill curve on one axis, and aggregate
throughput against stream count. Newest run per model, never best, because a
better old run would flatter the box.

Long model names are unreadable under vertical bars, so those two charts are
horizontal. The name gutter is measured from the widest label in the set: it
was capped, which fitted every name but the longest, and that one ran under
the start of its own bar.

The version string

v0.1.5 and v0.1.6 were both tagged and both shipped VERSION = "0.1.4", so
the app printed 0.1.4 in its banner, in --version and in the selfcheck, and
stamped bench_version: 0.1.4 into every result file it wrote. Anyone reading
a result file written before 19 September should treat that field as
unreliable; the measurements themselves are unaffected. A test now holds
bench.py and tiiny-app.json to the same number.

Two things learned about the device

/openapi.json stopped being per-service in TiinyOS 1.0.0: the AI Gateway
moved behind port 80 and that path is served by the front proxy, so every
vhost spelling returns the same device management spec. Confirm a route
without running it by asking with GET, where a POST-only route that exists
answers 405 and one that does not answers 404.

A route whose class of model is not loaded answers 503 with an envelope
nothing else on the device uses. Code that reads it as a transport failure
reports a missing model as a broken box.

173 tests, offline, no device needed.

TiinyBench 0.1.7

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 19 Sep 16:09

Four more kinds of model can be measured, and the run page will offer them.

Four new test types

Twenty-nine models are installed across nine classes and the suite had tests
for five of them, so nine models wore a type with nothing behind it.

Class Test The figure
ASR asr seconds of audio transcribed per second of wall clock
Image-to-Text ocr seconds per page, and whether the digits came back
Music Generation music seconds of audio produced per second of wall clock
Text Reranking rerank query-document pairs scored per second

The audio clip and the page of digits are generated in code rather than
committed: a fixture you cannot diff is one you cannot trust when a number
moves. Both are byte for byte identical every run, and the tests decode them
back and read the digits off the pixels.

OCR takes whichever of two paths answers and records which, because a box whose
OCR model is a vision language model has no gateway route and has to be read
through chat completions, and those two numbers are not comparable. Music
handles both a blocking call and a polled job for the same reason.

Transcription accuracy is deliberately not scored: it needs a clip of known
speech, which cannot be generated without a text-to-speech model loaded, and a
benchmark that depends on a second model is measuring two things. Reranking
does get a free sense check, because one of its four passages actually answers
the query.

Three lists that had drifted are now one each

CHAT_TYPES was a second hand-written copy of the class list and had gone
stale, which is how four classes came to have no test at all. It is a view of
SUITES now.

The leaderboard named the three result blocks it knew how to read, so any class
added later scored blank on the report. It derives them from the registry now.

The run page kept a third copy, and the narrowest: it greyed out every model
that was not a chat model, which was already wrong for the image, speech and
embedding tests. The suites travel with the catalogue now, so the page has no
list of its own, and it plans per model: the ticked tests narrowed to the ones
that mean something for that class, and the whole class suite when that leaves
nothing. A Text-to-Image model picked beside a chat model runs its own test
rather than refusing "thinking". The estimate counts tests that will really
run rather than models times ticked boxes.

The version is one number again

v0.1.5 and v0.1.6 were both tagged and both shipped with VERSION = "0.1.4",
so the app printed 0.1.4 everywhere and stamped it into every result file it
wrote. Fixed, with a test so it cannot drift again.

Two things learned about the device

/openapi.json stopped being per-service in TiinyOS 1.0.0. The AI Gateway
moved behind port 80 and that path is served by the front proxy rather than
proxied on, so every vhost spelling returns the same 44-path device management
spec. Confirm a route without running it by asking with GET: a POST-only route
that exists answers 405, one that does not answers 404.

A route whose class of model is not loaded answers 503 with an envelope nothing
else on the device uses:

{"error": {"message": "No suitable model is currently running.",
           "type": "service_unavailable"}}

Code that reads that as a transport failure reports a missing model as a broken
box. The fake device reproduces it and four tests hold the behaviour.

Also

A troubleshooting entry for a stale cached gateway port, which survives the
move to port 80 and which Detect again will not clear because it prefers what
it already has. Restart the app. The README's table of what the suite measures
listed four tests; it lists all eleven.

131 tests, offline, no device needed.

v0.1.6

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 18 Sep 14:06

TiinyBench 0.1.6

Problem. 0.1.5 folded the sidebar to icons, but the main column was capped at 1500 pixels and left-aligned, so on a wide window the space the fold gave back sat empty on the right and the chat stayed the same width. Jason, looking at the folded chat: "the chat section stayed exactly the same width left to right."

Fix. The cap is gone. The main column now fills whatever the sidebar leaves, folded or wide, on every page. Prose keeps its own readable measure (the page blurbs and the empty-chat note are capped in characters); tables and the chat grow with the window.

Tests. The 100 offline tests pass unchanged; the chat page rendered at 2000 wide with the sidebar open runs to the right edge (grid: 236 px sidebar, main the rest).

v0.1.5

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 18 Sep 13:34

TiinyBench 0.1.5

Problem. The Chat page is the first page in TiinyBench that wants the whole width of the window: a model card, the conversation and the loaded models side by side. The sidebar took 236 pixels of it for eleven labels the eye already knows by their glyphs. Jason, looking at a chat: "if I could collapse the left-hand nav to just icons, would I have a little bit more space?"

Fix. The sidebar folds to icons. A button in the brand row folds it to 64 pixels and back; the choice is remembered per browser and applied before the first paint, so a folded sidebar never flashes wide on load. Folded, each button carries its label as a tooltip, section headings become hairlines, and the footer keeps only the Tiiny mark. On a phone the sidebar is already a strip and the button is not offered. Nothing about the pages, the routes or the chat changes.

Tests. The 100 offline tests from 0.1.4 pass unchanged (the page's script order, the em dash rule, the chat and the catalogue); both states rendered at 1440 wide: 236 px and 64 px, the document exactly the viewport in each.

v0.1.4

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 18 Sep 13:04

TiinyBench 0.1.4

Problem. TiinyBench could tell you what a model costs and never let you talk to one. Its sibling on the same hardware, AINode Pocket, has a chat tab that reports every reply's cost under the reply and the device beside the conversation, and asked for it here. The two apps also derived the same numbers from the same two gateway blocks in two places, which is the arrangement that drifts the first time the gateway renames a field and then has a chat bar and a saved benchmark disagreeing about the same request.

Fix. A Chat page, Pocket's three columns adapted to one Tiiny. Under every reply: time to first token, decode tok/s, total wall time, tokens in with the cache hits, tokens out, which device and model answered, and why generation stopped, all read off the gateway's own timings and usage on the stream's final chunk plus this machine's clock; the numbers are saved with the message. Left: a model card for whatever is selected (type, params, size, NPU units, input and output, where it is loaded, the vendor's description, a catalogue link) and the last twenty conversations, kept in the browser. Right: every model loaded on the box with its units and status and an Unload button, the NPU budget bar, and a Load panel that refuses in a sentence when a model is not installed, cannot chat, or would not fit, because the device accepts a load that does not fit and rolls it back without saying so. Replies render as markdown with highlighted code blocks and a Copy button; a reasoning model's chain of thought sits in a collapsible block behind the Thinking toggle; temperature, max tokens and a system prompt are controls. New routes in serve.py: POST /api/chat (streamed, ending with an event: stats before [DONE]), GET /api/model_card, GET /api/instances, POST /api/instances/load and /unload. All of it goes through this release's own transport negotiation, so the chat reaches the gateway on whichever port and vhost the firmware put it behind rather than assuming one. derive_stats and chat_stats are copied verbatim into bench.py with a comment naming their origin, and bench.chat now reads that derivation instead of its own second copy of the same arithmetic, so the page, a saved run and Pocket all compute one set of numbers. The Models page's own load goes through the same guard, minus the chat requirement, so both doors onto the device's start give the same refusals and neither can drift into letting a silent rollback through.

Tests. tests/test_chat.py and tests/test_manage.py, sixty tests on stdlib unittest against a fake device in tests/fake_device.py, offline, beside the discovery tests this release already had: a hundred in total, all passing on 3.9 and 3.14 and through the invocation CI uses. They hold the derivation against a completion measured on real hardware, the SSE framing (a blank line after every frame, event: stats then [DONE], one of each), the three load refusals word for word, a load that goes loading to running, the device taking an oversized load and rolling it back in silence, a stream that fails part way reporting nothing rather than zeroes, the model card's fields, enable_thinking arriving on the wire where the runtime reads it, and the page carrying no external script, stylesheet or font. Verified against a real Tiiny as well: chats to Ornith-1.0-35B reconcile with the gateway's raw blocks to the rounding, and the rail lists Ornith at 50 units, the embedder at 1 and the TTS at 7 against a 58 of 100 budget. The browser check that fails if the page does nothing now loads the Chat page too, at 2000, 1440 and 390 wide, with no page errors and nothing off the right edge.

Also. The Models page's catalogue is read off the box on every call and nothing about it is pinned, but you could not tell that from the page, and when the model store did not answer the read came back as an empty list and the page printed "0 in the catalogue" as though the store were empty. bench.online now raises instead, /api/manage carries online_error and online_read_at beside online, and the page says "catalogue: 50 entries, read from the box at 07:49" with a Refresh beside it, or "the box did not answer for the catalogue: ..." in red when it did not. Neither place prints a count it does not have.

Four transition: all rules now name what they animate, the two full-height rules are 100dvh so a mobile browser's shrinking chrome does not make the sidebar jump, and the two glyph verdicts on Can I Run It read as words, because the colour was already saying fits or does not fit. The dead-page bug this branch also carried a fix for is not listed: it never shipped, and the fix for it landed upstream first and better, along with the phone-width grid track.

Riding along in the same release is the field work that landed while this was being built: firmware 1.0 moved the gateway, so the app now finds the box over USB, LAN and the TiinyOS proxy names and negotiates the port and vhost per service instead of assuming one; discovery works on Windows, which it did not, so a tester with a Tiiny on the same network is no longer told there is none; every text read and write names its encoding, which is what made that Windows failure hard to see; the footer says which box it is before a key is entered; the connection tiles fit the values they hold, including a full serial and a vhost name; and CI loads the page in a real browser on every push and fails if the controls are dead.

TiinyBench 0.1.3

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 17 Sep 02:57
07b2f0c

Running 0.1.2 from the launcher at about 2000px wide, the Connection page showed REACHED AT as 172.17.7.1 on a box that is at 172.17.7.177, and HOW as given, port 80 as p8800.api. with the rest gone. The tiles were drawn in the face the dashboard numbers use, at one size whether the value was three characters or twenty.

  • The connection tile values now get a face sized for a short line of text, they wrap instead of running off the card, and hovering one shows the whole value.
  • The tiles take the width their own value needs, so 1 TB and given, port 80 as p8800.api.tiiny are no longer made the same size.
  • The discovery file on :39218 is asked once per address and fills in whatever the management API could not answer, so a box handed over by the launcher stops reading "found, unnamed" over a column of dashes.
  • The tiles say what found the box, "the launcher" or "the farm", rather than naming this suite's own environment variables.

The browser job in CI now loads the page at 2000 as well as 1440 and 390, and fails if any tile value is wider than the box holding it.

TiinyBench 0.1.2

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 16 Sep 23:23
1097455

Two testers reported the same thing in different words, and it was one cause: the page never ran.

The confirm dialog added in 0.1.1 sat above the script's own scope and called $ on its first line, so the whole <script> stopped on a ReferenceError before it asked for anything. Nothing below it executed. No navigation, no request for the device, no find button. The app came up looking finished with every control dead, on Windows and on a Mac alike.

Fixed in this release

  • The page runs. Connection opens, the find button finds, the dashboard fills in.
  • Discovery works on Windows. It used to shell out to ip and then ifconfig, and Windows has neither, so the candidate list was empty and a Tiiny on the same network could not be found. It is stdlib now and the same on every platform.
  • The scan asks the responder the farm CLI asks, a GADGET_DISCOVER_V1 datagram on port 39217, which is the only probe that finds a box whose address moved.
  • The Connection panel says what it is doing. It paints "Looking for your Tiiny" before the first request leaves, says so plainly when nothing answers, and offers the address box and the key box either way.
  • The dashboard no longer dies on a fresh install. Asking for the leaderboard without a key was dropping the connection.
  • tiiny-bench --where prints what it looked for and exits cleanly when there is no Tiiny attached.
  • At 390 pixels wide the navigation used to push every panel off the right edge behind a sideways scroll.
  • Text files are read and written as UTF-8 everywhere, so the report comes out right on Windows.

New

Continuous integration on Linux, macOS and Windows, with 28 unit tests. None of it needs a Tiiny.

A Tiiny on a different network is still reached by typing its address into the Connection panel. A discovery datagram does not cross a router, and the panel now says so.