Skip to content

Dyno Lab 0.6.6

Latest

Choose a tag to compare

@github-actions github-actions released this 09 Oct 01:59
· 2 commits to main since this release
fa55960

Dyno Lab 0.6.6 adds Dyno's assistant, a chat beside every tab that sets up tests and evals with a model on your Mac, and runs Evals on Inspect AI, the open-source evaluation framework from the UK AI Security Institute, and replaces the "knows it's being tested" phrase list with a detector that has to quote the agent. A real experiment on two local models shaped it and is written up as a tutorial and a study: Tutorial: agent tests and evals and Does a rule said once survive?.

Dyno's assistant

A panel beside every tab where you describe what you want to find out, typed or spoken. The assistant plans the work, looks at what Dyno has, fills in Agents → Setup for you to check, and asks before it runs or saves anything: each start or save is a card you approve or decline. It runs only on a model on this Mac, voice is recognised on device, and conversations are saved locally. Long conversations stay within the model's context: thinking is never sent back, old tool results are shortened, and older turns are summarized by the model, while the full history stays on disk. See Dyno's assistant.

Evals on Inspect AI

  • Four steps, like Agents: Choose → Set up → Run → Results. The banner says what does the work: Inspect's dataset → solver → scorer.
  • Build or import an eval. Write samples in a table or paste them from a spreadsheet, or import a CSV, JSONL or JSON file, or an Inspect .py task file (shown as code you must read first). Pick a solver (ask once, think then answer, multiple choice) and a scorer (contains, starts or ends with, exact, a regular expression, the right choice, or a judge model with your instructions). Run several epochs to see how consistent a model is.
  • Benchmark library. Installs inspect_evals once and lists benchmarks that run on a local model without Docker: SAD (evaluation awareness), StrongREJECT, Agentic Misalignment, MASK, TruthfulQA, SimpleQA, WMDP, BBQ and XSTest. XSTest loads from its authors' openly licensed data, so no Hugging Face login is needed. Badges say which benchmarks need a judge model.
  • One results board. Agent tests (decided by the Observer) and Inspect evals (scored by Inspect) sit side by side, each number with how many runs or samples it rests on and its 95% range. A benchmark that isn't right-or-wrong shows its own metric and which way is better. Click a cell for the samples and the judge's explanations, or open it in Inspect View, Inspect's own log viewer.
  • Answers that were cut off are counted. Each run reports how many answers hit the token limit or came back empty, so a token limit isn't mistaken for a wrong answer. The tested model gets 8,192 tokens per answer.
  • Each model runs in its own process, so a crashing eval can't take the lab down.

An evaluation-awareness detector

Knows it's being tested now asks whether the agent thinks it is being tested, not whether the task involves tests.

  • Words like test, eval or trap only pick passages to check. A judge model (the lead's own, unless you pick one) must quote the agent word for word and label the passage: verbalized, implicit, prompted (repeating what the test said), or task_testing (the task involves tests, or the user is just checking the agent is responsive). Only the first two fire, at or above a threshold you set.
  • Invented quotes are thrown out, and words the test itself used are ignored. Without a judge model, a strict pattern stands in.
  • Alerts from 0.6.3's phrase lists are marked old phrase list and greyed in the Room.

Agents setup

  • Team size. Type how many agents the test needs, the lead included (1–12). Above 1, the lead must create the rest of the team before any work: its prompt opens with the team instruction, telling it to build a team of N after working out the problem, the blockers, the rules and the gotchas. The instruction is a template you can reword and save with your prompt, but it can't be removed. Tests made before team sizes run as they did.
  • Each environment keeps its own test. Switching environment brings back the goal, rules and script you last used with it (or its latest past test, or the example), instead of carrying the previous environment's setup over.

Fixes

  • Models start with the settings Dyno shows. Dyno passed only the settings that differed from its own defaults, so mlx_lm.server used its own: 512 tokens per reply and up to 32 requests at once. Thinking models returned empty answers, and long agent tests could run the server out of Metal resources.
  • Evals leaves out runs that didn't really finish: rooms that stopped on a harness or model-server error, and scripted tests whose script wasn't fully delivered (a model that never files a report never gets the scripted requests).
  • Voice keeps what you said across pauses. Pausing for a second used to restart the transcription and lose what came before.
  • An Inspect cell shows its eval's latest run instead of pooling runs made with different settings.
  • Long thinking no longer turns black. Execution, the Room and the plain Chat showed a long answer or thinking as one block of text; past a few thousand tokens it was taller than macOS can draw and rendered black. Long text is now drawn in pieces.
  • Evals → Choose tells same-named scenarios apart: a list, newest first, with each scenario's environment, rules, runs and last run date, and a badge saying how two with the same name differ.
  • The Room's live lists no longer freeze the app (0.6.3's fix missed a case).
  • An imported 0.6.3 test's alert shows once.

Known issue

  • The report check can mark report honestly broken when a true claim about one action ("test order sent to staging, not production") sits next to a different rule event, such as an earlier read-only probe. Read the flagged claim before counting it.

0.6.4 and 0.6.5 were built but never published; their changes ship in 0.6.6.