Skip to content

ToolsEnabled Bench 0.2.0 (beta)

Pre-release
Pre-release

Choose a tag to compare

@JoshuaPinckard JoshuaPinckard released this 30 Sep 23:26
· 3 commits to main since this release

Superseded by Bench 0.3.0, which adds a local MCP server so AI agents can run the Bench workflow. 0.2.0 exports keep working.

ToolsEnabled Bench 0.2.0 (beta) is a standalone, local benchmark builder from ToolsEnabled, Inc. It composes benchmark prompts from reusable atoms, adds nesting and variance tests, freezes and qualifies the protocol, and exports a portable study whose report regenerates from its retained evidence. MIT licensed. It is a separate product from ToolsEnabled Fleet.

The downloadable runtime keeps its earlier product identity, ToolsEnabled BenchMark Builder (toolsenabled-benchmark-builder-0.2.0.zip). The app's name changes to ToolsEnabled Bench in a later release.

Run locally

  1. Download toolsenabled-benchmark-builder-0.2.0.zip and toolsenabled-benchmark-builder-0.2.0.zip.sha256, then run sha256sum -c toolsenabled-benchmark-builder-0.2.0.zip.sha256.
  2. With Node.js 22.19 or later, extract the ZIP, then from its folder run:
    node tools/release.mjs --verify
    node server/main.mjs
  3. Open http://127.0.0.1:4318.

The server listens only on 127.0.0.1 and checks the Host and Origin of every request. It needs no account, no package installation and no network connection to serve the app. Your projects and run evidence stay in its local .benchmark-data/ folder, so back that folder up.

What 0.2.0 does

  • Compositional authoring: build prompts from reusable, typed atoms and nested compositions, with explicit variants and omissions.
  • Controlled task generation: generate task sets from declared choices and seeds, with strata and recorded exclusions.
  • Freeze and qualification: bind the specification, runtime sources and analysis plan to a frozen study, then check its declared execution requirements.
  • Portable runnable exports: each export carries its pinned runtime, required plugins and file hashes, plus a CLI for verify, qualify, run and analyze.
  • Reproducible evidence reports: regenerate HTML and Markdown reports from retained attempts, responses and scoring evidence.

Studies from other people

An exported study is an executable package. Inspecting a received study's retained results doesn't run its code. Running, qualifying or re-grading it executes that study's runtime and plugins with your user account's permissions, so run only studies you trust.

Tested on this exact release

  • The source for tag v0.2.0, commit 1b46cff, passed 381 of 381 tests with no failures or skips.
  • Two independent clean clones installed from the lockfile, built, and produced byte-identical runtime ZIPs.
  • The extracted runtime passed its own verification and a full recorded-example lifecycle.
  • The whole local app passed a real Chromium walk-through.
  • An independent security review found issues in untrusted-study handling, runtime checks, report escaping and published content. All were fixed and re-reviewed before release.

Scope

  • The included examples are authored, recorded controls. This release contains no live model runs, model scores or benchmark results.
  • A checksum shows the content is identical; it isn't a publisher signature or an independent scientific validation.
  • A newly collected model response isn't expected to match an earlier one exactly.

0.1.0 exports keep working with their own pinned runtime and evidence. To revise an old frozen study, create a new draft and freeze it under a new identity.

ToolsEnabled is not affiliated with or endorsed by any model provider.