Skip to content

Releases: ToolsEnabled/toolsenabled-bench

ToolsEnabled Bench 0.3.2 (beta)

Pre-release

Choose a tag to compare

@JoshuaPinckard JoshuaPinckard released this 02 Oct 06:08

ToolsEnabled Bench 0.3.2 (beta) is a bug-fix and packaging release of the local benchmark builder and its MCP server. It fixes ways the web app could lose your work, gives clearer messages when local files are damaged or missing, and packages Bench as a Claude Code plugin. MIT licensed.

The runtime download keeps the name ToolsEnabled BenchMark Builder (toolsenabled-benchmark-builder-0.3.2.zip).

Data-loss fixes

  • A project that fails to open is never overwritten. In 0.3.1, if a saved project couldn't be opened (for example, a draft an MCP client saved in a shape the editor couldn't load), the app could keep the previous project selected while showing the failed one's content, and your next edit saved that content over the previous project. Now a failed open saves nothing, leaves no project selected, and says why. When the failed project was opened from Runs with "Inspect results", the project picker no longer keeps showing the previous project.
  • Routing edits stay put. In 0.3.1, changes typed into "Advanced: this routing as JSON" could be silently replaced by the next change in the routing tree or the rules table, even though the app said "Saved locally". The JSON box, the tree and the rules table now always share the same copy, so every edit survives autosave and reload. A related problem found during 0.3.2 testing, where a tree change could undo rules-table edits, was fixed before release.

Other fixes

  • One damaged project file no longer hides the rest. If a project file can't be read or is corrupt, the app lists your other projects and shows how many files it skipped. Error messages no longer show full file paths from your computer. A failed save keeps your last saved version and leaves no temporary files behind.
  • The second-Inspect issue from 0.3.1 is fixed. If a second Inspect fails, the first inspected study no longer stays selected, so Submit can't run it by mistake.
  • Your data folder must be outside the Bench folder. The web app and the MCP server now refuse to start without an absolute BENCHMARK_DATA_DIR outside the Bench folder, instead of quietly creating data inside it. The message names both BENCHMARK_DATA_DIR and the Claude plugin's data_directory option.

New: Claude Code plugin packaging

  • The ZIP is also a Claude Code plugin, toolsenabled-bench, with the 12 Bench tools and a workflow skill. On install it asks for a private data folder (data_directory), which must be outside the plugin's install and cache. See docs/CLAUDE-INSTALL.md.
  • No npm install when you install the plugin. The runtime ships no dependencies, package lock or install scripts, so installing or starting it runs no npm and needs no network.
  • The same folder also carries a Codex plugin manifest. Codex loads it with all 12 tools once BENCHMARK_DATA_DIR is set.

Unchanged: the 12 MCP tools and their schemas, execution confirmation (confirm = the exact study ID, plus trust: true for studies frozen elsewhere), the one-process-per-data-folder lock, and the print-only setup helper (node tools/mcp-config.mjs --client claude|codex|cursor|claude-desktop|deepseek), which now needs BENCHMARK_DATA_DIR set first.

Run locally

  1. Download toolsenabled-benchmark-builder-0.3.2.zip and toolsenabled-benchmark-builder-0.3.2.zip.sha256, then run sha256sum -c toolsenabled-benchmark-builder-0.3.2.zip.sha256.
  2. With Node.js 22.19 or later, extract the ZIP, then from its folder run node tools/release.mjs --verify.
  3. Pick a private data folder outside the Bench folder, for example export BENCHMARK_DATA_DIR=/absolute/path/bench-data.
  4. For the web app: node server/main.mjs, then open http://127.0.0.1:4318.
  5. For your agent: node tools/mcp-config.mjs --client <your client> and add what it prints.

Only one process can use a data folder at a time. Stop the web app before an agent uses the same folder, and give agents that run at the same time their own folders.

Upgrading from 0.3.1: 0.3.1 kept data in .benchmark-data inside the Bench folder by default. Move that folder somewhere outside the new Bench folder, set BENCHMARK_DATA_DIR to it, and run the setup helper again. Your projects and frozen studies come along: they list, export and analyze in 0.3.2. To qualify or run a study frozen with 0.3.1, use 0.3.1, because running a study needs the exact runtime it was frozen with, or freeze a new copy in 0.3.2.

Tested on this exact release (source tag v0.3.2, commit 04cf89d; runtime ZIP SHA-256 94923d535c5929013928d20c11e1a81a81a219bb96f74b03be60f37f73cd9502)

  • 548 of 548 tests passed headless, plus 53 real-browser regression checks and the full browser journey with no page errors.
  • Clean builds from the reviewed source and from this public source produced byte-identical ZIPs.
  • Real agents, through MCP only: Codex ran the offline recorded example end to end with Bench's tools: edit the composition, freeze, export, qualify, run, analyze and report. It produced study export SHA-256 92c4dbd049b8a5f474c3ed85cac6d79e0a3c12f3c4ba86af7926e3a25a111327, with 2 of 2 recorded tasks completed and passed. Claude Code 2.1.287, with the ZIP installed as a plugin and an absolute external data folder, ran the same example: 9 Bench calls, 2 of 2 tasks passed, and a byte-identical export. A DeepSeek Harness run on this exact ZIP got through freeze and export with the same export hash, but was cut short by an unrelated sandbox restart before qualify and run, so it is not counted.
  • Codex plugin loading, no sign-in: in a fresh, empty Codex profile, Codex installed the ZIP as a plugin and listed all 12 tools. With the data folder missing or not passed through, it started no tools and created no data.
  • Independent security reviews: a full review of the 0.3.2 changes, two follow-up reviews and a final check. Fixes along the way include both data-loss bugs above, file paths in error messages, and data written inside the install folder. The final check found nothing at Medium severity or above.

Limits

  • Local only. ChatGPT needs a hosted server with sign-in, which is a later phase.
  • The Claude Code marketplace entry isn't published yet. Until it is, register Bench with node tools/mcp-config.mjs --client claude.
  • The Claude Desktop bundle (.mcpb) isn't part of this release. Native Desktop install on macOS and Windows still needs testing and production signing.
  • This release was tested on Linux. macOS and Windows weren't re-tested.
  • The installed Codex plugin passes only BENCHMARK_DATA_DIR to Bench, so it suits recorded and offline studies. For a study that calls a provider, use a direct registration and forward the key's variable name yourself, as docs/CLAUDE-INSTALL.md shows.
  • Codex asks for approval before Bench's write and execution tools. For non-interactive codex exec, grant it with Codex's own per-server approval setting.
  • The included examples are authored, recorded controls. This release contains no live model runs, model scores or benchmark results.
  • A checksum shows the content is identical; it isn't a publisher signature or an independent scientific validation.
  • Running a study from someone else executes its code with your account's permissions. Run only studies you trust.
  • Known issue (Low, fix planned): the web app skips a damaged project file, but the MCP projects.list tool still fails until that file is fixed or moved out of the projects folder. Opening healthy projects by ID still works.

0.3.1, 0.3.0, 0.2.0 and 0.1.0 exports keep working with their own pinned runtimes.

ToolsEnabled is not affiliated with or endorsed by Anthropic, OpenAI, DeepSeek, Cursor or NVIDIA.

ToolsEnabled Bench 0.3.1 (beta)

Pre-release

Choose a tag to compare

@JoshuaPinckard JoshuaPinckard released this 01 Oct 20:49

Superseded by v0.3.2. Use 0.3.2 for new installs.

ToolsEnabled Bench 0.3.1 (beta) is a bug-fix release of the local benchmark builder and its MCP server. The web app now shows the right version, never writes on its own while you look around, saves every editor reliably, and shows studies you froze through MCP or the CLI. MIT licensed.

The runtime download keeps the name ToolsEnabled BenchMark Builder (toolsenabled-benchmark-builder-0.3.1.zip).

Fixes

  • Correct version everywhere. The sidebar, overview and Methods citation now take their version from the package instead of a hard-coded 0.2.0.
  • Looking doesn't write. Opening, reloading and navigating a saved project no longer changes its bytes or revision. Real edits still save locally.
  • Every editor saves. All retained editor drafts (routing, composition generation, nesting, variance, prompt set, pipeline and checks) now take part in saving, including long builder operations. A save conflict stays visible, and a revert queued during a save is kept.
  • Viewing never causes save conflicts. Choosing what to look at (tasks, snippets, audit cases, the Nesting composition) is no longer saved, so two open tabs or an MCP client no longer see false conflicts.
  • Freeze & review shows MCP and CLI studies. Studies frozen outside the web app appear with their full SHA-256. Opening one is an explicit inspection: it's verified against its expected hash before anything is shown, only one opens at a time, and inspecting never selects a study to run.
  • Older studies keep their citations. Analyzing a study frozen by 0.3.0 or earlier keeps that study's original generator citation, and new freezes cite 0.3.1.
  • Inspect is safer. A failed or busy Inspect keeps your own frozen study and its unexported run evidence, and nothing stale becomes runnable.

Unchanged: the 12 MCP tools and their schemas, execution confirmation (confirm = the exact study ID, plus trust: true for studies frozen elsewhere), the one-process-per-data-folder lock, and the print-only setup helper (node tools/mcp-config.mjs --client claude|codex|cursor|claude-desktop|deepseek).

Run locally

  1. Download toolsenabled-benchmark-builder-0.3.1.zip and toolsenabled-benchmark-builder-0.3.1.zip.sha256, then run sha256sum -c toolsenabled-benchmark-builder-0.3.1.zip.sha256.
  2. With Node.js 22.19 or later, extract the ZIP, then from its folder run node tools/release.mjs --verify.
  3. For the web app: node server/main.mjs, then open http://127.0.0.1:4318.
  4. For your agent: node tools/mcp-config.mjs --client <your client> and add what it prints.

Tested on this exact release (source tag v0.3.1, commit edc87ca; runtime ZIP SHA-256 0d43181d68293434d52191bcc693f68a231b112244b6f8f6949096f7767a77ae)

  • 483 of 483 tests passed headless, plus 33 real-browser regression checks and the full browser journey with no page errors.
  • Two clean builds, one of them from this public source, produced byte-identical ZIPs.
  • Real agents, through MCP only: Claude Code, Codex and DeepSeek Harness (DeepSeek V4.1 Flash through NVIDIA's OpenAI-compatible endpoint) each ran the offline recorded example end to end with Bench's tools: edit the composition, freeze, export, qualify, run, analyze and report. All three produced the identical study export (SHA-256 36dc653531d7bd2ac4e19894fa71cc3c7b54643b45117fae308dee2b57fb552e), with 2 of 2 recorded tasks completed and passed.
  • Independent security reviews: a full review of the 0.3.1 changes and three follow-up reviews. Fixes along the way include a data-loss bug in unsaved editor drafts, view-only selections that caused false save conflicts, and a failed Inspect that discarded the user's own evidence. The final review found nothing at Medium severity or above.

Limits

  • Local only. ChatGPT needs a hosted server with sign-in, which is a later phase.
  • Codex asks for approval before Bench's write and execution tools. For non-interactive codex exec, grant it with Codex's own per-server approval setting.
  • The included examples are authored, recorded controls. This release contains no live model runs, model scores or benchmark results.
  • A checksum shows the content is identical; it isn't a publisher signature or an independent scientific validation.
  • Running a study from someone else executes its code with your account's permissions. Run only studies you trust.
  • Known issue (Low, fix planned for 0.3.2): if you make a view-only selection between two Inspects and the second Inspect fails, the first inspected study can stay selected, and Submit would run it. Check the selected study before you submit.

0.3.0, 0.2.0 and 0.1.0 exports keep working with their own pinned runtimes.

ToolsEnabled is not affiliated with or endorsed by Anthropic, OpenAI, DeepSeek, Cursor or NVIDIA.

ToolsEnabled Bench 0.3.0 (beta)

Pre-release

Choose a tag to compare

@JoshuaPinckard JoshuaPinckard released this 01 Oct 09:34

ToolsEnabled Bench 0.3.0 (beta) adds a local MCP server, so an AI agent can run the whole Bench workflow for you in Claude Code, Codex, Cursor, Claude Desktop, DeepSeek Harness or any other MCP host. Bench is a standalone, local benchmark builder from ToolsEnabled, Inc. It composes benchmark prompts from reusable atoms, freezes and qualifies the protocol, and exports a portable study whose report regenerates from its retained evidence. MIT licensed.

The runtime download keeps the name ToolsEnabled BenchMark Builder (toolsenabled-benchmark-builder-0.3.0.zip).

What's new: Bench as an MCP server

  • node server/mcp.mjs runs Bench as a stdio MCP server on the same local .benchmark-data/ store as the web app. It opens no network listener.
  • 12 tools: projects.list, project.get, atoms.list, atom.add, composition.update, tasks.generate, study.freeze, study.qualify, study.export, study.run, study.analyze and report.get (summary, status, Markdown or HTML). See docs/MCP.md for the exact schemas.
  • Execution needs your confirmation. study.run and study.qualify execute a study's runtime and plugins, so each call must repeat the exact study ID in confirm. A study that wasn't frozen in your own store also needs an explicit trust: true.
  • One Bench process per data folder. The web app and the MCP server can use the same folder, but not at the same time; a second process refuses to start and names the one already running. Use a different BENCHMARK_DATA_DIR to run both at once.
  • Setup for your agent: node tools/mcp-config.mjs --client claude|codex|cursor|claude-desktop|deepseek prints the exact registration for that client. It only prints; it changes nothing. None of the snippets auto-approves the execution tools.

Run locally

  1. Download toolsenabled-benchmark-builder-0.3.0.zip and toolsenabled-benchmark-builder-0.3.0.zip.sha256, then run sha256sum -c toolsenabled-benchmark-builder-0.3.0.zip.sha256.
  2. With Node.js 22.19 or later, extract the ZIP, then from its folder run node tools/release.mjs --verify.
  3. For the web app: node server/main.mjs, then open http://127.0.0.1:4318.
  4. For your agent: node tools/mcp-config.mjs --client <your client> and add what it prints. For example, Codex: codex mcp add bench --env BENCHMARK_DATA_DIR=<folder>/.benchmark-data -- node <folder>/server/mcp.mjs.

Tested on this exact release (source tag v0.3.0, commit f193d5e; runtime ZIP SHA-256 a1398fb35c205e8f1dd9366a4f1ebd7fcec71b9e98440634d19df70338e44ba9)

  • 440 of 440 tests passed: the 381 from 0.2.0 plus 59 new MCP, lease and packaging tests, with no failures or skips.
  • Three independent clean builds, one of them from this public source, produced byte-identical runtime ZIPs. The extracted runtime ran the full stdio lifecycle offline, with no installed packages and with network access denied.
  • On this exact runtime ZIP, three real agents ran Bench's offline example end to end through the MCP server, using only Bench's tools: compose, inspect, freeze, export, qualify, run, analyze and report.
    • DeepSeek Harness with DeepSeek V4.1 Flash, through NVIDIA's OpenAI-compatible endpoint.
    • Codex.
    • Claude Code.
    • All three exported studies have the same SHA-256.
  • An independent security review of the MCP server found one Medium issue: two Bench processes on one data folder could lose an edit. It was fixed with the per-folder lock above and re-reviewed before release.

Limits

  • Local only. ChatGPT needs a hosted server with sign-in, which is a later phase.
  • Codex asks for approval before Bench's write and execution tools; in non-interactive codex exec, grant it with Codex's own per-server approval setting.
  • The included examples are authored, recorded controls. This release contains no live model runs, model scores or benchmark results.
  • A checksum shows the content is identical; it isn't a publisher signature or an independent scientific validation.
  • Running a study from someone else executes its code with your account's permissions; run only studies you trust.

0.2.0 and 0.1.0 exports keep working with their own pinned runtimes.

ToolsEnabled is not affiliated with or endorsed by Anthropic, OpenAI, DeepSeek, Cursor or NVIDIA.

ToolsEnabled Bench 0.2.0 (beta)

Pre-release

Choose a tag to compare

@JoshuaPinckard JoshuaPinckard released this 30 Sep 23:26

Superseded by Bench 0.3.0, which adds a local MCP server so AI agents can run the Bench workflow. 0.2.0 exports keep working.

ToolsEnabled Bench 0.2.0 (beta) is a standalone, local benchmark builder from ToolsEnabled, Inc. It composes benchmark prompts from reusable atoms, adds nesting and variance tests, freezes and qualifies the protocol, and exports a portable study whose report regenerates from its retained evidence. MIT licensed. It is a separate product from ToolsEnabled Fleet.

The downloadable runtime keeps its earlier product identity, ToolsEnabled BenchMark Builder (toolsenabled-benchmark-builder-0.2.0.zip). The app's name changes to ToolsEnabled Bench in a later release.

Run locally

  1. Download toolsenabled-benchmark-builder-0.2.0.zip and toolsenabled-benchmark-builder-0.2.0.zip.sha256, then run sha256sum -c toolsenabled-benchmark-builder-0.2.0.zip.sha256.
  2. With Node.js 22.19 or later, extract the ZIP, then from its folder run:
    node tools/release.mjs --verify
    node server/main.mjs
  3. Open http://127.0.0.1:4318.

The server listens only on 127.0.0.1 and checks the Host and Origin of every request. It needs no account, no package installation and no network connection to serve the app. Your projects and run evidence stay in its local .benchmark-data/ folder, so back that folder up.

What 0.2.0 does

  • Compositional authoring: build prompts from reusable, typed atoms and nested compositions, with explicit variants and omissions.
  • Controlled task generation: generate task sets from declared choices and seeds, with strata and recorded exclusions.
  • Freeze and qualification: bind the specification, runtime sources and analysis plan to a frozen study, then check its declared execution requirements.
  • Portable runnable exports: each export carries its pinned runtime, required plugins and file hashes, plus a CLI for verify, qualify, run and analyze.
  • Reproducible evidence reports: regenerate HTML and Markdown reports from retained attempts, responses and scoring evidence.

Studies from other people

An exported study is an executable package. Inspecting a received study's retained results doesn't run its code. Running, qualifying or re-grading it executes that study's runtime and plugins with your user account's permissions, so run only studies you trust.

Tested on this exact release

  • The source for tag v0.2.0, commit 1b46cff, passed 381 of 381 tests with no failures or skips.
  • Two independent clean clones installed from the lockfile, built, and produced byte-identical runtime ZIPs.
  • The extracted runtime passed its own verification and a full recorded-example lifecycle.
  • The whole local app passed a real Chromium walk-through.
  • An independent security review found issues in untrusted-study handling, runtime checks, report escaping and published content. All were fixed and re-reviewed before release.

Scope

  • The included examples are authored, recorded controls. This release contains no live model runs, model scores or benchmark results.
  • A checksum shows the content is identical; it isn't a publisher signature or an independent scientific validation.
  • A newly collected model response isn't expected to match an earlier one exactly.

0.1.0 exports keep working with their own pinned runtime and evidence. To revise an old frozen study, create a new draft and freeze it under a new identity.

ToolsEnabled is not affiliated with or endorsed by any model provider.