Skip to content

ToolsEnabled Bench 0.3.1 (beta)

Pre-release
Pre-release

Choose a tag to compare

@JoshuaPinckard JoshuaPinckard released this 01 Oct 20:49
· 1 commit to main since this release

Superseded by v0.3.2. Use 0.3.2 for new installs.

ToolsEnabled Bench 0.3.1 (beta) is a bug-fix release of the local benchmark builder and its MCP server. The web app now shows the right version, never writes on its own while you look around, saves every editor reliably, and shows studies you froze through MCP or the CLI. MIT licensed.

The runtime download keeps the name ToolsEnabled BenchMark Builder (toolsenabled-benchmark-builder-0.3.1.zip).

Fixes

  • Correct version everywhere. The sidebar, overview and Methods citation now take their version from the package instead of a hard-coded 0.2.0.
  • Looking doesn't write. Opening, reloading and navigating a saved project no longer changes its bytes or revision. Real edits still save locally.
  • Every editor saves. All retained editor drafts (routing, composition generation, nesting, variance, prompt set, pipeline and checks) now take part in saving, including long builder operations. A save conflict stays visible, and a revert queued during a save is kept.
  • Viewing never causes save conflicts. Choosing what to look at (tasks, snippets, audit cases, the Nesting composition) is no longer saved, so two open tabs or an MCP client no longer see false conflicts.
  • Freeze & review shows MCP and CLI studies. Studies frozen outside the web app appear with their full SHA-256. Opening one is an explicit inspection: it's verified against its expected hash before anything is shown, only one opens at a time, and inspecting never selects a study to run.
  • Older studies keep their citations. Analyzing a study frozen by 0.3.0 or earlier keeps that study's original generator citation, and new freezes cite 0.3.1.
  • Inspect is safer. A failed or busy Inspect keeps your own frozen study and its unexported run evidence, and nothing stale becomes runnable.

Unchanged: the 12 MCP tools and their schemas, execution confirmation (confirm = the exact study ID, plus trust: true for studies frozen elsewhere), the one-process-per-data-folder lock, and the print-only setup helper (node tools/mcp-config.mjs --client claude|codex|cursor|claude-desktop|deepseek).

Run locally

  1. Download toolsenabled-benchmark-builder-0.3.1.zip and toolsenabled-benchmark-builder-0.3.1.zip.sha256, then run sha256sum -c toolsenabled-benchmark-builder-0.3.1.zip.sha256.
  2. With Node.js 22.19 or later, extract the ZIP, then from its folder run node tools/release.mjs --verify.
  3. For the web app: node server/main.mjs, then open http://127.0.0.1:4318.
  4. For your agent: node tools/mcp-config.mjs --client <your client> and add what it prints.

Tested on this exact release (source tag v0.3.1, commit edc87ca; runtime ZIP SHA-256 0d43181d68293434d52191bcc693f68a231b112244b6f8f6949096f7767a77ae)

  • 483 of 483 tests passed headless, plus 33 real-browser regression checks and the full browser journey with no page errors.
  • Two clean builds, one of them from this public source, produced byte-identical ZIPs.
  • Real agents, through MCP only: Claude Code, Codex and DeepSeek Harness (DeepSeek V4.1 Flash through NVIDIA's OpenAI-compatible endpoint) each ran the offline recorded example end to end with Bench's tools: edit the composition, freeze, export, qualify, run, analyze and report. All three produced the identical study export (SHA-256 36dc653531d7bd2ac4e19894fa71cc3c7b54643b45117fae308dee2b57fb552e), with 2 of 2 recorded tasks completed and passed.
  • Independent security reviews: a full review of the 0.3.1 changes and three follow-up reviews. Fixes along the way include a data-loss bug in unsaved editor drafts, view-only selections that caused false save conflicts, and a failed Inspect that discarded the user's own evidence. The final review found nothing at Medium severity or above.

Limits

  • Local only. ChatGPT needs a hosted server with sign-in, which is a later phase.
  • Codex asks for approval before Bench's write and execution tools. For non-interactive codex exec, grant it with Codex's own per-server approval setting.
  • The included examples are authored, recorded controls. This release contains no live model runs, model scores or benchmark results.
  • A checksum shows the content is identical; it isn't a publisher signature or an independent scientific validation.
  • Running a study from someone else executes its code with your account's permissions. Run only studies you trust.
  • Known issue (Low, fix planned for 0.3.2): if you make a view-only selection between two Inspects and the second Inspect fails, the first inspected study can stay selected, and Submit would run it. Check the selected study before you submit.

0.3.0, 0.2.0 and 0.1.0 exports keep working with their own pinned runtimes.

ToolsEnabled is not affiliated with or endorsed by Anthropic, OpenAI, DeepSeek, Cursor or NVIDIA.