Skip to content

OpenEval v0.1.0

Choose a tag to compare

@RasputinKaiser RasputinKaiser released this 15 Jul 22:55
c21f9fa

OpenEval v0.1.0

OpenEval is a free, local-first evaluation dashboard for agent CLIs. It turns runs, live traces, grader evidence, cost, performance, and accuracy data into an inspectable record you can use to decide what to ship.

Highlights

  • Run repeatable evaluation cases across Claude Code, Codex, ncode, and descriptor-defined harnesses.
  • Combine deterministic checks, trace contracts, visual evidence, optional LLM judges, and manual review.
  • Inspect live sessions with honest measured, inferred, missing, and malformed provenance.
  • Compare runs by quality, cost, tokens, duration, and case-level regressions.
  • Audit oracle coverage, known-bad rejection, evidence tiers, and weak proof surfaces.
  • Keep operator history, transcripts, workdirs, and SQLite data local by default.

Launch film

This release includes the final 29.5-second OpenEval launch film, a video poster, and a feed-readable 1280×640 social preview as downloadable release assets. Production source stays outside the product repository.

The application was developed with Noumena Code, Claude Code (including Fable 5 and Sonnet 5), and OpenAI Codex. Fable 5 contributed substantially to OpenEval itself. GPT-5.6 Sol (High) led the launch-film edit, composition, visual QA, and release pass, with additional iteration support from GPT-5.6 Luna (xhigh).

Production used Codex, ImageGen, TouchDesigner, HyperFrames, GSAP, CDP Recorder, and FFmpeg.

Huge thanks to @KingBootoshi from Right to Intelligence.

Release assets