pytest/jest, but for LLM agents. Write test cases whose inputs are prompts and whose assertions are built-in checks on what the agent did — which tools it called, with which parameters, and what it finally said or computed.
eval-core is an agent testing framework: prompts in, assertions on behavior out; bring your own
harness. You implement ONE method that runs a prompt against your agent and returns what it did;
you write a tiny suite of assertions (in RON or inline); the framework runs every case, times it,
isolates panics, and hands you a report (plus a self-contained HTML dashboard). It is
game-agnostic and has zero dependency on any LLM/engine crate, so a case file is just a prompt and
a list of expectations.
One method: run a prompt, return what the agent did as [RunArtifacts]. Build the artifacts with
the chainable with_* builders and ToolCall::new.
use eval_core::{Agent, EvalError, RunArtifacts, ToolCall};
use serde_json::json;
struct MyAgent; // wraps your real LLM loop / tools
impl Agent for MyAgent {
fn run(&self, instruction: &str) -> Result<RunArtifacts, EvalError> {
// ... drive your real agent here, collecting the tool calls it made
// and its final reply. (Return `Err(EvalError::agent(...))` on a backend failure.)
Ok(RunArtifacts::new()
.with_tool_calls(vec![
ToolCall::new("calculator", json!({ "op": "add", "a": 2, "b": 2 })),
])
.with_final_text("The answer is 4.")
.with_tokens(17)) // optional; forwarded to the report's token stats
}
}Every RunArtifacts field is optional — a minimal agent can return
RunArtifacts::new().with_final_text(...) and nothing else.
A case is a name, an instruction (the prompt), and a list of expect predicates. A file may
hold a single case or a list of them:
[
(
name: "adds-two-numbers",
instruction: "what is 2 + 2?",
expect: [
CalledToolWith(tool: "calculator", args: { "op": "add" }),
FinalNumberEquals(value: 4.0),
],
),
(
name: "no-tools-for-chitchat",
instruction: "hello there",
expect: [
NoToolCalls,
FinalTextContains(text: "hello", case_insensitive: true),
],
),
]There is no setup field on the easy path — it defaults to ().
use eval_core::{load_cases, run_suite};
use eval_core::report_html::generate_report;
use std::path::Path;
# fn demo(agent: &impl eval_core::Agent) -> Result<(), eval_core::EvalError> {
let cases = load_cases(Path::new("suite/"))?; // every *.ron in the dir, sorted
let report = run_suite(agent, &cases);
println!("{report}"); // the human-readable summary table (progress goes to stderr)
// ...or persist run records and render a self-contained dashboard:
let _html = generate_report(Path::new("results/")); // writes results/report.html
# Ok(())
# }That's the whole loop: Agent::run → author Expectations → run_suite. No World, no
Setup, no Scorer to implement.
Prefer inline cases for a first run? Build a
Vec<EvalCase<(), Expectation>>directly — seeexamples/calculator.rsfor a complete, dependency-free agent (no real LLM) tested entirely with the built-in assertions.
eval-core ships a ready-to-run baseline capability suite (18 cases). Hand it straight to
run_suite:
# fn demo(agent: &impl eval_core::Agent) {
let report = eval_core::run_suite(agent, &eval_core::baseline());
println!("{report}");
# }The baseline covers three capability areas:
| File | Cases | What it checks |
|---|---|---|
arithmetic.ron |
6 | Mental/calculator math (number assertions, plus an opt-in tool subset). |
language.ron |
5 | Instruction-following over the agent's final text (fully portable). |
tool_use.ron |
7 | Tool-calling with the right name/args/count/order. |
Portable vs. adapt-me, honestly: the number/text assertions only inspect the agent's final
reply, so they hold whether the agent uses tools or reasons inline. The tool_use.ron cases (and
a small, clearly-labelled subset of arithmetic.ron) additionally assert specific tool names and
args — they assume a documented convention (a search, calculator, and send_email tool) that
you must adapt to your own agent. Do not read a tool-name mismatch there as a capability failure.
To start from the baseline as a template, dump the raw embedded files into your own suite
directory and edit them — same files that baseline() runs:
# fn demo(my_suite_dir: &std::path::Path) -> std::io::Result<()> {
for (name, contents) in eval_core::baseline_files() {
std::fs::write(my_suite_dir.join(name), contents)?;
}
# Ok(())
# }Every built-in Expectation variant, its exact RON form, and what it asserts over the run's
RunArtifacts. A case passes iff the run did not error AND every expectation holds.
| Variant | RON syntax | Asserts |
|---|---|---|
CalledTool |
CalledTool(tool: "search") |
The agent called tool at least once (any args). |
DidNotCallTool |
DidNotCallTool(tool: "search") |
The agent never called tool. |
CalledToolWith |
CalledToolWith(tool: "calc", args: { "op": "add" }) |
The agent called tool at least once with args that superset the given args (subset match — see below). |
ToolCallCount |
ToolCallCount(tool: Some("search"), min: Some(1), max: Some(1)) |
The number of calls is within [min, max] (each optional). tool: Some(name) counts only that tool; omit tool (defaults None) to count all calls. Either bound may be omitted. |
CalledToolsInOrder |
CalledToolsInOrder(tools: ["search", "send_email"]) |
The named tools appear as a subsequence of the call order — in this relative order, but not necessarily contiguous (other calls may interleave). Empty tools trivially holds. |
NoToolCalls |
NoToolCalls |
The agent made no tool calls at all (pure-reasoning / refusal check). |
FinalTextContains |
FinalTextContains(text: "paris", case_insensitive: true) |
final_text contains text. case_insensitive defaults to false (exact-case substring). Fails when there is no final text. |
FinalTextEquals |
FinalTextEquals(text: "OK") |
final_text equals text exactly, after trimming surrounding whitespace on both sides. Fails when there is no final text. |
FinalTextMatches |
FinalTextMatches(regex: "(?i)^\\s*(yes|no)\\b") |
final_text matches regex anywhere. Fails when there is no final text. A malformed regex is an authoring error (a hard EvalError::Regex), surfaced as a clearly-labelled failed predicate rather than a silent miss. |
FinalNumberEquals |
FinalNumberEquals(value: 3.33, tolerance: 0.01) |
The last number in final_text equals value within tolerance (see below). tolerance defaults to 0.0 (exact). Fails when there is no final text or it contains no number. |
NoError |
NoError |
The run reported no error. (The runner already fails a case on any run error; this is an explicit, labelled "the run was clean" predicate.) |
CalledToolWithargs are a SUBSET match. The expectedargsJSON must be a subset of the actual call's args: objects recurse key-by-key, and every other JSON value (string/number/bool/null/array) must match exactly. So{ "op": "add" }matches a call made with{ "op": "add", "a": 2, "b": 2 }, but{ "op": "sub" }does not, and an extra key the call didn't have ({ "op": "add", "c": 9 }) does not. Arrays are compared whole, not element-subset —{ "at": [1, 2] }does not match a call with"at": [1, 2, 3]. This keeps positional args like a[x, y, z]coordinate predictable.FinalNumberEqualsextracts the LAST number. The last numeric token infinal_textis taken as the agent's answer (models typically end with the answer), then compared tovaluewithintolerance. "Number" = an optionally-signed integer or decimal, with ASCII thousands-separator commas tolerated inside the integer part (-1,024.50→-1024.5). A lone-/.is not a number.
FinalTextContains.case_insensitive→falseFinalNumberEquals.tolerance→0.0(exact)ToolCallCount.tool/.min/.max→None(no restriction / unbounded)
- One file, one or many cases. A
.ronfile may hold either a singleEvalCase(...)or a list[EvalCase(...), EvalCase(...)]of related cases. Group a whole capability in one file without one-file-per-case sprawl. setupdefaults to()on the easy path, so a case omits it entirely (see the examples above). On the advanced path it defaults to yourSetup::default().load_cases(dir)reads every*.ronindirin sorted (filename) order — deterministic across runs and machines — and flattens single- and multi-case files together, so a directory may freely mix the two shapes. Non-.ronentries and subdirectories are ignored.- Fail-loud. A malformed
.ronis a hardEvalErrornaming the offending file (so a typo fails the load, and any CI load test, rather than silently dropping a case).
The easy path scores what the agent did (tool calls, params, final text/number). When you need to assert on state the agent actually changed — not just the calls it emitted — there is an escape hatch:
- Implement [
Harness] over your ownWorld+Setup:setup(&setup) -> Worldbuilds a fresh world per case,run(&instruction, &mut world) -> Result<RunArtifacts>runs the prompt against it (mutating the world). - Implement a custom [
Scorer] over the sameWorld:score(expect, &artifacts, &world)returns(label, passed)for each of a case's predicates, inspecting the post-run world. - Run with [
run_eval] (orrun_eval_with_meta) instead ofrun_suite.
use eval_core::{Harness, RunArtifacts, Scorer};
struct MyHarness;
impl Harness for MyHarness {
type World = World; // whatever state your agent mutates
type Setup = Setup; // per-case starting configuration
fn setup(&self, setup: &Setup) -> World { /* build a fresh world */ }
fn run(&self, instruction: &str, world: &mut World) -> anyhow::Result<RunArtifacts> {
/* run the prompt against `world`, return what the agent did */
}
}
struct MyScorer;
impl Scorer for MyScorer {
type World = World;
type Expect = MyPredicate; // your own predicate type
fn score(&self, expect: &MyPredicate, _artifacts: &RunArtifacts, world: &World) -> (String, bool) {
/* inspect `world` to decide if the predicate held */
}
}This is the advanced ~10%. The reference consumer (the AetherCore voxel engine) uses it for
"did the world actually change" checks — e.g. asserting N solid voxels were placed after a build
command, by scoring the post-run voxel world rather than the agent's tool calls. See
examples/minimal.rs for a complete, dependency-free Harness + Scorer.
run_suite / run_eval return an EvalReport:
- Accuracy —
passed()/total()cases, plusaccuracy()in0.0..=1.0. - Latency —
mean_latency(),p50_latency(),p95_latency()(nearest-rank quantiles). - Tokens —
total_tokens()/mean_tokens()over cases that reported a count (Nonewhen none did, so the summary never prints a misleading0). - Per-case detail — each
CaseOutcomecarriespassed, the per-predicate(label, passed)list (pinpointing which predicate failed),latency,tokens, the tool-call display strings,final_text, any runerror, and the transcript.
println!("{report}") prints an aligned summary table with the failed predicates spelled out per
case. For comparing many models/runs at a glance, persist RunRecords as JSON to a directory and
call report_html::generate_report(dir) — it writes a single, fully self-contained report.html
(no CDN, no external scripts; opens offline by double-click) with a sortable leaderboard, a
model × case heatmap, and per-case transcript expanders.
Pass a persist target on the run metadata and every run saves itself — the run JSON is written and
report.html is regenerated as part of the run, no extra wiring:
use eval_core::{run_suite_with_meta, RunMeta};
let meta = RunMeta::new(0.0, "local: my-model", system_prompt)
.persist_to("eval/results", "my-model") // dir + grouping key
.backend_kind("local") // shown in the report's Backend column
.cases_dir("eval/cases");
let report = run_suite_with_meta(&agent, &cases, meta);
// → eval/results/my-model_<timestamp>.json written, eval/results/report.html regeneratedWithout persist_to, runs are compute-only (no disk I/O) — the bare run_suite / run_eval paths
stay pure, so examples and unit tests don't write files.
A run can also POST itself to the EvalForge dashboard (evalforge.ai) after it finishes, so results show up online with no manual export. You just supply a project id + API key on the run metadata, and nothing is sent unless you configure it:
Getting your credentials:
- Sign in at evalforge.ai, then create or open a Project and copy its
Project ID — a UUID. This is the only non-secret value, and the one you hardcode (it is the
project_idin the calls below). - Mint an API key (format
sk-eval-…) under your account's API-key settings. Treat it like a password: set it as theEVALFORGE_API_KEYenvironment variable so the secret stays out of source control..upload_from_env(project_id)reads exactly that variable. .upload_to(project_id, api_key)takes the project id first, then the key — use it when you supply the key yourself; prefer.upload_from_env(project_id)to keep the key in the environment.
use eval_core::{run_suite_with_meta, RunMeta};
// Upload (and persist locally too, if you like — they share one record / dedup key):
let meta = RunMeta::new(0.0, "remote: my-model", system_prompt)
.persist_to("eval/results", "my-model")
.upload_to(project_id, api_key); // project UUID + account API key (sk-eval-…)
let report = run_suite_with_meta(&agent, &cases, meta);
// → uploaded run to evalforge: id=… deduped=false
// Or read the key from the EVALFORGE_API_KEY env var (upload-only, no local persist):
let meta = RunMeta::new(0.0, "remote: my-model", system_prompt)
.upload_from_env(project_id)
.upload_model("my-model"); // record identity, since there is no persist targetThe endpoint is fixed to evalforge.ai — there is no URL to configure, only a project id and key.
Uploads are independent of persist_to (they work with or without it), and an upload failure is
warned, never fatal, so it can never drop the eval signal. Re-uploading the same run is safe: the
server dedups on the run's (project, model, timestamp) and reports deduped: true.
cargo add eval-coreThis is a pre-release 0.3 — the API may still shift between minor versions.
MIT OR Apache-2.0.
Young but usable; runs can now be uploaded to the hosted EvalForge dashboard (evalforge.ai) — built in and configured at runtime — on top of the existing self-contained HTML report.