Skip to content

Benchmarks and results

Ginks edited this page Aug 22, 2026 · 1 revision

Benchmarks & results

GenOS does not treat a successful demo as a benchmark. The repository contains a benchmark fleet that plans campaigns, assigns specialist agents, retains execution traces, and refuses to publish a public score when the evidence requirements are incomplete.

The results below describe local deterministic fixtures in this repository. They are neither a leaderboard against other frameworks nor a measure of production LLM quality.

What is evaluated

ID Area What the test is meant to establish Status
B01 Replay State reconstruction fidelity from recorded events Locally executable
B02 Isolation Separation of identities, events, files, and processes Locally executable
B03 Resilience Fault injection, circuit breaking, kill switch, and recovery Locally executable
B04 Performance Fork, replay, and telemetry latency distributions Locally executable
B05 MCP safety Policy, quarantine, approval, and safe command behavior Locally executable
B06–B09 Public suites SWE/Terminal-Bench, BFCL/tau-bench, WebArena/OSWorld, safety Not claimable without approved datasets
B10 Observability Control-plane and evidence comparison Executable with limitations

The source of truth is the versioned portfolio and backlog.

Headline result: safe parallel debugging

The benchmark runs three approaches on the same fixture, ten times each:

Approach Success Median wall time Replay verified Promotion approved
One fixed attempt 0/10 42.02 ms — —
Sequential retry 10/10 122.20 ms — —
GenOS with three isolated worlds 10/10 1,004.71 ms 10/10 10/10

GenOS is deliberately slower on this small fixture: every run performs fifteen evidence and safety operations. The result demonstrates isolation, comparison, replay, and controlled promotion—not a speed advantage. No model is called, so model calls, tokens, and model cost are all zero.

See the raw report and the explained scenario.

Resilience and MCP safety

The resilience benchmark verifies six invariants: circuit opening, quarantine of destructive calls, continued read-only access, half-open recovery, kill-switch enforcement, and reset. Across 1,500 local fault injections, it also verifies causal rollback, dependency-aware revert, and model fallback. Reported in-process control-cycle latency is 4.75 µs at p50 and 59.959 µs at p99 on the recorded machine.

The MCP benchmark verifies 11/11 deterministic safety predicates and 5/5 required command suites. It remains a local layered test: direct MCP caller identity and per-call human approval are not yet supported.

Why public scores remain empty

GenOS marks a result not_claimable when it lacks an approved dataset, checksum, exact runtime identity, reproduction commands, comparison conditions, or human review. B06, B07, and B09 therefore retain null scores instead of inventing performance.

This discipline is part of the product idea: missing evidence becomes explicit, inspectable state.

Reproduce the campaign

node benchmarks/audit-results.mjs
./benchmarks/run-fleet.sh
./benchmarks/run-tasks.sh B02 B05
./benchmarks/run-fleet.sh --studio

cargo run --release -q -p genos-store --bin replay_benchmark -- \
  --iterations 500 --events 100 --warmups 20 \
  --output benchmarks/results/replay-fidelity-report.json

Continue with Examples & evidence and the benchmark README.

Clone this wiki locally