-
Notifications
You must be signed in to change notification settings - Fork 0
Benchmarks and results
GenOS does not treat a successful demo as a benchmark. The repository contains a benchmark fleet that plans campaigns, assigns specialist agents, retains execution traces, and refuses to publish a public score when the evidence requirements are incomplete.
The results below describe local deterministic fixtures in this repository. They are neither a leaderboard against other frameworks nor a measure of production LLM quality.
| ID | Area | What the test is meant to establish | Status |
|---|---|---|---|
| B01 | Replay | State reconstruction fidelity from recorded events | Locally executable |
| B02 | Isolation | Separation of identities, events, files, and processes | Locally executable |
| B03 | Resilience | Fault injection, circuit breaking, kill switch, and recovery | Locally executable |
| B04 | Performance | Fork, replay, and telemetry latency distributions | Locally executable |
| B05 | MCP safety | Policy, quarantine, approval, and safe command behavior | Locally executable |
| B06–B09 | Public suites | SWE/Terminal-Bench, BFCL/tau-bench, WebArena/OSWorld, safety | Not claimable without approved datasets |
| B10 | Observability | Control-plane and evidence comparison | Executable with limitations |
The source of truth is the versioned portfolio and backlog.
The benchmark runs three approaches on the same fixture, ten times each:
| Approach | Success | Median wall time | Replay verified | Promotion approved |
|---|---|---|---|---|
| One fixed attempt | 0/10 | 42.02 ms | — | — |
| Sequential retry | 10/10 | 122.20 ms | — | — |
| GenOS with three isolated worlds | 10/10 | 1,004.71 ms | 10/10 | 10/10 |
GenOS is deliberately slower on this small fixture: every run performs fifteen evidence and safety operations. The result demonstrates isolation, comparison, replay, and controlled promotion—not a speed advantage. No model is called, so model calls, tokens, and model cost are all zero.
See the raw report and the explained scenario.
The resilience benchmark verifies six invariants: circuit opening, quarantine of destructive calls, continued read-only access, half-open recovery, kill-switch enforcement, and reset. Across 1,500 local fault injections, it also verifies causal rollback, dependency-aware revert, and model fallback. Reported in-process control-cycle latency is 4.75 µs at p50 and 59.959 µs at p99 on the recorded machine.
The MCP benchmark verifies 11/11 deterministic safety predicates and 5/5 required command suites. It remains a local layered test: direct MCP caller identity and per-call human approval are not yet supported.
GenOS marks a result not_claimable when it lacks an approved dataset, checksum, exact runtime identity, reproduction commands, comparison conditions, or human review. B06, B07, and B09 therefore retain null scores instead of inventing performance.
This discipline is part of the product idea: missing evidence becomes explicit, inspectable state.
node benchmarks/audit-results.mjs
./benchmarks/run-fleet.sh
./benchmarks/run-tasks.sh B02 B05
./benchmarks/run-fleet.sh --studio
cargo run --release -q -p genos-store --bin replay_benchmark -- \
--iterations 500 --events 100 --warmups 20 \
--output benchmarks/results/replay-fidelity-report.jsonContinue with Examples & evidence and the benchmark README.
GenOS is active pre-alpha research software. Verify maturity and evidence before relying on a capability. · Repository · License · Security
Start
Inside the system
- Core concepts
- Architecture
- Isolation, replay & provenance
- Hallucination reduction
- Evaluation, evolution & memory
- Orchestrator & 77 strategies
- Organizations, teams & swarms
Use GenOS
See it in action