Skip to content

Releases: yuvin-labs/consequencebench

ConsequenceBench 0.1.0 - Development Preview

Choose a tag to compare

@yuvin-labs yuvin-labs released this 27 Jul 15:52

ConsequenceBench evaluates whether consequential AI agents investigate the right evidence, choose the right action, preserve legitimate effects, recover from faults, and prove final source state without unsafe or duplicate effects.

Included

  • 100 canonical scenarios across banking, healthcare, cybersecurity, energy, and software delivery
  • 300 deterministic base, causal-sister, and invariance-sister lifecycle worlds
  • JSONL adapters for arbitrary candidate agents
  • evaluator-owned tools, synthetic source state, effects, and deterministic scoring
  • separate direct-capability, governance-conformance, and frozen-candidate A/B study contracts
  • public result-submission and independent-audit workflows

Validation

  • 153 public tests passed
  • 143 source files bound by the public integrity manifest
  • zero manifest-to-Git blob mismatches
  • credential-scanned deterministic source archive

Evidence boundary

This release is DEVELOPMENT_PREVIEW_NOT_QUALIFIED. Published leaderboard rows are self-reported local development evidence. They are not model ratings, production-safety certifications, or independent qualification results. All external effects are simulated and evaluator-observable.

Start with Run a Candidate, then read Claims and Evidence before publishing a result.