Skip to content

v0.2.0 — Reliability: benchmark framework, packs ecosystem, developer experience

Latest

Choose a tag to compare

@Chloride233 Chloride233 released this 31 Jul 12:52

v0.2.0 (Reliability)

Benchmark framework (v0.2)

  • Unified runner: python benchmarks/runner.py --domain ecommerce
  • Six question types: semantic understanding, entity discovery, relationship reasoning, metric dependency, evidence grounding, safety validation
  • SRB score with hard safety gate (multiplicative factor)
  • Per-domain datasets: model.yaml / questions.json / expected_context.json / expected_metrics.json / safety_scenarios.json / evaluation.json

Ecosystem & experience

  • Five built-in semantic packs: ecommerce, saas, finance, game, healthcare
  • Streamable HTTP transport for the MCP server (--http)
  • docker compose quick start, Claude Desktop / Cursor config example
  • SafetyProvider extension point for external safety engines (JoinLint adapters)
  • Killer demo with root-cause evidence (examples/ecommerce/demo.py --pack <name>)

Verification

  • 125 tests passing (local + CI), ruff clean
  • Benchmark: metric dependency 1.000, safety validation 1.000, SRB 0.581