Skip to content

v0.3.0 - Auditable research evidence and SQL evaluation

Latest

Choose a tag to compare

@ahines99 ahines99 released this 28 Sep 03:09

This release addresses the portfolio audit with stronger semantic, benchmark, domain and live-model evidence while preserving the existing release policy and all earlier results.

  • Versioned reference-grounded semantic advisory with structured claims, equivalent units, citation support, known word-digit disclosures and explicit uncertainty. It cannot override release gates.
  • 120 AI-authored traces across support and SQL, source-frozen first execution, per-dimension precision/recall/F1/coverage, blinded review exports and dual-review/adjudication tooling. All eleven deterministic-scorer disagreements are retained. No independent human validation is claimed.
  • Twelve real SQLite cases exercised through persistent evaluation and release gating. Three controls demonstrate eligible, rejected and critical-block outcomes; installed-wheel/container verification includes the SQL domain.
  • Completed predeclared study: two pinned Claude profiles, three interleaved rounds, 72 invocations, 138 paid requests and twelve actually triggered/recovered failure probes. New estimated token cost $0.484913; cumulative retained reservations $8.323672 under the $12 study ceiling.
  • Sonnet passed 22/30 baseline observations; Haiku passed 24/30 but had two critical synthetic-disclosure findings. All review/block decisions remain intact. The exploratory ten-case comparison does not establish a general model ranking.

Read the audit disposition and full results, roadmap, and verification record.

Validation: local Windows/Python 3.14 passed 392 tests with 16 PostgreSQL-only skips and 95.77% line coverage; Ruff and mypy passed. The exact release CI run verifies Linux Python 3.12/3.13/3.14, Windows Python 3.12, real PostgreSQL, installed wheels, Docker/Compose, process restart and private HTTPS behavior.

The wheel and sdist attached here come from that CI run's Linux/Python 3.12 artifact. The verification ZIP includes source identity, CI job metadata, five JUnit reports, behavioral delivery evidence, frozen benchmark/study inputs and preserved experimental outputs. SHA256SUMS covers all three binary artifacts. Original v0.2.x tags and release assets remain unchanged.

Scope: AI-assisted portfolio engineering with synthetic data, local-first operation and an authenticated shared-server option. No public evaluation SaaS, external adoption, customer-impact metric, independently human-labeled benchmark or production SLA is claimed. The requested five-agent team exceeded runtime capacity; three available subagents were reused across the five documented workstreams with coordinator-owned SQL/integration work.