Skip to content

Releases: ahines99/agent-eval-redteam

v0.3.0 - Auditable research evidence and SQL evaluation

Choose a tag to compare

@ahines99 ahines99 released this 28 Sep 03:09

This release addresses the portfolio audit with stronger semantic, benchmark, domain and live-model evidence while preserving the existing release policy and all earlier results.

  • Versioned reference-grounded semantic advisory with structured claims, equivalent units, citation support, known word-digit disclosures and explicit uncertainty. It cannot override release gates.
  • 120 AI-authored traces across support and SQL, source-frozen first execution, per-dimension precision/recall/F1/coverage, blinded review exports and dual-review/adjudication tooling. All eleven deterministic-scorer disagreements are retained. No independent human validation is claimed.
  • Twelve real SQLite cases exercised through persistent evaluation and release gating. Three controls demonstrate eligible, rejected and critical-block outcomes; installed-wheel/container verification includes the SQL domain.
  • Completed predeclared study: two pinned Claude profiles, three interleaved rounds, 72 invocations, 138 paid requests and twelve actually triggered/recovered failure probes. New estimated token cost $0.484913; cumulative retained reservations $8.323672 under the $12 study ceiling.
  • Sonnet passed 22/30 baseline observations; Haiku passed 24/30 but had two critical synthetic-disclosure findings. All review/block decisions remain intact. The exploratory ten-case comparison does not establish a general model ranking.

Read the audit disposition and full results, roadmap, and verification record.

Validation: local Windows/Python 3.14 passed 392 tests with 16 PostgreSQL-only skips and 95.77% line coverage; Ruff and mypy passed. The exact release CI run verifies Linux Python 3.12/3.13/3.14, Windows Python 3.12, real PostgreSQL, installed wheels, Docker/Compose, process restart and private HTTPS behavior.

The wheel and sdist attached here come from that CI run's Linux/Python 3.12 artifact. The verification ZIP includes source identity, CI job metadata, five JUnit reports, behavioral delivery evidence, frozen benchmark/study inputs and preserved experimental outputs. SHA256SUMS covers all three binary artifacts. Original v0.2.x tags and release assets remain unchanged.

Scope: AI-assisted portfolio engineering with synthetic data, local-first operation and an authenticated shared-server option. No public evaluation SaaS, external adoption, customer-impact metric, independently human-labeled benchmark or production SLA is claimed. The requested five-agent team exceeded runtime capacity; three available subagents were reused across the five documented workstreams with coordinator-owned SQL/integration work.

v0.2.1 - Enforce version-wide release blocks

Choose a tag to compare

@ahines99 ahines99 released this 28 Sep 00:37

Portfolio evidence addendum: the owner explicitly replaced the independent-human benchmark requirement with transparent AI-authored characterization. The 40-trace report preserves three false positives and three false negatives, all first-run evidence, and an AI-inspected walkthrough. Independent human validation has not been performed on this set. Portfolio finalization is complete under the revised scope.

The addendum source is 2f318dc; all six CI jobs passed: 292 tests per OS/Python matrix job plus 16 PostgreSQL tests. Pages and the dependency audit also passed. benchmark-addendum-0.2.1.zip and BENCHMARK-SHA256SUMS contain the separately source-bound evidence. The original v0.2.1 tag, software packages, SHA256SUMS and verification ZIP remain unchanged.

Original corrective software release:

This corrective release fixes approval of an older pending review after another run blocked the same agent version. Upgrade from 0.2.0 to enforce the documented version-wide block consistently.

  • Current release decisions and accepted regression baselines honor all committed blocks of an agent version, including older passing or approved runs. Historical gate artifacts and approvals remain intact; reports identify the effective blocking runs.
  • Gate publication and approval serialize on the agent's database row across SQLite and PostgreSQL. A stale gate computed before a concurrent block is refreshed atomically with its audit hash and pause state.
  • New regressions cover the original bypass, historical acceptance, baseline exclusion, stale PASS/REVIEW publication, corrupt evidence, failed monitoring and both block/approval commit orders.
  • README formatting and portfolio completion accounting are corrected. Existing human approval remains recorded; the original independent benchmark was pending at this software snapshot; see the subsequent scope decision above.

Validation: all six CI jobs passed at the released source revision, with 290 tests passing per OS/Python matrix job and 16 additional PostgreSQL tests. The full local Windows/Python 3.14 run passed all 306 tests with PostgreSQL enabled and 95.55% line coverage. Installed-wheel, container, persistent Compose/MCP, private HTTPS and dependency checks passed.

Release CI | Correction record | Demo

No database schema change or backfill is needed; the corrected implementation reads existing committed gates. This is enforcement of the existing gate-policy/1.2 rule; scoring thresholds and the scorer are unchanged. Previously accepted runs of a blocked version now report an effective blocked decision.

At the original software release, independent human evidence was still pending. The subsequent owner-approved scope change and AI-authored evidence addendum above supersede that portfolio completion status. The existing 33-row/35-label corpus remains separately human-reviewed and development-exposed, with all nine disagreements preserved. The original independence criterion was replaced, not fulfilled.

Assets: wheel, source archive, verification ZIP and SHA256SUMS. The checksum file covers all three artifacts; the verification manifest binds test reports, CI records, operational evidence and benchmark status to the exact source commit. Original 0.2.0 packages and evidence remain available as historical artifacts.

v0.2.0 — Verified local-first portfolio release

Choose a tag to compare

@ahines99 ahines99 released this 27 Sep 23:26

Correction: upgrade to v0.2.1. Version 0.2.0 contains a stale-approval defect in enforcement of version-wide blocks. The corrective release fixes it; its later evidence addendum records completion under the owner-approved AI-authored benchmark scope, without claiming independent human validation. The historical packages, evidence assets and tag below remain unchanged.

This release turns the evaluation harness into a runnable, documented local-first portfolio project with an authenticated shared-server option.

  • Complete trace verification, fenced execution, atomic checkpoints, comparable baselines and authenticated tenant isolation.
  • Passing Linux Python 3.12/3.13/3.14 and Windows Python 3.12 CI; 281 tests pass per matrix job, with eight additional PostgreSQL contracts passing in the database job.
  • Verified installed wheel, Docker behavior, persistent Compose/MCP restart and private HTTPS through Caddy with explicit CA trust.
  • Static 150-second walkthrough, engineering case study, operational checks and real OpenTelemetry Collector evidence.

Interactive demo | Original release verification | Release CI

The first live Sonnet 5 experiment is preserved: 21 paid requests, $0.10417 estimated token cost, 1/10 cases passed, and a review-required gate. These results expose fixture/latency limits and are not a general model-quality score. No further paid runs were needed.

Alexander Hines approved all 33 scorer challenge rows and 35 labels without corrections on September 27, 2026. Nine semantic disagreements remain documented; this release does not claim independently human-validated scoring accuracy or a publicly hosted evaluation backend.

Assets include the wheel, source archive, SHA256SUMS and a verification ZIP containing test reports, coverage, source provenance, GitHub run records and bounded live/operational evidence. SHA256SUMS covers the original wheel, source archive and verification ZIP; those four assets remain unchanged.

Human-review addendum: review sheet and successful review CI. The separate human-review-addendum-0.2.0.zip and HUMAN-REVIEW-SHA256SUMS bind the reviewed corpus, unchanged scorer results, completed review sheet and test reports to commit 517f96c. All six CI jobs pass: 282 tests per OS/Python matrix job and eight additional PostgreSQL tests. The corpus remains development-exposed, and all nine disagreements are preserved. The existing label review is complete. Independent human evidence was pending at this historical snapshot. The owner subsequently approved AI-authored characterization instead; see the current v0.2.1 evidence addendum. Independent human validation is not claimed.