Skip to content

Agentic Evidence Lab v0.1.0-alpha.2

Pre-release
Pre-release

Choose a tag to compare

@ryuhmanov-m ryuhmanov-m released this 12 Aug 11:18
· 4 commits to main since this release

Agentic Evidence Lab v0.1.0-alpha.2

This alpha publishes Agent Skills Season 1 activation calibration: a bounded,
machine-readable answer to whether twelve exact public skill snapshots can run
and activate in the controlled Codex adapter. It is not an effectiveness
leaderboard.

Included

  • ten independent Season 1 protocols and exact source locks for twelve public
    skill trees from four upstream repositories;
  • a healthy ten-task public calibration pack, revisioned through two retained
    benchmark invalidations and one format-only transition;
  • 22 valid final-revision Codex run records, ten measurement sets, and ten
    evidence receipts;
  • explicit activation evidence for ten treatment skills and inconclusive
    activation for Anthropic mcp-builder and webapp-testing;
  • source-lock validation and non-executing caller-checkout verification;
  • exact study-revision resolution in the evidence validator;
  • a public first-wave screening boundary that keeps private screening and
    confirmation packs outside the Git worktree;
  • a release-canary guard against accidental private-pack publication.

Decision

The runner and evidence pipeline are ready for study-specific screening design.
The ten activated snapshots are execution-compatible candidates only. MCP and
webapp activation redesign is deferred to their later study milestones and does
not block the declared three-study first wave.

No evidence in this release establishes that a skill improves correctness,
debugging, testing, security, design, cost, latency, transfer, or production
outcomes. Shared acceptance on one public mechanics task cannot establish
equivalence. No universal or cross-study ranking is published.

Integrity corrections

Activation found two benchmark defects that static pack-health checks missed:

  1. an underdetermined truthful-completion state prefix;
  2. an external-resource false positive on an inline SVG namespace.

The affected raw runs remain private and content-addressed; public invalidation
records explain their exclusion. Final evidence uses only corrected revisions.

Security boundary

The Codex process can read the maintainer's reusable ChatGPT credential. The
owner accepted that property for the exact pinned and reviewed Season 1
snapshots in maintainer-controlled fixtures. This is not a public submission
service: arbitrary or changed third-party content remains blocked pending its
own pin, review, and explicit acceptance or a stronger credential boundary.

Compatibility

The package version is 0.1.0a2; the Git tag is v0.1.0-alpha.2. Contract v0
remains pre-stable. Study references now resolve by exact ID and revision;
duplicate identical revisions fail validation.

Verify the release

Run the Python 3.11–3.13 CI matrix, public evidence validation, source-lock and
manifest checks, public calibration reproduction, Docker isolation smoke,
release-tree scan, package build, and clean-wheel validation against the exact
release SHA. Distribution assets include SHA-256 checksums.