Skip to content

Releases: deghosal-2026/agent-self-edit

v0.2.0

Choose a tag to compare

@deghosal-2026 deghosal-2026 released this 02 Sep 08:40

AgentSelfEdit v0.2.0

Release date: 2026-09-02

What's new

  • Staged analyzer pipeline — 4-stage (summarize → select → synthesize → score) replaces the monolithic analyzer
  • Model role separation — executor, analyzer, and judge roles with per-role provider configuration
  • Multi-domain support — classification, extraction, generation, and mixed-domain suites
  • Runtime scorer selection — ExactMatch, StructuredExtractionScorer, LLMJudgeScorer auto-selected per task type
  • Rejection context — staged analyzer receives structured rejection feedback to avoid repeated failed edit families
  • 40-task promotion corpus — replaces the fragile 5-task set for statistical power
  • Hermetic CI suite — zero LLM call tests safe for CI
  • Row-identity trace ack — immutable row-ID based processing with in-flight reservation

Correctness fixes

  • Gate edit-distance measured against full candidate prompt, not fragment
  • Gate drift measured against original prompt baseline
  • A/B alpha semantics: p < (1 - confidence_level)
  • Candidate prompt materialization uses full prompt
  • Staged analyzer rejection context threaded through all prompt stages
  • CLI exit codes deterministic; init writes runnable config; validate checks runnability

Docker

  • Multi-domain Docker tests (classification, extraction, generation, staged analyzer)
  • Multi-stage Dockerfile with reduced image size

Field test

  • Gate FP rate: 0% across all runs
  • Multi-domain suites: 4/4 executed
  • Adversarial edits: 5/5 caught, 0 false negatives
  • Strongest analyzer: Mistral Small 24B (OpenRouter) — directional gains but no promotable edit
  • Conclusion: framework is correct, observable, and safe. Analyzer quality is the bottleneck.

Known limitations

  • Coverage at 81% (target 92%)
  • No verified prompt self-improvement yet — analyzer search quality is the bottleneck

Installation

pip install agent-self-edit==0.2.0

Full changelog: https://github.com/deghosal-2026/agent-self-edit/blob/main/CHANGELOG.md

v0.1.0 — Self-Improving Prompt Optimization

Choose a tag to compare

@deghosal-2026 deghosal-2026 released this 01 Sep 02:53

AgentSelfEdit v0.1.0 Release Notes

Date: 2026-09-01

What's New

AgentSelfEdit is a sidecar that observes execution traces and rewrites its own system prompt through a closed loop:

Agent executes task → trace stored → Feedback Analyzer (LLM) reviews →
A/B test candidate vs current prompt → Promotion Gate (6 deterministic checks) →
promote or reject → versioned Registry

Core Features

  • F-01 Execution trace ingestion — SQLite-backed trace storage with validation, batching, and cleanup
  • F-02 Feedback analyzer — LLM-powered analysis of failed traces with structured prompt-edit proposals
  • F-03 A/B test engine — Paired comparison of prompt variants with bootstrap CI, permutation test, effect size
  • F-04 Promotion gate — 6 deterministic fail-fast checks: sample floor, effect size, confidence, frozen sections, edit distance, drift
  • F-05 Prompt registry — Versioned prompt storage with full lineage, diff, rollback
  • F-06 Frozen core sections — User-annotated sections the analyzer cannot modify
  • F-07 Edit-distance limit — Configurable max lines changed per edit
  • F-08 Prompt diff visualization — Side-by-side and inline diff between prompt versions
  • F-09 CLI — 10 commands: init, run, status, diff, rollback, guardrails, lineage, propose, ingest, validate
  • F-10 Held-out task set management — YAML-based task sets for A/B evaluation
  • F-11 Near-miss logging — Rejected edits logged with guardrail reasoning
  • F-12 Rollback — One-command rollback to any previous prompt version
  • F-13 Config file — YAML/TOML configuration for all thresholds and paths
  • F-14 Docker support — Multi-stage Dockerfile and docker-compose

Field Test Results

The v0.1.0 field test ran 15 iterations of the self-edit loop against a local Qwen3.5-4B-4bit model on Apple Silicon.

Metric Result
LLM calls 4,150
Total tokens 716,580
Wall time 37 minutes
Cost $0.00 (local 4B)
Gate false positive rate 0%
Gate false negative rate 0%
Docker tests 9/9 pass
All tests 443/443 pass

Key finding: The promotion gate works correctly. It rejects edits that produce real but statistically underpowered improvement (p=0.23 >= 0.05). Zero false positives, zero false negatives over 15 iterations.

Known Issues

  • Coverage: 89% (target 92%) — CLI modules tested via Docker not unit tests
  • Improvement: 0% over 15 iterations — analyzer proposes same edit every time, needs rejection feedback
  • Non-LLM hermetic tests: CI-safe tests not yet run in CI

Upgrade Guide

v0.1.0 is the first release. No upgrade path from previous versions.

Installing

pip install agent-self-edit

Quickstart

agent-self-edit init
agent-self-edit run --once
agent-self-edit status

See the README for full documentation.

Credits

Built by Debashish Ghosal. Powered by Qwen3.5-4B-4bit via MLX on Apple Silicon.