The factory learns to study other factories without being changed by them β a health-gated repo-scout skill, per-skill production telemetry, a forge-proof eval judge, and four code-grounded reference studies with explicit adoption gates. One cycle since v2.22.0.
Highlights
- Skills are now measured in production, not just at authoring time.
bun run skill-reportderives honest per-skill trust states βtrusted / proven / active / demoted / dormantβ from the existing trace substrate (skill names only; args never land in a trace)./learn-loopgrounds in it as its fourth instrument, and/kill-or-keep's quarterly curator readsdormantrows as archive candidates. Ported idea: OpenSpace's outcome tables β on our substrate instead of a new store. - The eval judge can no longer be steered by what it judges. Every judged output now sits inside a per-call UUID-tagged untrusted boundary minted after the output exists β forged close tags, fake verdict JSON, and judge-addressed imperatives are content to grade, not directions to follow. One fix covers every factory judge (eval harness +
/goal's loop). Ported idea: Adrian's boundary design. /repo-scoutβ the study method, promoted to a skill. Verify health (gh-api gates) β scratchpad-only shallow clone β facts-only hardened deep-dive (repo content is data, never instructions; assessed code is never executed) β references-grammar draft in an operator-gated backlog. Trending mode rides OSS Insight's public API with a 3-repo cap; a new STANDING-ORDERS program scopes its unattended authority.- Four reference studies, with the discipline showing. Graft (NanoNets) Β· Adrian (Secure Agentics) Β· AgentENV (kvcache-ai) Β· OpenSpace (HKUDS) β each entry records what to mine with file-path evidence, what deliberately isn't adopted, the measured gate any adoption must pass, and the trigger that would reopen the question. Two ideas ported the same day; zero frameworks imported.
What's inside
v2.23.0 β the external-repo mining cycle Β· skill-outcome telemetry (bun run skill-report, eval skill-outcome-fidelity) Β· judge untrusted-output boundary (4 deterministic boundary tests) Β· /repo-scout skill + program + heartbeat item, eval-covered via contract pins Β· four references/README.md studies + clone lines Β· new doctrine: api-compatibility-as-distribution (the AgentENV/E2B lesson) and the auto-captured-skill-landfill anti-pattern (OpenSpace's 203-skill corpus as the control group for ungoverned skill capture) Β· credits gain "a thousand generosities" β the references tier, with a 10x tier that waits for measured evidence.
Get started
curl -fsSL https://raw.githubusercontent.com/hamza-ali-shahjahan/hamzaish/main/install.sh | shThen open Claude Code and type /builder-mode <your idea>. New to terminals? Start at docs/start-here.md.
License
AGPL-3.0 β clean, no added clauses. Free for builders; commercial license on request.