Submission snapshot for the micro1 Agentic Workflows Hackathon.
- Agent pipeline (judge + two probes + deterministic verifier) on the Claude Agent SDK, TypeScript
- 10 evaluation runs on 30 human-annotated SWE-bench tasks; every number in the README regenerates from results/
npm run screenfor any task;/fairtaskagent skill (skills.sh, Claude Code plugin, Codex plugin)- Three adversarial code reviews and an evals-methodology audit, all documented in the README
- Start with REPRODUCE.md