Autoresearch Sentinel: a controller-driven experiment loop beyond LLM training #645
Shamanbenny
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I built Autoresearch Sentinel, a fork inspired by autoresearch’s evaluate-and-keep-or-reject loop. I wanted to use that loop with projects whose progress can be measured by a repeatable command, such as a benchmark, simulation, or test-based evaluator.
The main change is a Python controller that runs one bounded agent attempt at a time. The controller creates an isolated workspace, runs the configured build and evaluator commands, checks an edit allowlist, records the result, and decides whether to retain the candidate. It also records enough state to continue after an interruption. The current agent adapter uses Codex.
There is a CPU-only maze pathfinder example with the evaluator, experiment records, plot, and recorded results. On its fixed maze set, the recorded baseline explored 1,930 nodes; the latest approved version explored 577 while the evaluator checked route correctness.
I’d welcome feedback on the controller design and on which other measurable tasks would make useful examples.
I'm hoping for contributors to look into how they can reshape Autoresearch Sentinel to support other CLI models like Claude Code or Gemini-CLI, while keeping Codex entirely usable as is. I think there's alot of potential here in removing that unpredictability of the Autoresearch paradigm and making it more stable and consistent by incorporating Sentinel as its controller!
All reactions