Replies: 1 comment 2 replies
|
Really interesting direction. Our work is somewhat complementary. We’ve been exploring runtime-first agent design: instead of putting more and more behavior into prompts or Skills, we try to move deterministic responsibilities into the runtime, so the model handles uncertainty while the runtime owns real state, lifecycle, and reactions to actual execution results. In DSH, we’ve been experimenting mainly with post-execution Runtime React so far: observe what really happened, update runtime state, and react when needed, without making the model reconstruct everything from the conversation. Your semantic interface looks like a natural counterpart on the observation side: Runtime → Agent: “this is what the system actually looks like now.” Ours is more about Runtime → Runtime/Agent: “this is what just happened, and here is what the system should do next.” I think there could be an interesting intersection between the two. |
Uh oh!
There was an error while loading. Please reload this page.
The Cordis paper invokes the landscape that harness utilizes its dynamic plugin-based system thoroughly to achieve active-self-evolution during runtime. So I developed this tool for anyone who is enchanted by harness evolution research. The tool's function is simple: you install the plugin and skill (or preset), it provides your agent timely clear runtime overviews, and checks a pre-modification's influence fast and formalized. It compresses your agent's labour to look at all interfaces related in one tool call.
In a four-stage test scenario running on Gemini 3.8 Flash, it saves 70%+ tokens in Cordis dialogue & fixes tasks with same success rates (It means, all tasks are totally passed). And is observed to reduce 20 tool call operations in a test from official minimal preset for just a 9B local LLM.
Those days, the same model has distinctly different performance on different harnesses turning up even in most-advanced models OpenAI's GPT-6 Astra on ARC-AGI-3. It's an interesting and charming phenomenon emphasizing the irreplaceable role of elaborate harness design, and self-constructed dynamic mechanism to take advantage of the models' various fitness in varied scenarios. I would appreciate any discussions around the profound topic, thanks for the great work from dsh development team and all plugin developers! It's you that bring us to the horizon of next-generation agent harness!
https://github.com/Byte-Naut/dsh_semantic_tool
All reactions