Repository navigation
Agents orchestrating across machines, without giving up the human control surface #3427
a-alphayed
started this conversation in
Ideas
Replies: 1 comment
|
Have you investigated any overlap with vscode agents window; which is based on agent host protocol? |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I want to float an idea rather than propose a feature, and I'm genuinely
interested in whether you think it belongs in herdr or somewhere else entirely.
Where multi-agent work seems to be heading. Increasingly the useful unit
isn't one agent in a terminal, it's several with different jobs — one holding
the plan, one implementing, one or two reviewing independently. herdr already
supports this well on a single machine:
agent start,send,readandpane runare enough for one agent to drive others, and the pane model meanseach gets its own context.
Where it stops. That pattern hits a wall the moment the agents need to live
on more than one machine, and the reasons they do are ordinary: a laptop that
sleeps, a box with real cores for builds, an always-on host for long jobs, a
machine that holds particular credentials or hardware. Once the fleet spans
hosts, an orchestrating agent can only see and drive what's on its own machine.
Everything else requires a human to switch sessions and relay by hand — which
removes most of the value of having an orchestrator at all.
The obvious fix is the wrong one. You could solve this with a headless
control plane: a scheduler, an API, agents as jobs. That's a well-understood
shape and I think it's a bad fit here, because it throws away the thing that
makes agent work survivable — watching.
Agent work fails in ways that need eyes on them. An agent sits at a confirmation
prompt nobody answers. A reviewer misreads the task and confidently produces
nonsense. Something loops. A task drifts scope. You catch these by looking at
what the agent is doing and then reaching in — typing a correction, answering a
prompt, killing it. If the orchestration layer is headless, your only window is
logs after the fact, and you find out an hour late.
What herdr already is, that most tools aren't. herdr is simultaneously a
programmatic surface and a human one, over the same object. The pane an agent
drives with
agent sendis a pane I can look at, type into, and take over —with no mode switch and no separate tooling. tmux is a great human surface that
knows nothing about agents; orchestration frameworks are great APIs with no
human surface at all. herdr sits in an unusual spot where both are true at once,
and that duality is what makes supervising a handful of agents actually workable.
So the idea is just: extend that property across machines. Not remote
execution, and not merging machines into one pretend workspace — but letting the
runtime span a fleet while the human keeps a single control surface over all of
it. An orchestrating agent addresses any agent anywhere through one API; a human
sees them all in one place and can drop into any of them. The same duality,
one scope wider.
Concretely, the properties that seem to matter:
persistence. The local node aggregates metadata and routes commands; it does
not own remote runtime.
so an orchestrating agent sees one address space instead of per-machine
special cases.
it's the part a headless design gives up.
own. It needs the hosts to be reachable over SSH; SSH is already the
authentication. How you get that reachability — a mesh VPN, a bastion, a
private subnet — is left to the operator. This is squarely a single-operator
design: whoever runs the local node can drive every agent on every host it
reaches, and it isn't trying to be multi-tenant.
Equally important is what it deliberately isn't, since "multi-machine" can
mean an enormous amount of surface area and most of it I think herdr should stay
out of:
orchestrating agent already knows which host it wants.
the host's own herdr wouldn't run for a local caller.
hosts keep their own workspaces and the local view is an aggregate, not a
union.
there's no broker, no daemon, and nothing new listening.
and no reconciliation between machines, because nothing is ever owned twice.
That last one is the constraint that keeps the whole thing small. Most of the
hard problems people associate with distributed systems come from shared
ownership, and there isn't any here.
I have this running. I've been using a private fork along these lines daily
for a few months across a handful of machines, so I can speak to how it behaves
in practice rather than just on paper. It's substantial — roughly 14k lines
across files that don't exist here — and I'm not asking anyone to adopt that.
I'm asking whether the direction is one herdr wants to go, because that
determines whether it's worth me writing up the design properly or just
continuing to run it privately.
The one real UX problem I hit, and I think it's worth naming because it's
the honest weak spot: typing latency to a geographically distant host. Locally
and on nearby machines it's a non-issue, but on a link with ~175 ms RTT every
keystroke you type into a remote-attached pane waits a full round trip to echo.
Everything else about the model held up under daily use — state stayed
consistent, agents didn't get lost, failures were recoverable — but that one is
felt constantly.
It's a well-understood problem with a well-understood answer: Mosh-style
speculative local echo, where the client predicts the echo and reconciles when
the server's version arrives. Worth flagging that it isn't a small addition in
herdr's current shape — in remote-attach mode the client is essentially a byte
pump with no local terminal state to predict against, and the protocol has no
input sequence numbers or echo acknowledgement, both of which that technique
depends on. I mention it not as a request but because it seems like the natural
thing that would need solving if the runtime is ever expected to span long
links, independent of anything else here.
Two things I'd genuinely like opinions on, regardless of whether any code
ever changes hands:
herdr should have, or is one-machine-per-server a deliberate boundary you
want to keep? Both are defensible; I'd rather know which you intend.
correctness of remote mutations — timeouts, ordering, and what you can
truthfully tell a caller when a request may or may not have landed on the far
side. I'd be glad to write up what I learned there even if nothing else comes
of it, because I think anything doing remote I/O in the request path runs
into the same wall.
Entirely fine if the answer is "interesting, not for core." That's useful to
know and I'll stop wondering.
All reactions