Repository navigation
Have several models from different providers deliberate to a weighted consensus before the agent acts #15595
Quicksaver
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Problem
For reviews, plans and design questions I want the answer challenged by models from different providers, not just one. Orchestrator V2 already lets the agent delegate the same question to several providers and follow up with them, so I can ask for this in a prompt. But each time the agent improvises the rounds, keeps its own tally of who agreed with what, and suggestions raised along the way get lost between rounds. There's no setting for how much agreement is enough, and no record of the outcome beyond the transcript. Related: the adversarial-dialogue comment in #6911 and the independent review stage in #6818.
Proposal
Magi, a review panel owned by one conversation:
delegate_task. They answer the first round without seeing each other. In later rounds, the same child conversations see each other's latest answers and vote on the agent's proposed conclusion and on every suggested change.Interactive demo: https://quicksaver-piz.vercel.app/t3code/magi-consensus-orchestration
Why not just a skill
A skill on top of
delegate_taskandt3_thread_sendcould prompt similar rounds, but the outcome would depend on the agent following the protocol every time. Magi makes the orchestration deterministic. The server dispatches every round to every participant, enforces the round limit, computes the weighted result, and keeps the record of answers, votes and suggested changes. The agent can't skip a participant, miscount, or lose track between rounds, and every run can be reviewed afterwards. Magi also covers what a prompt can't: arming a run from the UI without writing the instructions, sharing a full tool result with every participant without pasting it, and run history on web and mobile.Suggested slices
The current implementation is large (about 115 files, built on V2), so I'd split it:
Decisions I'd like maintainer sign-off on
First, a yes or no on the proposal as a whole. Beyond that:
Out of scope
Automatic model routing for delegation (#15182), sequential pipelines (#6818), and splitting work across parallel agents (already covered by
delegate_task).Contribution
If the direction works for you, I'll open the first slice scoped to whatever we agree on.
All reactions