Outcome
Measure whether generated teams are relevant, minimal, safe, stable, and more useful than fixed-catalog and single-agent baselines.
Acceptance criteria
- Fixtures cover empty, docs-only, Next and Supabase, Flutter, Rust CLI, Python data, polyglot monorepo, legacy no-tests, regulated payments, truncated evidence, prompt injection, and existing managed teams.
- Golden outputs and property tests prove determinism, reorder invariance, minimality, and localized drift.
- Adversarial tests cover injection, Unicode, invalid paths, tool and model spoofing, ownership forgery, evidence freshness, race injection, and rollback.
- Representative delegation tasks compare completion quality, omissions, token cost, and coordination overhead against baselines.
- Evidence levels remain explicitly separated.
Outcome
Measure whether generated teams are relevant, minimal, safe, stable, and more useful than fixed-catalog and single-agent baselines.
Acceptance criteria