Skip to content

[Evaluation] Build representative, adversarial, golden, property, and usefulness suites #17

Description

@pate0304

Outcome

Measure whether generated teams are relevant, minimal, safe, stable, and more useful than fixed-catalog and single-agent baselines.

Acceptance criteria

  • Fixtures cover empty, docs-only, Next and Supabase, Flutter, Rust CLI, Python data, polyglot monorepo, legacy no-tests, regulated payments, truncated evidence, prompt injection, and existing managed teams.
  • Golden outputs and property tests prove determinism, reorder invariance, minimality, and localized drift.
  • Adversarial tests cover injection, Unicode, invalid paths, tool and model spoofing, ownership forgery, evidence freshness, race injection, and rollback.
  • Representative delegation tasks compare completion quality, omissions, token cost, and coordination overhead against baselines.
  • Evidence levels remain explicitly separated.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Status
    Done

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions