-
Notifications
You must be signed in to change notification settings - Fork 0
Evals
Evals are the inferential, LLM-based feedback control. Where a sensor checks structure deterministically, an eval judges meaning: is this PRD clear, is this design's trade-off real, does this plan actually follow from the PRP. An eval is an LLM-judge scored 0-10 against a rubric, with a pass threshold of 8.0 and up to 3 retries.
No regex can tell you a problem statement is vague, a hypothesis is unfalsifiable, or a design hid its hardest trade-off. That is semantic judgment, and only a language model can do it at this granularity. The cost: it is probabilistic, so you calibrate it with a rubric and weights rather than trusting a bare "rate this 1-10".
A rubric is a markdown file: weighted dimensions, each with a 0/5/10 anchor, a threshold, and a strict output format.
# Eval: Design Quality
Type: LLM-judge
Mode: quality gate
Threshold: weighted total >= 8.0
## Rubric
### Requirements rigor (weight 15%)
Numeric targets + back-of-envelope math?
- 10: numeric targets + sizing math
- 5: some numbers, no math
- 0: prose only
...
## On failure (total below 8.0)
Retry. Regenerate lowest-scoring sections only. Max 3 attempts.
## Output format
{ "scores": {...}, "weighted_total": 0.0, "feedback": ["dimension: issue with line ref"] }weighted_total = sum(score x weight%) / 100. The judge must cite a line or section when it scores a
dimension below 7, so the feedback is actionable, not a vibe.
generate -> sensors (hard gate) -> eval (score)
^ |
| below 8.0 v
+------ regenerate weak sections ---+ (max 3 attempts, then blocker)
Critically, the retry regenerates only the low-scoring dimensions, not the whole artifact, and the judge's per-dimension feedback says which. After 3 failed attempts the stage returns a blocker rather than shipping something sub-threshold.
The weights are where you express priorities. In the PRD-quality rubric, metric completeness and clarity carry the most weight because a PRD lives or dies on a crisp problem and measurable success. In the design-quality rubric, architecture + deep dives and trade-off discipline dominate because that is what separates a staff design from a box diagram. Tuning weights is a "human on the loop" action: you change what the gate cares about once, and it applies to every future artifact.
An eval that always passes is worthless. Two practices keep it honest:
- Anchored scales. Each dimension defines what 0, 5, and 10 look like, so the judge is not guessing what "7" means.
-
Adversarial framing where it matters. The system-architect's
design-review-deptheval scores a review and explicitly fails one that rubber-stamps: a review with no named gaps, flat severity, or generic "consider improving" advice scores low by design.
| Stage | Evals |
|---|---|
prd |
prd-quality, prd-readiness
|
prp |
prp-quality, prp-context-readiness
|
plan |
plan-quality |
dev |
dev-quality |
pr |
pr-quality |
| system design |
design-quality, design-review-depth
|
The spec-driven loop (/sse:sdd) uses a different eval shape. Instead of a 0-10 score, spec-satisfied
returns PASS/FAIL against the PRP's Success criteria (verifiable) and Validation gates. A
FAIL re-enters the dev↔test loop with a next_iter_focus hint; the loop runs at most 3 iterations.
It runs in a fresh session with no worker context, so the judge is not biased by the implementer's
narrative. This is an eval used as a goal predicate rather than a quality score.
Order matters: sensors run before evals. There is no point spending an LLM call scoring the prose of an artifact that is missing half its sections. The deterministic gate clears the cheap, objective failures first; the eval judges only well-formed artifacts.
- Sensors — the deterministic gate that runs first
- Guides — feedforward that reduces how often evals fail
- Pipeline and Stages — where evals sit in a stage
The harness
- Harness Engineering
- References
- Guides
- Sensors
- Evals
- Pipeline and Stages
- Golden Path
- Agents
- Agent Pipelines
- Designer Skill
v5: autonomy
System design