Repository navigation
Jev and System One
Every stage in harness-kit ends at an eval: a judge reads the artifact against a weighted rubric and the stage passes at 8.0. When the judge is a model from the same family that wrote the artifact, the gate measures familiarity along with quality. Jev, a System One model by TypeSafe AI, is the second judge harness-kit offers for that reason. How to switch judges and read the output is on Evals. The research and the model design behind that choice are below.
Zheng et al. set the baseline in Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv:2306.05685, 2023). GPT-4 as a judge agreed with human preferences over 80% of the time, the same level at which humans agree with each other. The same paper measured three biases. Position: with the two answers swapped, GPT-4 gave a consistent verdict in 65.0% of cases, raised to 77.5% with few-shot examples. Verbosity: an answer padded with a rephrased copy of its own list beat the original for Claude-v1 and GPT-3.5 in 91.3% of 23 cases, and for GPT-4 in 8.7%. Self-enhancement: GPT-4 favored its own answers with a 10% higher win rate and Claude-v1 with 25%, though the authors say their data was too limited to confirm the effect.
Later work confirmed it. Panickssery et al., in LLM Evaluators Recognize and Favor Their Own Generations (arXiv:2404.13076, 2024), found that models can tell their own output apart from others', and that the strength of self-preference rises linearly with that self-recognition; fine-tuning a model to recognize itself better made it prefer itself more. Wataoka et al., in Self-Preference Bias in LLM-as-a-Judge (arXiv:2410.21819, 2024), traced a mechanism: judges give higher scores to text with lower perplexity, text that is more familiar to them, whether or not they wrote it.
For a Claude pipeline graded by Claude this is the worst case. The PRD under review is text the judge finds maximally familiar. A fresh evaluator with a clean context removes the writer's reasoning, and the weights that find the prose familiar stay the same.
Verga et al. tested the obvious counter in Replacing Judges with Juries (arXiv:2404.18796, 2024). A panel of three smaller judges from different model families agreed with humans better than a single GPT-4 judge across their datasets, showed less intra-model bias, and cost over seven times less. harness-kit is one step toward that: one judge from a different family, with Claude as the fallback. It is not a panel yet.
The names echo Kahneman's two modes of thought: System 1 is fast and intuitive, System 2 is slow and deliberate. An LLM works in the second mode. It generates reasoning token by token and answers in prose that your code then parses, and the parse can fail.
TypeSafe's System One models do the first kind of work: the judgment a knowledgeable person makes in a second given the right context. You send a state (the content to judge) and a set of typed questions, and Jev returns a probability distribution over the options you supplied. It never produces a value outside them, so the answer cannot be malformed. TypeSafe trains it with a method it calls RLCD to return calibrated probabilities, and the same weights serve every account.
Jev does not write text, call tools or hold a conversation, and the coding agents page says plainly that it cannot replace the model behind Claude Code. The split in harness-kit follows that line. Claude writes the PRD, which is System Two work. Whether sections.success_metrics names a baseline is a System One question.
Every question has an id, a type, and instructions. The id stays in your code and never reaches the model.
| Type | Answer | Read it as |
|---|---|---|
| Noul | noul |
probability that the answer is yes |
| Score |
score, probabilities, confidence
|
position on ordered levels you describe |
| Choice |
choice, probabilities, confidence
|
one option from an unordered set |
A Noul value is the answer and the certainty in one number. The docs show recorded answers to "Is the customer asking for a human agent?": 0.02 for "Thanks, that fixed it!", 0.40 for "Are you a bot?", 0.84 for "Is there any way to speak to someone about my invoice?". A value near 0.5 is the model saying it does not know.
A Score is the probability-weighted mean of the level numbers. For a bug report split between "workaround exists" and "no workaround", the probabilities are 0.0, 0.57 and 0.43, so the score is 0 x 0.0 + 1 x 0.57 + 2 x 0.43 = 1.43. Levels must describe situations: the same report scored 0.55 with bare levels "0", "1", "2" and 0.0 at full confidence with descriptive ones.
A Noul is a probability and not a scale. Asked "Is the candidate strong in Python?", a candidate with two years of daily use got 0.81 and one with eight years got 0.92. The gap measures how sure the model is that "strong" applies, and says nothing about how much more experience the second one has. When you want degree, use a Score.
Calibrated means the probabilities match frequencies: across many answers of 0.8, about 80% should turn out to be yes. That property is what lets code put thresholds on the numbers.
Choice and Score answers come with a confidence from 0 to 1, computed from the returned distribution. A Noul has none, since its single probability already describes a two-outcome distribution.
Noul confidence = |2p - 1|
Choice confidence = (p_max - 1/n) / (1 - 1/n)
Score confidence = max(0, 1 - sum_i p_i |i - m| / MAD_unif)
MAD_unif = (1/n) sum_i |i - (n - 1)/2|
The Noul form is the distance from a coin flip: 0 at p = 0.5, 1 at p = 0 or 1. The Choice form measures how far the top option sits above an even split of 1/n, so 0 is a uniform guess and 1 is certainty; it reads only the top probability, which is why (0.6, 0.3, 0.1) and (0.6, 0.2, 0.2) both get 0.4.
The Score form counts how far probability sits from the most likely level m, measured in levels, and compares that with the same spread for an even distribution. Probability on a neighboring level costs less than probability at the far end: over three levels, (0, 0.5, 0.5) gets 0.25 and (0.5, 0, 0.5) gets 0. The bug report above gets 1 - 0.43 / (2/3), about 0.35. TypeSafe's confidence page treats these as sensible defaults and returns the full probabilities so you can compute another measure if your case needs it.
The Jev docs ask for one snap judgment per question. A Noul that asks two things at once, such as "Is the customer angry and asking for a refund?", forces the model to judge both together, and the value means less. Split it into two Nouls and combine them in code with weights you control.
The eval research reached the same conclusion from the other side. CheckEval (arXiv:2403.18771) decomposes each criterion into binary checklist questions and raised average agreement across evaluator models by 0.45. TICK (arXiv:2410.03608) has an LLM write a checklist per instruction, and judging against it raised exact agreement between LLM judgments and human preferences from 46.4% to 52.2%.
harness-kit applies both. A rubric dimension such as metric completeness used to ask "every metric has baseline, target, horizon? guardrails listed? kill criteria numeric?" in one breath. It now carries one - check: line per condition, and jev-judge.py sends each line as one Noul:
### Metric completeness (weight 20%)
- check: Every metric in `sections.success_metrics` has a baseline value.
- check: Every metric in `sections.success_metrics` has a target value.
- check: `sections.success_metrics` states a kill criterion with a numeric threshold.All checks of a rubric go in one request. Jev evaluates them in parallel, so more questions barely change the response time; TypeSafe's parallel questions cookbook measured 13 questions in one call at 11.5 times cheaper and 9.6 times faster than 13 separate calls, with the same answers.
A text judge emits one score token, the most likely one, and throws away the rest of its distribution. G-Eval (arXiv:2303.16634) kept it: it weights each possible score by the probability of its token and sums them. On the SummEval benchmark, G-Eval with GPT-4 reached a Spearman correlation of 0.514 with human ratings using probabilities, against 0.502 without, and the weighted version also broke the ties that whole-number scores produce.
Wang, Zhang and Choi made the general case in Improving LLM-as-a-Judge Inference with the Judgment Distribution (arXiv:2503.03064, 2025). Across their settings, the mean of the judgment distribution beat the mode, which is what greedy decoding returns. They also found that chain-of-thought prompting can collapse the distribution toward one answer, which removes the information the mean relies on.
Jev returns the distribution directly, and its Score field is already the expected value. jev-judge.py uses the same idea for checks: a dimension scores 10 times the mean p(yes) across its checks, the expected share of checks met. A check at 0.6 contributes 0.6, not a rounded 1. A dimension with no checks falls back to one Score over the rubric's own 0/5/10 anchors, a weaker mode that no weighted rubric in the repo uses any more.
TypeSafe publishes the jagged edges of jev-1.13: it reads instructions literally, does not count reliably, handles numbers and dates as text, loses accuracy with indirection and with large state full of irrelevant detail, can be moved by adversarial content, may lean toward the first option in a Choice, and does not generate text. harness-kit routes around each one it meets.
| Weakness | What harness-kit does |
|---|---|
| Counting, arithmetic |
- absent: regexes run in code; eval-score.py computes the weighted total and applies 8.0 |
| Numbers and dates | checks ask whether a value is present, never compare or compute values |
| Large, distracting state | artifact sent as sections; each check names sections.x; a missing section fails in code with no call |
| Context limit | artifact over about 30k tokens (120,000 characters) escalates to Claude |
| Literal reading | each check states one exact condition, phrased so yes means good |
| Choice option order | no Choice questions are used |
| Generation | feedback is the failed check text, written by code |
Two edges remain. A check like "Every metric in sections.success_metrics has a baseline value" asks Jev to apply a condition across a list, which sits close to the counting weakness when the list is long; splitting it per metric in code would be stricter. Adversarial content is also open: the artifact is model-written text, and a sentence that argues for its own quality can move an answer. Escalation catches the uncertain cases, and nothing catches a confident wrong one.
Jung et al., in Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement (arXiv:2407.18370, 2024), let a judge decide only when its confidence clears a threshold, and pass the rest to a stronger judge. With the threshold calibrated on human labels, the method guarantees a chosen level of agreement with humans on the cases the judge keeps.
harness-kit applies the shape without the calibration. An answer with 0.3 < p < 0.7 counts as unsure. jev-judge.py recomputes the weighted total twice: the low bound counts only answers at 0.7 or above as met, which forces every unsure answer to no, and the high bound counts every answer above 0.3, which forces them to yes. When one bound reaches 8.0 and the other does not, the unsure answers decide the verdict, and the script exits 3 so the Claude evaluator judges the artifact. Unsure answers that cannot flip the result do not escalate.
flowchart LR
A[artifact and rubric] --> J[jev-judge.py]
J -->|no key, over 32k tokens, API error| C[Claude evaluator]
J --> B[low and high bounds]
B -->|both sides of 8.0| C
B -->|same side| S[eval-score.py]
The same exit sends the run to Claude when the judge is set to local, the key is missing, the artifact is too large, or the API fails. Escalated runs are logged with the verdict escalated. The 0.3 and 0.7 bounds are a chosen convention; Jung et al. set theirs from human data, and harness-kit has none yet.
jev-1.13 costs $0.042 per million input tokens, and output tokens are free. A request of 10,000 input tokens costs $0.00042. The context holds 64k tokens per request, of which the state plus the longest single question may use 32k, and the rate limits are 100K tokens per second and 80 requests per second. TypeSafe's docs say most queries complete in about 100 ms, and the launch post gives 70 to 500 ms.
Cheap is still paid. Jev runs on the developer's machine with a key from an env var or the env block of .claude/settings.local.json, and never in CI. The alias jev-latest moves when a new release ships, so the answers behind it can change; every logged run records the versioned model id that answered, and you can pin one with TYPESAFE_MODEL once thresholds are tuned against it.
Neither judge has been measured against human labels on harness-kit artifacts. Jev's calibration is TypeSafe's claim on TypeSafe's data, the 0.3 to 0.7 band is a convention, and 8.0 is a convention. The research above says atomic checks, expected values and a judge from another family each move agreement in the right direction. None of it says by how much on a PRD.
Hamel Husain, in Using LLM-as-a-Judge for evaluation, describes the way out: a domain expert labels examples pass or fail with a written critique, about 100 per failure mode, and you measure how often the judge agrees. Raw agreement flatters a judge when most answers are yes. Two graders who each say yes 90% of the time agree 82% of the time by chance alone (0.9 x 0.9 + 0.1 x 0.1), which is why chance-corrected measures such as Cohen's kappa are the usual report.
Every Jev run appends its per-check probabilities, bounds, verdict and model id to .claude/runtime/outputs/evals/jev-judge.jsonl. Labeling those records turns the log into that set, and until then the failed checks are worth more than the number.
Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023. Panickssery et al., LLM Evaluators Recognize and Favor Their Own Generations, 2024. Wataoka et al., Self-Preference Bias in LLM-as-a-Judge, 2024. Verga et al., Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models, 2024.
Liu et al., G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, 2023. Wang, Zhang and Choi, Improving LLM-as-a-Judge Inference with the Judgment Distribution, 2025. Lee et al., CheckEval, 2024. Cook et al., TICK: Generated Checklists Improve LLM Evaluation and Generation, 2024. Jung et al., Trust or Escalate, 2024. Liu et al., Lost in the Middle, TACL 2024, on why a long distracting state costs accuracy.
TypeSafe AI, Introducing System One models and Jev, and the TypeSafe docs: primitives, Noul, Score, confidence, composite scoring, state, models, and jev-1.13 jaggedness. Hamel Husain, Using LLM-as-a-Judge for evaluation.