Releases: markpollack/agent-judge
Release list
Agent Judge 0.17.0
Agent Judge 0.17.0
Teaches the result format to say "this criterion does not apply here" without that being
mistaken for "I could not decide", to name the cause of every instrument failure in a countable
way, to stop charging its own failures to the subject, and to record enough structure that a
stored composite result can be read correctly on its own.
This is one coherent change to the result format. It is a breaking change for every Java consumer
and for every reader of stored results.
⚠️ A requirements judge with no requirements no longer passes
EarsJudge.create and Rfc2119Judge.create reject an empty roster, and the rollup refuses one
independently.
A roster is a denominator, and a conjunctive rollup over an empty one is vacuously true: nothing
failed, nothing was unestablished, so the implementation "satisfies the specification". Released
0.16.0 returns PASS for that, and the result is indistinguishable in every stored field from a
specification that was genuinely met.
A stored PASS with a criteriaTotal or constraintsTotal of 0 identifies a run affected
before this fix. Parsing is unchanged: a document that matches nothing still yields an empty list,
because that is a fact about the document. The judge is what refuses.
NOT_APPLICABLE is its own status
Five statuses now, and the difference between the last three is what a denominator does with them:
| Status | The question | A denominator |
|---|---|---|
PASS |
asked, answered yes | counted, in the numerator |
FAIL |
asked, answered no | counted, against the subject |
ABSTAIN |
applied, undecided | counted; no vote cast |
NOT_APPLICABLE |
should not have been asked | excluded, and counted separately |
ERROR |
never reached | excluded from the subject denominator |
ABSTAIN no longer means "not applicable" anywhere in the documentation, and the "not applicable"
wording has been removed from every strategy Javadoc that used it to justify excluding abstentions.
Missing evidence is still ABSTAIN.
fromWire stays exact and throws on an unknown name, and the compile break in exhaustive switches
over JudgmentStatus is intended.
Exclusion must be declared in advance
Exclusion is the one outcome that removes a judge from its own denominator, which makes it the one
outcome an instrument could use to dodge a question it does not like the look of. So:
JudgeMetadatagainsnotApplicableWhen: the condition under which this judge may exclude.Judges.notApplicableCapability(Judge)is the one lookup, and it walks wrapper chains outward in,
so a deduplicating rename cannot strip a capability off the judge underneath.- A
SimpleJuryrefuses a capable seat under a strategy that refuses exclusions, andJuries.meta
refuses a possibly-excluding member, both at construction. - At vote time, an exclusion from an undeclared seat becomes
ERROR undeclared_not_applicable, a
judge-origin error the configuredErrorPolicygoverns. Jury.aggregateMayBeNotApplicable()is the conservative bound a parent checks; it defaults to
false, so an opaque jury makes no pre-spend guarantee and is checked at runtime.
NotApplicablePolicy
REFUSE (the default), EXCLUDE, TREAT_AS_FAIL. Every built-in strategy takes one, and
VotingStrategy.notApplicablePolicy() declares it so composition can be validated before anything
runs.
Conditional criteria in the requirements judges
EarsCriterion and Rfc2119Constraint gain an optional applicability clause. The prompt offers
NOT_APPLICABLE only for criteria that carry one, naming them; an exclusion of an unconditional
criterion, or one with no reason, is a protocol error rather than a finding, and is never counted as
an authorized exclusion. An excluded criterion is not a Check. Both judges declare the capability,
naming the ids.
Every ERROR names its cause
Judgment gains reasonCode, a closed vocabulary in two families:
- instrument codes —
judge_failed,judge_metadata_unreadable,judge_reported,
undeclared_not_applicable,errors_propagated,not_applicable_refused,aggregation_failed,
stage_failed,no_tier_decided— required on everyERROR; - the subject code
subject_empty— optional on aFAIL. AFAILwith no code is an uncoded
rejection, by deliberate choice.
errors_propagated must carry its origin: a non-empty errorCodeCounts block naming the terminal
causes it propagated, flattened through nested wrappers. Judgment.propagatedError builds it
atomically, and construction refuses the code without it.
Machinery failure never supplies rejection evidence
An error from a configured judge is the error policy's business. An error from the library's own
composition or reduction is not, and is never converted into a failing contribution — not even
under TREAT_AS_FAIL, which propagates instead. Charging a broken reduction to the subject would
produce a rejection indistinguishable from a real one in every stored field.
Containment, and one cascade rule
Why this matters, measured rather than argued. In one stored evaluation corpus, 20 of 39 runs
had a jury that errored or abstained — and every one of them is recorded as a subject that did not
pass. The instrument failed and the subject was charged for it, silently, in more than half the
corpus. That is the defect this section removes.
- A strategy that throws, returns null, or returns an aggregate it was not entitled to produce now
yields anERROR aggregation_failedverdict marked undecided, with every judge's own result
intact. Previously it discarded the whole jury and, inside a cascade, the enclosing tier.
Erroris not caught. - A meta-jury member that throws, returns an undecided verdict, or excludes itself without declaring
it may, is a stage failure: excluded from strategy input, its actual verdict kept on the
attempt, the aggregateERROR stage_failed. - A cascade tier that did not produce a usable determination but had already established a genuine
individualFAILstops —REJECT_ON_ANY_FAILonly.ACCEPT_ON_ALL_PASSescalates past it,
because a broken reduction cannot demonstrate that a subject is fine. - When it stops that way, no
FAILand no score is manufactured. The rejection is carried by the
decision; the root aggregate stays a machinery error.
A judge whose metadata cannot be read
A JudgeWithMetadata whose metadata() returned null or threw used to escape SimpleJury.vote():
the seat's name was read outside the failure handling, so every other judge's result was lost and an
enclosing cascade recorded the tier as JURY_EXECUTION_FAILED.
The jury now reads every seat's key once, on the caller's thread and before any judge runs. Such a
seat becomes ERROR judge_metadata_unreadable, keyed Judge#N, naming its position and the failure,
for the ErrorPolicy to resolve — and the judge does not run, so it cannot exercise a capability the
jury was never built with. describe() refuses the seat loudly and by name, including a NamedJudge
whose wrapped judge's metadata cannot be read, and Juries.fromJudges fails at construction naming
the position.
Configuration is refused at construction, not at the first vote
Cheap configuration mistakes are now build-time errors. Each of these used to fail later, after
judges had already run, or not at all:
| Was | Now |
|---|---|
null ErrorPolicy — every aggregation threw NullPointerException out of vote(), even when every judge passed |
IllegalArgumentException("errorPolicy must not be null") from all ten strategy constructors |
null TiePolicy — failed only at the first tie, so a jury could run for months and then break on one input |
MajorityVotingStrategy refuses it at construction |
null NotApplicablePolicy |
refused at construction |
A NaN or infinite judge weight — NaN < 0 is false, so the sign check let both through |
SimpleJury.Builder rejects any non-finite weight |
An invalid threshold is still reported first. The no-argument constructors already passed
PROPAGATE (and TiePolicy.FAIL) and are unchanged.
This is a behaviour change, not only a message change: a jury that was constructed with a null
TiePolicy and never tied used to build and run. It now fails to build.
Weighted average survives an overflowing weight total
WeightedAverageStrategy threw when finite weights summed past Double.MAX_VALUE: the total became
Infinity, the score NaN, and the infinite evidence value was then refused by Judgment
construction.
Scores lie in [0, 1] and weights are non-negative, so any overflow shows up as an infinite total,
and only then are the eligible weights rescaled by a power of two before averaging. Every total
that fits runs the original arithmetic, so results that worked before are bit-identical — pinned by
a table captured from the unmodified sources and by a seeded differential test against the old
arithmetic.
Recorded residual: an overflowing total is reported in the aggregation evidence as
Double.MAX_VALUE. inputWeight and eligibleWeight are therefore a saturating view, not an exact
sum, and a reader must not treat Double.MAX_VALUE there as a measured total.
A jury can be described before it votes
A verdict records what a jury did; nothing recorded what it was configured to do, so a jury that
scored with fewer judges than it lists could not be caught by comparing the two.
Jury.describe() and VotingStrategy.describe() return that structure before any vote, from a new
@NullMarked description package. SimpleJury, CascadedJury, MetaJury and all seven strategies
describe themselves; a consumer jury or strategy that does not override is described as opaque or
undeclared, never as empty. Judges.describe(Judge) looks through NamedJudge — which now
exposes `delega...
Agent Judge 0.16.0
Agent Judge 0.16.0
Adds requirements judges that run a written specification back against the implementation,
two non-compensatory aggregation strategies, and stops the file judges passing on things they
never examined.
⚠️ Behaviour change — the file judges no longer fail open
FileExistsJudge and FileContentJudge previously returned PASS for inputs they had not
actually examined. Each of these is now refused:
| Input | Before | 0.16.0 |
|---|---|---|
| empty path | could pass | refused |
. |
could pass | refused |
| a directory | could pass | refused |
| absolute path outside the workspace | could pass | refused |
parent traversal (../) |
could pass | refused |
If you have suites quietly passing on a mistyped or empty path, they will start failing. That is
the fix working.
This is the failure mode judges are most prone to, and it is worth stating plainly: a JUnit test
has an oracle, but a judge is the oracle. Nothing checks the checker. A test that cannot run
reports a failure; a judge that examines nothing still returns a status, and PASS on an
unexamined thing is indistinguishable from PASS on a verified one.
Requirements judges — agent-judge-ai-core
New. These take a document somebody wrote before the code existed — acceptance criteria,
architectural constraints — and answer whether the implementation satisfies it.
| Type | Purpose |
|---|---|
EarsCriterion |
One acceptance criterion, parsed from the specification |
EarsJudge |
Answers every acceptance criterion |
Rfc2119Constraint |
One architectural constraint, with its RFC 2119 keyword |
Rfc2119Judge |
Answers every architectural constraint |
Observation |
Non-binding evidence noticed while establishing a judgment |
The document supplies the roster, and that changes the aggregation. Elsewhere a jury samples a
population and an ABSTAIN is dropped as not applicable. Here the roster is fixed: the document
says the requirement applies, so "could not be established" is not "does not apply."
any ERROR -> ERROR
else any FAIL -> FAIL
else any ABSTAIN -> ABSTAIN
else -> PASS
ABSTAIN, not PASS. The specification
has not been shown to hold — a different statement from having been shown to fail, and both
different from success. Know this before composing these with anything else: a strategy that drops
abstentions will silently convert unverified into fine.
Both judges also now report what they actually did at INFO, so a run is auditable without a
debugger.
Two non-compensatory strategies — agent-judge-core
AverageVotingStrategy is compensatory: a high score on one criterion offsets a low score on
another. Right when criteria trade off — a slower solution that is markedly clearer may be better.
Two additions for when they do not trade off:
ConjunctiveStrategy— reduceseffectiveScore()withmin. A single low assessment decides
the aggregate and nothing lifts it. Correctness is not offset by elegance.AllMustPassStrategy— the same conjunction over outcomes. Every applicable judgment must
bePASS; a mixed jury is a rejection. Use it as a gate, where the question is admissibility
rather than quality.
AggregationEvidence records how a strategy reached its result.
Averaging a jury whose criteria are independently necessary is the most common way a rubric
reports a healthy number for unusable work. If any single criterion failing should sink the
result, an average will not say so — it reports the mean and looks reasonable doing it.
Judges combinators
The Boolean boundary of the Judges combinators is pinned rather than widened, so composing them
no longer quietly changes how non-Boolean statuses are treated.
Documentation
A practitioner's manual for constructing judges, covering the decisions a default otherwise makes
for you: the question, the denominator, how the pass mark is derived, compensatory versus
conjunctive aggregation, and what gets persisted.
One rule from it worth repeating here: aggregate at read time, never at write time. Parts cannot
be recovered from an aggregate — once you store the mean, the individual judgments are gone.
Dependencies
agent-client moves 0.29.3 → 0.30.0, the current GA. No other dependency changes.
Agent Judge 0.15.2
⚠️ Correction, 2026-08-27 — the affected range is wider than this note originally saidThis note originally scoped the defect to 0.15.0 and 0.15.1. That was wrong. The affected code
is byte-identical in every release from 0.11.0 through 0.15.1, verified by blob hash on
TextSources.java, and all six are published on Maven Central. The uncontained-throw path in
SimpleJuryis older still, present from 0.9.2.If you followed the original guidance you re-checked two releases and left four in place.
If you have collected scores with a parallel jury on any release from 0.11.0 through 0.15.1,
treat them as suspect. Upgrade to 0.15.2 or later.The description of the defect below is accurate; only the version range was understated.
Agent Judge 0.15.2
Correctness release. If you have collected scores with any release from 0.11.0 through 0.15.1, read this.
A jury could silently score with fewer judges than it listed
In every release from 0.11.0 through 0.15.1 a jury could drop a judge from the vote and still return a verdict. The
verdict did not fail, did not warn, and still named every configured judge — but the aggregate was
computed from the judges that survived. Any score produced by a parallel jury on any release from 0.11.0 through 0.15.1
may have been computed from fewer judges than it reports. Treat those scores as suspect and
re-run anything you are relying on.
Two independent defects combined to produce it.
Prompt templates were resolved on the rendering thread. TextSources.classpath(...) returned a
source that read the resource lazily, through the calling thread's context classloader, at render
time. A parallel SimpleJury renders on a ForkJoinPool.commonPool() worker, whose context
classloader is the system classloader rather than the application's. A classpath template that
loaded perfectly well from application code could therefore fail to load inside the jury — under
mvn exec:java, in a container, or anywhere the application's resources are not visible to the
system classloader.
A judge that threw took its jury down with it. SimpleJury let the exception escape, so one
failing judge discarded every other judge's result in the same jury, and inside a CascadedJury
collapsed the whole tier. Combined with the first defect, whether a judge counted depended on which
thread it landed on.
This was reported twice independently, once from inside the project and once downstream, where two
of six judges turned out never to have run.
What changed
TextSources.classpath(...)andTextSources.file(...)now read their text eagerly, on the
thread that calls the factory, and the returnedTextSourcereplays it. Rendering no longer
depends on any thread's context classloader. Classpath resolution tries this library's own
classloader first, then the context and system classloaders, covering container and child-first
arrangements.- A missing classpath resource or unreadable file now fails at construction, on the caller's
stack, next to the configuration that named it — instead of inside a judge on a pool thread. - Template text is frozen at construction, so a file edited mid-run can no longer have two
judgments in the same run made against different prompts. SimpleJuryconverts a judge that throws, or returns no judgment, into anERRORjudgment
naming the judge and the cause, in both parallel and sequential modes. Every configured judge is
represented in the returnedVerdict, and the voting strategy'sErrorPolicydecides what an
error means.Erroris deliberately still not caught.- The count that actually voted is readable from the
AggregationEvidenceblock on the aggregate
judgment: submitted, eligible, errored, and errors treated as abstain.
Upgrading
Drop-in for 0.15.1. No API changed and no signature moved.
Two behaviour changes are worth knowing about, both of them the fix:
- A template naming a resource that does not exist now throws from
fromClasspath(...)/
fromFile(...)rather than fromrender(...). If you were building templates eagerly and
rendering them lazily against resources you expected to be absent, you will see the failure
sooner. - A throwing judge no longer propagates out of
Jury.vote(...). It becomes anERRORjudgment.
Under the defaultErrorPolicy.PROPAGATEthe aggregate is anERRORverdict you can read,
with the working judges' results attached, rather than an exception.
Agent Judge 0.15.1
Agent Judge 0.15.1
Maintenance release aligning Agent Client and compatible Spring, Jackson, test, and analysis
dependencies and publishing consumer-correct CycloneDX SBOMs for every artifact.
Release 0.15.0
Agent Judge 0.15.0
Agent Judge 0.15.0 is a dependency-alignment and release-reproducibility update. It preserves the
0.14 normalized Judgment contract, the existing ten published modules, and the Java 21 baseline.
Consumer dependency floors
- Standalone consumers continue to receive Jackson 2.21.6 where Jackson 2 applies.
agent-judge-llmandagent-judge-ragnow export the accepted Jackson 3.1.6 floor through their
flattened POMs. This prevents a no-parent, no-AgentWorks-BOM consumer from falling back to the
older Jackson 3 versions selected transitively by Spring AI 2.0.0.- The provided Agent Client integration baseline advances from 0.25.0 to 0.26.0. Applications still
select and provide their Agent Client runtime.
Agent Sandbox alignment
agent-judge-exec now consumes agent-sandbox-core 0.10.0. That Core release avoids disclosing
environment values from ExecSpec.toString() and stops forwarding the entire parent process
environment when caller overrides are present.
Agent Judge's own three execution paths construct ExecSpec values with only a command and timeout;
they do not set environment variables, so those disclosure paths were not reachable from Agent Judge
itself. The update protects consumers that use the exported Sandbox API with their own environment-
carrying specs. Agent Judge does not add or provide the Agent Sandbox Docker or E2B modules.
Reproducible release infrastructure
- Build and snapshot workflows are pinned to the reviewed immutable Build Tools commit
35297f1ade5f47c2925d6dab42a7e2d43bd734d0. - The Build workflow can be dispatched manually, allowing exact-SHA hosted verification before a
release. - Unused non-Central repositories and the ineffective scheduled security workflow were removed.
Release security evidence uses the validated, immutable Trivy database snapshot and actual
standalone-consumer runtime JAR closures.
Compatibility and license
No Agent Judge public API is removed or changed in this release. All ten modules remain licensed
under the project-specific Business Source License terms in the repository root LICENSE file.
Release 0.14.0
Agent Judge 0.14.0
Agent Judge 0.14 replaces the former sealed Score hierarchy with a normalized Judgment contract built for portable evaluation results.
This is a breaking pre-1.0 release; consumers must recompile and follow the 0.13 to 0.14 migration guide.
Highlights
Judgmentnow carries requiredstatusplus optional normalizedscoreand optionallabelas independent facts.PASS,FAIL,ABSTAIN, andERRORhave explicit semantics. Mixed applicable consensus votes aggregate toABSTAIN; downstream gate or tier policy decides whether disagreement rejects or escalates.- All voting strategies expose an
ErrorPolicy, defaulting toPROPAGATE, and aggregate results include structured population and error-accounting evidence. - Judgment metadata is validated and recursively frozen at construction as ordinary JSON-compatible values. Rejected values report their exact metadata path.
- Result timing uses the portable integer metadata key
elapsedMillis;Judgment.elapsed()remains the JavaDurationview. - Model usage records optional provider-reported input, output, reasoning, cache-creation, cache-read, and total token quantities. It stores no volatile price estimate.
agent-judge-coreno longer depends on Reactor. It remains framework-neutral while declaring its actual Jackson Databind, SLF4J API, and JSpecify dependencies.- JSpecify makes
scoreandlabelvisibly nullable in the public result API, with NullAway enforcing the adopted package during compilation. - Composite results now expose complete ordered named attempts through
Verdict.compositeAttempts().
Each attempt contains either its returned childVerdictor the code-only failure
jury_execution_failed;CompositePaths.flatten(...)derives stable RFC 6901 paths for nested
results. MetaJuryrecords and continues after member execution exceptions. If any member fails, its
strategy is not invoked and the root isERROR, while successful evidence and every ordered
success/failure attempt remain available.
Removed API
The complete io.github.markpollack.judge.score package and ReactiveJudge are removed.
Use outcome-specific Judgment factories/builders, effectiveScore() where a numeric PASS/FAIL view is intentional, and a runtime-specific async wrapper when needed.
The unreleased composite projection subVerdicts is deleted rather than deprecated. Composite
consumers must move directly to compositeAttempts() or CompositePaths.flatten(...) and handle
the verdict-or-failure outcome explicitly. See MIGRATION_0.14.md.
Samples
The Agent Judge Tutorial contains the canonical executable examples.
Ten credential-free Maven modules cover the normalized API, jury behavior, model-backed judges, and evaluated-side Koog and LangChain4j bridges.
Dependencies
Use io.github.markpollack:agent-judge-*:0.14.0 for every Agent Judge module.
Runtime bridge dependencies remain provided where practical, so applications choose their framework runtime versions.
License
Agent Judge 0.14.0 is licensed under the project-specific Business Source License terms in the repository root LICENSE file.
Release 0.13.0
Release 0.11.0
Full Changelog: v0.10.0...v0.11.0
Release 0.10.0
Full Changelog: v0.9.2...v0.10.0
Release 0.9.2
Full Changelog: v0.9.1...v0.9.2