Make every agent's “done” earn a PASS.
ToppleCat is an open-source verification tool for Java/JUnit teams that let AI coding agents implement features but keep acceptance with a person. After the agent says done, ToppleCat rechecks the human-confirmed business rules from several independent angles and reports what happened in this run.
ToppleCat calls those runnable rules the Executable Contract. It records a
current-run PASS only when every required Gate passes, then gives the person
responsible for acceptance the evidence behind that result.
Ordinary Java acceptance tests and typed JSON or YAML case rows are the source of truth. Generated JSON and HTML explain what was checked; they never become a second specification.
A checkout Spec says that a 1,000-dollar order receives a 100-dollar discount.
The public example checks 1,000 -> 900, but an implementation could hard-code
900 and still pass it.
A Reviewer can prepare additional evidence before handing the work to the agent:
- reviewer-controlled Typed Case Rows (Hidden Tests), such as
2,000 -> 1,900or999 -> no discount; - a Property such as “the payable total is never negative”; and
- managed Mutation Testing that asks whether the public Acceptance Method notices a changed boundary or arithmetic operation.
After implementation, ToppleCat runs each enabled safeguard independently and
shows all results together. A selected AC whose public acceptance lets an
attributed altered program pass fails Mutation Testing and the aggregate run;
only when every required Gate passes does ToppleCat record PASS for the
current run. The human reads the evidence and makes the delivery decision.
ToppleCat cannot infer a missing rule, such as a VIP discount that the Spec
never stated.
human selects the Spec and prepares the executable contract
-> Spec Review: confirm what will be checked
-> AI agent implements with ordinary ./gradlew test feedback
-> toppleCatVerify: produce fresh formal evidence
-> Verification Report: show the current verdict and AC evidence
-> human decides what happens to the delivery
Both HTML reports are for the human Reviewer. The Implementation Agent receives the public contract and safe Gate-level feedback, never reviewer-owned cases or either HTML report.
The public project page is for human visitors, not an Implementation Agent handoff. Its clearly labelled red-team demonstrations may show fully synthetic report details to explain a safeguard; they are never evidence, feedback, or an approval for an actual delivery.
ToppleCat supplies commands, evidence, and reports. The team decides who runs them and whether they run locally, in CI, or in another workflow.
| Safeguard | Question | Gate |
|---|---|---|
| Hidden Tests | Do reviewer-controlled Typed Case Rows pass the same public Acceptance Method? | REVIEWER_JUNIT |
| Property-Based Testing | Does a human-approved invariant survive bounded generated inputs? | PROPERTY |
| Mutation Testing | Does each exact public Acceptance Method detect the managed-profile mutants it covered? | MUTATION |
The safeguards are independent: one never supplies evidence for another, and their results are not blended into a quality score. Contract Integrity confirms that the complete contract and verification policy still match the Mechanical Seal. Expected Consumption separately checks that authored expected values were actually asserted.
Formal Mutation Testing uses ToppleCat's fixed, versioned PIT profile and preserves PIT's official outcomes. Project-specific PIT tasks remain separate and never enter ToppleCat evidence. See the verification guide.
When an AC's unchanged public acceptance misses an attributed mutation, its
report section shows the detected and undetected counts followed by cards for
only the undetected changes. The cards explain the change, source location, and
that the acceptance still passed; exact replacements are shown only when the
PIT description and unambiguous source line support them. These reviewer-only
details stay out of safe agent feedback. A compact ⓘ control beside Mutation
Testing and its unfamiliar card terms gives a short explanation in the selected
report language. Hover or focus reveals it; clicking or tapping keeps it open
until the Reviewer dismisses it. The visible result and the existing technical
disclosures remain readable without opening help.
ToppleCat 0.2.0 requires JDK 21 or 25 and a compatible Gradle version. The published artifacts target Java 21 and are built on JDK 25. A consumer project may target Java 17, 21, or 25, but the Gradle/plugin execution environment must run on JDK 21 or 25. JDK 17-only ToppleCat execution is unsupported. The official artifacts are available from Maven Central, so consumer projects do not need to clone or build this repository.
settings.gradle.kts:
pluginManagement {
repositories {
gradlePluginPortal()
mavenCentral()
}
}
dependencyResolutionManagement {
repositories { mavenCentral() }
}build.gradle.kts:
plugins {
java
id("io.github.samzhu.topplecat") version "0.2.0"
}
dependencies {
testImplementation("io.github.samzhu.topplecat:topplecat-junit:0.2.0")
testImplementation("org.junit.jupiter:junit-jupiter:6.1.1")
testRuntimeOnly("org.junit.platform:junit-platform-launcher:6.1.1")
}
tasks.test { useJUnitPlatform() }Prepare and inspect the feature Spec before implementation, then seal the complete contract. After the agent's done claim, the normal command verifies the whole project:
./gradlew toppleCatCheck --spec specs/checkout/spec.md
./gradlew toppleCatReview --spec specs/checkout/spec.md
./gradlew toppleCatSeal
./gradlew test
./gradlew toppleCatVerifyFor a quick reviewer check of the work just delivered, Verify can scope the formal run with either one or more Spec files or repeated AC IDs. Do not combine the two forms:
./gradlew toppleCatVerify --spec specs/checkout/spec.md
./gradlew toppleCatVerify --ac AC-CHECKOUT-THRESHOLD --ac AC-CHECKOUT-VIPA scoped Verification Report repeats that its PASS covers only the listed ACs
and does not mean the complete executable contract passed. CI should use
toppleCatVerify without either option.
Reviewer HTML is English by default. To read ToppleCat-owned report prose in
Traditional Chinese, add --language zh-TW to toppleCatReview or
toppleCatVerify. The choice applies
only to that invocation's HTML; authored text and canonical values such as AC
IDs, Gate names, verdicts, and PIT outcomes stay exactly as recorded.
Start with the official technical documentation
or the getting-started guide. To see what
each safeguard catches, the optional JUnit cart-orders learning
project provides five synthetic routes: Public
Acceptance, Hidden Tests, Property-Based Testing, Mutation Testing, and
Contract Integrity. Each ./demo.sh <lesson> run keeps its local synthetic
HTML Verification Report at build/topplecat/demo-reports/<lesson>/index.html.
The sample README says which Gate to look for and what its expected failure
means.
Each Acceptance Condition has one public Java/JUnit Acceptance Method. The method describes one ordered Scenario; ordinary Java Stage methods perform the business calls and assertions.
@ToppleAcceptanceTest("AC-CART-COUPON")
@DisplayName("Apply a coupon to an order")
void appliesCoupon(ToppleCase c, ToppleScenario scenario, CouponStage coupon) {
scenario.given(coupon).a_cart(c.input("cart", Cart.class));
scenario.when(coupon).creates_an_order();
scenario.then(coupon).receipt_matches(c);
}Typed rows provide the inputs and expected results:
- caseId: coupon-at-threshold
acId: AC-CART-COUPON
inputs:
cart: {subtotal: 1000}
expected:
receipt: {total: 900}Public rows live under src/test; reviewer-owned rows use the same schema under
src/hiddenTest. See Authoring contracts for the
complete Java, Stage, case, expected-value, and Property rules.
| Artifact | Audience | Purpose |
|---|---|---|
build/topplecat/reports/review/index.html |
Reviewer | Spec Review before handoff: the complete selected Spec and bound executable material, not an execution result. |
build/topplecat/reports/verification/index.html |
Reviewer | Verification Report for the current formal run, with AC-first results and private diagnostics. |
build/topplecat/evidence.json |
Reviewer / automation | Machine-readable current-run verdict and Gate digests. |
build/topplecat/agent-feedback.json |
Implementation Agent | Safe Gate-level remediation without reviewer answers. |
Both Reviewer HTML reports accept the same invocation-only --language en or
--language zh-TW presentation choice on their report command. This changes headings, accessibility
text, controls, explanations, and HTML language metadata, but never report
JSON, evidence, safe feedback, the Mechanical Seal, or the contract itself.
Verification Report uses plain reader outcomes before canonical technical
evidence. Failed cases show input and expected/actual differences first;
Scenario Steps, raw failures, Gate verdicts, and PIT details remain available
for deeper inspection. Every AC begins with its key result visible; use its
reading control for one AC or the report-wide control for incremental all-AC
reading. A failed or incomplete safeguard is emphasized with its recorded
plain-language reason; the fixed safeguard order and technical evidence remain
unchanged. Technical evidence remains independently collapsed, and linked AC
or safeguard targets are revealed before the report moves to them.
The aggregate verdict is:
PASS: every required Gate passed for this run's ACs; a scoped run covers only its selected ACs;FAIL: completed verification found a blocking AC or Gate problem; orINCOMPLETE: ToppleCat could not obtain enough trustworthy current-run evidence.
The human keeps the final decision in every case.
ToppleCat starts at the executable acceptance boundary:
- Humans or external workflows select the Spec and remain responsible for complete rules and cases.
- Teams own task state, Spec lifecycle, execution placement, PR policy, organizational approval, and sign-off.
- Ordinary unit and QA testing, project-specific PIT, performance programs, and security programs remain project concerns.
- Reviewer Custody is plaintext mechanical storage, not encryption, a sandbox, CI isolation, or operating-system security.
Read the canonical product definition before proposing a new ToppleCat responsibility.
The repository also provides the
topplecat-acceptance skill
for turning selected ACs into executable Java acceptance methods, public and
reviewer case rows, and optional Properties. Humans or external workflow
automation remain responsible for running ToppleCat and deciding what to do
with its evidence.
