Skip to content

Testing en

shinyashen edited this page Sep 17, 2026 · 9 revisions

中文 | English

Testing

268 tests, all green (JUnit5 wired into gradlew build; the test sourceset compiles on a JDK 17 toolchain and never ships in the jar). This page explains how the suite is organized and why.

Methodology: differential testing

Correctness is not asserted by eyeballing — it is checked line-by-line against upstream semantics:

  1. The reference suite sets the capability face: TB-ThirdParty's Thunderbolt reference suite (12 scenario families × 3 stock modes MISSING/MINIMUM/UNBOUNDED) is the baseline. 1.21 concepts are encoded as real 1.12 pattern shapes: catalysts (returned) → equal same-key by-products; the durability (finiteUse) family was removed along with the amortization (see Differences & Roadmap); reusable host stock (returnedFrom) → slot substitute variants plus the host pool merged into network stock.
  2. Every portable upstream test family is ported — assertions matched to source semantics, not just shape.
  3. Exposure becomes a pin: every engine semantic gap the differential testing uncovered was fixed and turned into a regression test.
  4. The adjudication batch gate: every plan the planner judges executable must COMPLETE on the faithful CPU simulator; unavoidable engine divergences are exempted by name in an allowlist cited to AE2UEL source, and a stall anywhere else fails the build. Each scenario gets a FRESH planner instance and the gate is skipped on runner timeouts (the 1s deadline makes the case inconclusive) — a abandoned worker's late writes must not pollute the gate's read. The plan's task order is CONSTRUCTED by topological sort of the pattern dependency graph (P0): folded rings keep the firing order their priming floor was probed under (the cycle is cut at the priming edges). For exact plans any such order completes — a deadlock needs a cycle of mutual waits, and the greedy drain of shared outputs merely delays a consumer by a pass — so the gate is demoted to a fallback assertion; the only order-stall exemption (single-dag/fibonacci/unbounded) left the table with the construction, leaving five. The two FALSE_POSITIVEs the suite once carried (fuzzy/variant-route, host-variant routing) were resolved on 2026-09-17 by translating the fixture into its real 1.12 shape — the reusable host tool becomes a wildcard-durability CRAFTING pattern with substitution enabled, which the CPU's craftable branch executes natively. Since then the suite runs 36/36 SUPPORTED with zero timeouts and zero exempted divergences. Plans over two million crafts trust the construction (the replay would cost real server-thread time).

Capability conclusion

36/36 SUPPORTED — including the zero-missing closure of upstream's only FALSE_POSITIVE (carried since v1.9.6) — plus 38/38 boundary feasibility assertions, zero flakiness across repeated fresh JVM runs.

Test families

Family Cases Coverage
VmSemanticsTest / VmSemantics2Test 13 marker seeds, amplifier net-gain correction (seedless reports exactly 1), catalyst working-capital, sub-item stock awareness, conversion ring conservation, shortfall reporting, 11-level Fibonacci, JIT cross-request reuse, self-growth pruning, amount-1 boundary, FUZZY_SLOT, processing default fuzzy (the durability closed-form tests were removed with the amortization)
Capability suites (Reference + Boundary) 37 + 38 all reference-suite scenarios SUPPORTED (the durability family was removed); amount boundaries, fuzzy primary output × white/grey/no-variant stock, mid-chain stock at depth 10/20, craftable fluid partial stock
CrossRequestCacheTest 11 cross-request determinism matrix (same request / shortfall branch / multi-step stock drain / amount change / deep chain / diamond / dual-VM isolation / cold-hot equivalence / empty-stock recapture / 24-level 10^9 chain craftables never missing)
Recursion & rings 15 RecursionReference 6 + CatalystFeedbackLoop 6 (lossy reports exactly 2) + conversion ring, self-growth
Fuzzy family 13 v1.10.5 exact-slot vs substitute-slot demand separation, substitution group registration semantics, craftable sub-item with partial stock still schedules, NBT variants in usedItems
Task-order construction (TaskOrderingTest) 8 DAG producers first; ring units keep their probed firing order; bridge patterns fold into the SCC; self-loops inert; counts preserved; deterministic
JIT / reuse / plan gating 10 JitReuseTest 4 + PlanReuseGuardTest 3 (stock re-verification pure functions + version invalidation) + ScaledUnwrapTest 2, plus lifecycle cases
Dead-cycle pre-prune 10 mutual-ring detection, self-loop always pruned, seeded ring kept, external input feeds ring, sibling candidates don't feed, by-product inflow feeds, deep chain non-ring, pruning keeps healthy candidates, all-pruned fallback, end-to-end healthy path
Boundary family 42+2+2 etc. stock-aware sub-pattern off-by-one matrix (42 parameterized), craftable-fluid last-unit delivery (35 parameterized), fluid bucket boundary, x1/x2, multi-pattern solver regression 5, partial-stock chains 3 (GAP-2 audit)
Fluid compat 8 against the real AE2 Fluid Craft Rework classes (ObjectHolder fields hand-filled on plain JVM): drop NBT identity and mB carrying, long amount decoding, packet-key encoding and over-amount rejection, compile-time normalization, full-VM chain on real item keys
VMTest / PerfReportTest 6 + 5 bytecode builder / CALL_BY_KEY / request wrapping / DIV_ROUNDUP units; performance informational (see the performance page)

Harness infrastructure

  • Plain JVM, no bootstrap: IAEItemStack fakes, slot-level substitute recipe fakes, and a sandbox fake (non-destructive network view + type-exact insert cache, same semantics as production NetworkCraftingSandbox) — no Minecraft launch needed;
  • Vendored planner: TB-ThirdParty's pure-Java planner (23 files) is vendored into the test sourceset as the differential baseline (the Thunderbolt-Core runtime cannot run on 1.12 — see Differences);
  • TraceSimulationState: a per-scenario decision-trace tool — an X-ray machine for solver behavior.

Real bugs the differential testing exposed (all pinned by regressions)

  1. A self-returning catalyst seed could be "self-satisfied" by the pattern's own by-product → seed extraction moved before this bundle's by-product insertion;
  2. Processing default-fuzzy on exact slots swallowed cross-item substitution groups → split nbtFamilyOf vs fuzzyFamilyOf;
  3. Stock snapshots missed substitution groups / NBT variants when no IGrid handle existed → snapshotExecuteStartStock extended enumeration;
  4. New pattern registration blocked by a stale bundle (1.12 equivalent of upstream's PatternRefreshReuse) → bundle composite key + missingSubKeys;
  5. "Plan missing → restock → recalculate" replayed stale shortfalls → shortfallRetryable;
  6. Craftable substitutes for fuzzy slots were not scheduled → resolver T2.5 craftable-variant resolution.

Clone this wiki locally