Releases: SAY-5/ledgermesh
Release list
v6.0.0 — durable authorization ownership and compensation safety
LedgerMesh 6.0.0
This release repairs order deduplication, cancellation compensation and payment
authorization ownership in LedgerMesh's synthetic-provider demo. The previous
assigned-ID merge path could let two requests with one idempotency key create two
orders. Cancelled orders could retain card authorizations, and competing payment
creators or late approvals could leave surplus holds at the processor.
New records insert rather than overwrite an existing request or payment. Payment
cancellation is acknowledged and resent until answered; the operator resend endpoint
also covers older cancelled orders. Inventory reservations belong to an order and
SKU, so repeated or out-of-order releases do not credit stock that order never held.
Durable, order-bound authorization decisions serialize claim and retirement before
a fresh payment read. Claim, payment authorization and the completion outbox commit
together. Retirement is irreversible and commits before provider release, preventing
both a late claim of an already released hold and release of a concurrently claimed
hold. Failed releases retry from outstanding-provider listings after restart; retired
codes never become claimable again. This includes H2 and PostgreSQL race regressions,
first-insert conflicts, cancellation, transaction rollback and an additive-schema
upgrade check.
The chaos verifier inspects final orders, payments, stock holds and synthetic
processor authorizations, not only aggregate request counts. Cleanup is conditional
on provider/database availability, scheduler progress and listing throughput; grace
plus one interval is not an unconditional time guarantee. Concurrent attempts can
still produce temporary surplus authorizations; only one can be kept and the others
must be released. These tests are not evidence of a production payment deployment.
Upgrade requirements
- Stop and drain all old payment listeners, sweepers and callback-capable processes
before activating the new protocol. Do not use a mixed-version rolling upgrade. - Preserve authorization decision tombstones. An old binary is not a safe rollback
target while retired-code callbacks can still arrive. - Let old inventory orders and their compensations finish before upgrading inventory.
The new reservation ledger cannot reconstruct or automatically credit deductions
made by the 5.0.0 inventory service. Terminal order state alone is insufficient. - Replay older uncompensated cancellations with the documented bounded administrator
resend endpoint until none remain. Consult the operator upgrade guide.
Verification
The reconciled release tree passed 172 local tests: 144 unit/H2, ten PostgreSQL
ownership/upgrade and 18 end-to-end tests, with zero failures, errors or skips.
Independent review reproduced the original races before repair and reran the
ownership tests on both databases. Formatting, configuration and measured-artifact
drift checks, TypeScript/Vite build and all browser self-checks passed.
Final release PR CI
passed all seven jobs. ExactlyOnce
passed 50/50 fresh-stack runs (300 test executions); CompensationResend
passed 50/50 (50 test executions). Independent inspection of seed 74713's retained
chaos artifact found 1,200 submitted orders, 1,170 confirmations, 30 stock cancellations,
three service kills and no failed/stuck orders or money/stock violations. Every
confirmed payment retained exactly one matching authorization; 14 surplus holds
were released. The PR synthetic-merge checkout and final release tree are identical.
The merged commit also passed all seven jobs in main CI 36551110669, including a fresh chaos run.
v5.0.0
Every service now answers GET /ops/overview with its own health, the lag of each of its listener groups, the dead letter depth per topic, the state of every circuit breaker and, on the order service, the sagas that are still open and the ones that have already missed a deadline. The numbers come from the gauges the Prometheus endpoint already publishes, so the page an operator reads and the dashboards cannot disagree, and ledgermesh.saga.stuck is published alongside them. The chaos harness gained a tight profile that doubles the kills and restarts the victim after two seconds instead of five, reachable as make chaos-tight, and its summary now ends with the overview of all three services. The recorded run in the README is a real make chaos on this release: 1200 orders at 20 a second, three kills, 1162 confirmed, 38 cancelled for stock, none failed or stuck, with no lag, no dead letters and no stuck orders left behind. Host port overrides let the stack run beside something else already on 8081 to 8083.
v4.0.0
POST /orders now accepts an Idempotency-Key header. The first call stores the answer it returned under that key in the same transaction as the order, its outbox row and its timeline entry, so the key and the effect can never disagree, and a repeat is answered from the store byte for byte with Idempotent-Replay: true instead of placing a second order. Two calls racing on one key both try to store the answer; the primary key lets one of them through and the loser rolls back its own order and returns what the winner wrote. On the way out the relay stamps an attempt on an outbox row before it sends and marks it published only after the broker acknowledged it, so a row that comes back attempted but unpublished is exactly the crash window between the two and is sent again and counted in ledgermesh.outbox.resends. The redelivery changes nothing downstream because the consumer recognises the event id it already applied, which the new tests check against the stock ledger and the payment for the order.
v3.0.0
Every business topic now has a dead letter shadow. A record whose handler keeps failing is retried in place with exponential backoff and is then published to <topic>.dlq with its original coordinates and the exception in headers, so the partition moves on instead of stalling behind it. POST /admin/dlq/{topic}/replay reads dead letters back onto their source topic with a dedicated consumer group and commits only after the republish, and the consumers that had already applied the event ignore the replay by event id. A replay count travels with the record, so once it has used up ledgermesh.dlq.max-replays the replayer parks it instead of starting another lap. Consumer lag and dead letter depth are recomputed from broker offsets into gauges, and every failed delivery is counted per topic.
v2.0.0
Every saga step now has a deadline. A stuck order reaper cancels reservations that never answer (releasing stock through the usual compensation), re-drives a silent payment once over the new order.payment_requested topic, and cancels with PAYMENT_TIMEOUT if the second deadline passes too. The payment service answers a re-drive with the outcome it already has, or attempts the payment on the spot.
GET /orders/{id}/timeline returns every step an order went through with timestamps, including late or ignored events. Unit tests: 73 (was 59).
v1.0.0
First cut of LedgerMesh: order, inventory and payment services on Redpanda with a transactional outbox, idempotent consumers, a compensating saga, Redis write-through caching and Resilience4j breakers, retries and time limiters. The chaos harness submits 1200 orders at 20/s while killing a service three times and finishes with zero failed or stuck orders.
Unit tests: 59 on H2. Integration tests: 6 under Testcontainers.