Repository navigation
Wolverine 6.47.0
Wolverine 6.47.0 is a release that largely came from outside the core team. The biggest change in it, a hardening pass over the NATS JetStream transport, was reported, designed and written by a community member, and so was the fix for a global partitioning handoff that stranded scheduled messages. Thank you to everyone who sent code, reproductions, or diagnoses this time around.
Community contributions
- @PascalEschbach reported eight distinct problems in the NATS transport on #4845, three of them silent message losses, and then fixed every one of them in #4856. The report alone was unusually good: each item came with the mechanism, not just the symptom, and the PR arrived with a test per item that fails without its fix. Details under NATS below.
- @BlackChepo found and fixed the case where a scheduled message parked under a global partition slot was promoted by the node that had just given the slot up, which threw into a stopped listener and committed the rows under a live node so recovery never touched them again (#4823, GH-4822).
- @smoqmilus reported, reproduced and diagnosed two losses of a forwarded scheduled envelope (GH-4824), and is credited as a co-author on the fix (#4830).
- @chrisbbe added the Marten-backed Native AOT smoke lane (#4826) that exposed three startup failures no native image in CI had ever reached, and is credited as co-author on the direct construction of the internal agent routers (#4842).
NATS JetStream
Sending durably over JetStream, with an outbox in front of a stream, lost messages in several places without an error or a log line. All of them are closed (#4856, with follow-ups in #4859):
- A publish the server refuses now fails the send. NATS answers a refused publish (a full stream under
DiscardPolicy.New, a failedNats-Expected-*check, an oversized message) with aPubAckcarrying an error, not an exception. Wolverine read only the sequence, reported success, and the durable outbox deleted a message the stream never stored. It throws now, the outbox keeps the message, and a duplicateNats-Msg-Idis still a success. - A message Wolverine moves to the error queue is dead-lettered on Wolverine's decision, not only once JetStream's
MaxDeliveris used up. A buffered listener acknowledges on receipt, so under the old guard the message was simply gone. The dead letter copy carries its ownNats-Msg-Id, because the original's id had it discarded as a duplicate when the dead letter subject sat in the same stream. - Sharded subject streams use work-queue retention.
AsWorkQueue()has always meant interest retention, which drops a message published while no consumer is bound, which is exactly the startup and rebalancing window a sharded topology has. The streamsUseShardedNatsSubjects()declares now keep a message until it is acknowledged.AsWorkQueue()itself is unchanged; its docs now say what it does. - Dropped core NATS messages are logged. A full subscription channel silently dropped messages; the loss is now a warning, throttled per subscription.
ConfigureNatsOpts()gives you the last word over the NATS.Net client options, for the pending channel size, ping interval and anything else Wolverine does not surface. It applies to the shared and every tenant connection.- Listeners reconcile their named consumers on start, so changing
AckWaitorMaxDeliverin code just works.Provisioning(NatsProvisioning.CreateOnly | CreateOrUpdate | Verify)controls what startup does with declared streams and named consumers that already exist, withVerifyfailing the start and listing every deviation. Resource setup and the listener now write the same consumer settings, filter included, so aVerifyhost starts cleanly after resource setup. - Resource setup honors the JetStream domain, so a leaf node no longer creates streams in its own JetStream instead of the configured domain. Only a 404 counts as "missing" anywhere Wolverine looks a stream or consumer up; any other failure propagates with its real message.
InvokeAsync()fails at once when nobody is listening. A request now carries a per-request reply subject, so the server's "no responders" answer reaches the waiting call and it fails in milliseconds withWolverineRequestReplyExceptioninstead of waiting out its timeout. The responding side answers on that exact subject, which a NATS user limited toallow_responsesrequires.
Global partitioning and durability
An exploratory pass over global partitioning after the long chain of slot handoff fixes, looking for the next "runs on a node that does not own the slot" before it was reported (#4853). It found two:
- A slot handoff releases the backlog that came through the shard queue. Every earlier regression fixture published from a node that owned the slot, so every inbox row carried the companion address. In a real cluster that is the minority path: a non-owner's message goes into the shard queue, and the owner's pop wrote the inbox row at the slot's address, which the handoff release did not cover. There is now a real multi-node Balanced fixture for this path.
- The shard queue bridge reports its depth, so the back pressure it applies is visible.
And around it:
- Scheduled messages parked under a slot are promoted on the node that owns it now, not the one that just gave it up (#4823).
- A forwarded scheduled envelope is stored as outgoing before its inbox row is retired, which closes two losses: a poll by the owning node between the send and the delete, and a failed first send (#4830).
- Ejecting a stale node, the step that releases every envelope the dead node owned, was the one link in crash recovery with no log line. It logs a warning now, and each sub-threshold stale observation on the way to it at debug (#4854).
- A requeue from an external durable listener no longer double-counts
Envelope.Attempts; a database or broker queue saw attempts go 1 then 3 (#4855).
Native AOT
The 6.46.0 AOT work stopped at the first failure each publish cycle, because the smoke target did. That is fixed, and the lanes it now runs found the rest:
- Every Native AOT smoke lane runs and all the failures are reported together (#4839). The literal explanation in the 6.46 bug reports was "each one appears once the previous one is worked around".
- Marten- and Polecat-backed native lanes (#4826, #4828). The first native image in CI to boot a store-backed HTTP application found three startup failures:
SideEffectPolicycould not findExecuteonIStartStream, the HTTPIEndpointMetadataProviderclose was trimmed, and Marten could not mapEnvelopein a native image. All three are rooted or avoided now. - The message routers are de-genericized (#4849). A handler returning a value type, the ordinary
InvokeAsync<Guid>request/reply shape, killed a native image at startup because the routing warm-up closed a generic over it reflectively. The routers take the message type as an argument now, so the one reflective close every message type had to be rooted for is gone, along with everything that existed to prop it up. - The framework's own agent messages are constructed directly (#4842). They are internal, so generated code could never name them in a rooting block, and an application with a durable message store trimmed exactly those instantiations.
- The side-effect roots match the test the policy actually applies, the HTTP rooting order is guarded, and a Balanced native lane runs in CI (#4847).
Also in this release
- WolverineFx.InMemory is a new package that runs Wolverine over the
JasperFx.Events.InMemoryprototyping store, so a store-agnostic application's handlers run before anyone has chosen Marten, Polecat or Fisher. Only Wolverine's store-agnostic usage is supported (#4844).