Skip to content

Wolverine 6.42.0

Choose a tag to compare

@jeremydmiller jeremydmiller released this 29 Sep 22:13
· 243 commits to main since this release

A release driven almost entirely by what users reported. Five of the durability fixes below came from one reporter working a multi-tenant deployment hard enough to find them, and the pattern running through them is the same: a failure that belonged to one tenant was being paid for by every tenant, or was not being reported at all.

⚠️ Read this first: a durable publish that cannot be persisted now fails the caller

GH-4662. PublishAsync/ScheduleAsync made outside a handler and outside any transaction, to a local queue made durable by UseDurableLocalQueues(), used to return without an error when the message store was unreachable. Four attempts inside ~400 ms, one line at Information, and the message was gone — not even recovered after the database came back, because no inbox row was ever written.

The cause was the shape of the retry rather than the retry itself. RetryBlock.PostAsync runs one attempt inline and posts the rest to its own worker, so the task the caller awaits completes the moment the first attempt fails. That is the right shape once an envelope is durable somewhere and recovery can find it; it is the wrong shape for the write that makes it durable, because there is no other copy.

DurableWriteRetry keeps the same 50/100/250 ms budget, awaits it inline, and lets the last exception reach the caller.

Who is affected is narrower than it sounds. If you enlist a real outbox (EF Core, Marten, DatabaseEnvelopeTransaction) nothing changes — that path never reached this code. Inside a handler nothing new escapes either; the publish is buffered and a failure is still caught, logged and recorded as a discard, just at Error with a real tracking event instead of a buried Information line. Only the un-transacted publish from outside a handler can now throw.

One tenant's outage stops costing you every tenant

  • A tenant-only inbox failure no longer pauses the listener for everyone (GH-4658), including the single-store batch path that the first fix missed (GH-4659).
  • A deferred tenant-scoped failure no longer spins. It was being redelivered with no delay, so one tenant's outage turned into a hot loop. There is now a per-tenant write brake.
  • An envelope whose completion could not be written no longer stays owned by the live node until that node stops (GH-4664) — the completion retries until the store returns.
  • PostgreSQL: the envelope tables' timestamp default is now a true UTC instant (GH-4663). It depended on the session's time zone, so InboxStaleTime handed back envelopes that were still running. ⚠️ Changing a column default does not migrate already-deployed tables — see the upgrade notes.
  • Documented what actually happens to a durable listener when one tenant's database is down (GH-4660).

Partitioning

A scheduled message no longer runs on a node that does not own the slot (GH-4673). Under global partitioning a scheduled envelope took the local shortcut regardless of slot ownership, so two nodes could process the same group id concurrently — the exact thing partitioning exists to prevent. A promoted scheduled envelope also no longer still reports its status as Scheduled.

HTTP

Stream large multipart uploads through a MultipartReader endpoint parameter, contributed by @erdtsieck (#4674) — no buffering the whole upload to handle it. F# code generation emits the new frame too, and the F# codegen coverage is now gated against a committed baseline so a new unsupported frame cannot slip in unnoticed.

Diagnostics, metrics and operations

  • Dead letters from local queues are counted again. Every local:// destination was being classified as System traffic, so a user's own local queue never reached the per-type/per-tenant DeadLetters metric (GH-4665).
  • Scheduled-message listings carry the envelope's tenant (GH-4666).
  • BrokerResource.Check no longer reports a thrown check as a missing resource (GH-4693), which was hiding the real cause behind "Missing known broker resources".
  • A projection agent that can never start is no longer retried forever (GH-4676), using ShardStartException.Reason.
  • SQLite: SqliteNodePersistence honors the table prefix (GH-4668); node shutdown failed outright on a prefixed store.
  • SQL Server tenant stores start against a named-instance connection string (GH-4685).

Dependencies

  • JasperFx 2.76.1 — carries the named-instance fix above.
  • Weasel stays at 9.35.1. The upgrade to 9.37.0 was prepared and deliberately held back: it regresses Oracle, where a node_number column that is already GENERATED BY DEFAULT AS IDENTITY is read back as drift, so every migration emits an ALTER ... MODIFY that Oracle refuses. It will ship once that is fixed upstream.

Thank you

@framos-varajo Reported five of the durability issues above — GH-4658, 4659, 4662, 4663 and 4664 — with repros. Most of this release is theirs.
@erdtsieck (Anne Erdtsieck-Wijnen) Contributed the MultipartReader streaming upload support in #4674
@alexandrefresnais Reported the global-partitioning scheduled-message ownership bug, GH-4673
@frankvdb7 Reported the SQL Server named-instance startup failure, GH-4685