Repository navigation
v2.18.0
Minor release. Frames read behind a lame-duck INFO in the same chunk are no longer applied to the connection the failover opened (#182). Before, the old server's PING was answered on the new connection, a stale INFO or PONG was applied to it, and an -ERR failed the read or dropped the replayed reply inbox although the failover had succeeded.
Upgrading is recommended if your clients fail over on lame-duck mode (reconnect enabled, more than one server in the pool). There are no API changes. symfony-nats-messenger 5.4.1 works with this release unchanged.
Fixed
- Frames read behind a lame-duck INFO were handled on the new connection (#182). The lame-duck failover (#47) runs inline in whichever read brings the
INFO, and the dispatch went on with the rest of the chunk afterwards: aPINGgot itsPONGon the new connection; a staleINFOoverwrote the new server's info and a second lame-duckINFOstarted a second failover; aPONG, or a late reply of the old server on the reply inbox, confirmed the replayed reply inbox before the new server had answered anything; an-ERRdropped the still unconfirmed inbox or fired a replayed subscription's rejection handler and failed the read; when the failover failed, or theLameDucklistener closed the connection itself, thePONGwritten on the closed transport failed the read withThe connection is gone. The dispatch now records the connection it reads for and, once that connection is replaced or ended, treats the rest of the chunk as the old server's:PING,PONGandINFOare dropped, an-ERRis reported through the error listener asThe server the connection left sent an error frame: ...and not applied, and messages are still queued and delivered, a reply on the reply inbox without counting as the new server's confirmation of it. README Reconnect Behavior gained an item on the lame-duck failover and on this rule.
Quality gates
- PHPStan level 8.
- 2846 unit tests (with data sets) on PHP 8.2 to 8.5, 149 live integration tests and 48 Behat scenarios.
- All 45 runnable examples executed against a live server.
- Statement coverage 99.00% from the unit suite (95% floor enforced in CI) and 99.06% combined (97% floor).
- Infection covered MSI 93.3% over the 3985 mutants it ran (90% floor enforced in CI); it skipped 3844 of the 7829 it generated as too slow, so the score does not include them.
- Every fix has tests that fail on the old code: the
PONGon the second connection, the stale server id, the second dial, the inbox kept on a stale confirmation, the read failing with the stale-ERRor with the closed transport. Against two real nats-server instances put into lame-duck mode with theldmsignal, the failover and a request on the second server succeed; the server sends the lame-duckINFOalone in its chunk, so the coalesced orders are covered by the scripted transports. - The change was reviewed for semantics and for interleavings by two independent reviewers, and their findings were applied.
- Two unit tests that guarded against a spinning operation by bounding CPU time now count the operation's own looks instead (
CountingCancellationfor a request,CountingBufferfor a queue poll): the CPU bounds had failed the Mutation job under Xdebug coverage on the shared runners with the code correct, while a real spin stayed under them locally; the look counts are 0 to 6 on the current code and in the hundreds or thousands on spinning mutants. A process-wide count of event-loop registrations was tried as well and dropped, since reconnect attempts of clients that earlier tests leave behind register their 500 handshake polls in a burst. Dev-only. - symfony-nats-messenger 5.4.1's suites pass against this release: PHPStan level max, 528 unit tests, 52 functional scenarios against live NATS and 5 examples.
The full record is in the CHANGELOG.