Skip to content

v2.14.0

Choose a tag to compare

@bpacholek bpacholek released this 07 Oct 02:14
· 26 commits to main since this release

Minor release. An operation whose own read is the first to meet a lost connection now waits for the reconnect only within its own deadline, and returns as soon as its result arrives during the outage (#178). Before, that read ran the whole reconnect itself: a poll with a 1 s timeout returned only once the server was back, 6.65 s into a 3.5 s outage, and a request whose reply a held-up delivery brought 50 ms after the drop returned it after the outage.

Upgrading is recommended if your application can lose its connection while a poll, request, fetch or pull consumer is the only reader. There are no API changes; the documented wait of such an operation changes, below. symfony-nats-messenger 5.4.1 works with this release unchanged.

Upgrade notes

  • Operations wait for a reconnect within their own timeout, also when their own read started it. request()/requestMany(), SubscriptionQueue polls, fetchBatch()/fetchNext(), directGetBatch(), the pull consumer, Key/Value keys()/history(), flush() and rtt() whose own read is the first to notice a lost connection now start the reconnect on its own fiber and wait for it as they wait for one another fiber started: within their timeout, and not at all with waitForReconnect: false, where they fail at once with Connection is not open. An operation that times out before the reconnect gives up no longer sees Reconnect attempts exhausted; the Closed event carries it. Your own processIncoming() and readIncoming(), a serving loop's read and the heartbeat still run the reconnect themselves and wait for all of it.
  • flush(), rtt() and drainSubscription() fail earlier. A flush whose own read meets the drop fails with Connection lost before the server answered the PING as soon as the reconnect's first attempt clears the pong slots, rather than after the reconnect; a drainSubscription() whose flush does resolves then, its subscription removed, while the reconnect runs on.

Fixed

  • An operation whose own read ran the reconnect overran its deadline by the outage (#178). The reconnect now runs on its own fiber and the operation's read waits for it only within its deadline and wake-up: a result that a delivery still under way brings during the outage ends the wait at once, the deadline ends it with the operation's timeout, and the reconnect carries on either way.

Quality gates

  • PHPStan level 8.
  • 2775 unit tests (with data sets) on PHP 8.2 to 8.5, 149 live integration tests and 48 Behat scenarios.
  • All 45 runnable examples executed against a live server.
  • Statement coverage 99.04% from the unit suite (95% floor enforced in CI) and 99.10% combined (97% floor).
  • Infection covered MSI 93.8% over the 3812 mutants it ran (90% floor enforced in CI); it skipped 3931 of the 7743 it generated as too slow, so the score does not include them.
  • Every CI job passed on the first attempt. Every fix has tests that fail on the old code (70 new cases in two test classes): the operations returned after the whole outage, where they now return at their deadline or with their result. On a real server with AmpSocketTransport a poll with a 1 s timeout returned at 1.0 s during a 3.5 s outage and got the first message published after the reconnect, where 2.13.0 returned at 6.65 s.
  • The change was reviewed for semantics and for interleavings by two independent reviewers, and their findings were applied.
  • symfony-nats-messenger 5.4.1's suites pass against this release: PHPStan level max, 528 unit tests, 52 functional scenarios against live NATS and 5 examples.

The full record is in the CHANGELOG.