Skip to content

v2.10.3

Choose a tag to compare

@bpacholek bpacholek released this 05 Oct 10:39
· 87 commits to main since this release

Patch release. It fixes a regression from 2.10.1: an -ERR the server sends while keeping the connection open ended a healthy connection. Examples are maximum subscriptions exceeded and a permissions violation for a publish with a reply subject. It also fixes the bugs found while reviewing that fix, in reconnect announcements, the request reply inbox, drains, services and INFO handling.

Upgrading is recommended for everyone on 2.10.1 or 2.10.2, which have the regression, and for anyone on 2.10.0 or earlier, since most of the other fixes are for long-standing bugs. symfony-nats-messenger 5.4.1 requires it. Read the upgrade notes first. Several fixes change when connection listeners are called and what the error listener receives.

Upgrade notes

  • -ERRs that keep the connection open. maximum subscriptions exceeded, Invalid Publish Subject and permissions violations other than ... for Publish to ... and ... for Subscription to ... no longer end the connection. The read that brought one fails with it, as in 2.10.0, and no Disconnected, Reconnected or Closed follows. The 2.10.1 note now covers only an -ERR the server closes the connection after, such as Stale Connection.

    • The heartbeat, a reconnect's replay, a service's run() and the flushes of drain(), drainSubscription() and a service's drain() report such an -ERR to the error listener and carry on. Watch that listener, not Closed, for these errors.
  • Drains past such an -ERR. drain() and drainSubscription() now read on past it to their PONG.

    • If the server then goes silent, drain() waits out its budget, where it used to end at the -ERR.
    • If the server closes the connection instead, drainSubscription() runs the reconnect itself. With reconnect on, it returns only once that reconnect ends, past requestTimeoutMs if need be. 2.10.1 and 2.10.2 waited for that reconnect only within the budget.
  • Reply inbox at the subscription limit. Once the server rejects the shared reply inbox of request()/requestMany() for the subscription limit, requests fail fast until a slot is free. Each attempt costs a round trip and nothing throttles them, so a caller that retries should pause between attempts.

  • Transport doubles. The write that subscribes the reply inbox now ends with a PING (SUB _INBOX.<inbox>.* <sid>\r\nPING\r\n). A test double must answer every PING, not only a write that is exactly PING\r\n.

  • When Reconnected is announced. A reconnect now announces the new connection once it is over, not from inside it.

    • Operations that waited for the reconnect go on without waiting for the listener. disconnect() and drain() no longer wait for a listener that is still running.
    • An operation the listener calls that finds the connection gone reconnects again, and the listener hears those events while that call is still running. Do not hold a lock across an await on the same connection inside a listener.
    • A Reconnected call always finds the connection Open. A connection lost before the listener was told is not announced.
  • Publish retry against a drain. A drain() that waited for a reconnect can now close the connection before a publish() is retried, if that publish's failed write ran the reconnect. This happens while the Reconnected listener, or a logger that suspends as it records the event, is still busy. That publish then fails and is not sent. An error from this library's own write is thrown as is. A built-in transport's raw stream error is wrapped in a TransportClosedException ("Transport is not connected").

  • Closed with reconnect off. The Closed event of a lost connection now carries the error that ended it and is logged at warning level. A failed read is therefore logged twice.

  • Services. A service's run() now reports what its reads used to swallow, to the error listener and at error level. That includes another subscription's handler that throws, once per failure. A handler's CancelledException no longer stops the service.

  • Header names. NatsHeaders::fromWireBlock() and fromWireBlockMulti() declare their keys as int|string, since a header named 1 comes back as an int key. Static analysis now reports two kinds of code:

    • code that hands such a key to a function taking a string, such as strtolower($name) in a loop over the map; under strict types that call already threw a TypeError whenever a publisher sent such a header;
    • code that passes the whole map where string keys are declared, such as Symfony Messenger's SerializerInterface::decode().

    Cast such keys to string there. NatsHeaders::get() takes the map as it is.

  • First INFO. A first INFO that is valid JSON but not an object fails the connect, as broken JSON does, with INFO payload is not a JSON object. With reconnect off, connect() fails with that message. With reconnect on (the default), it first retries like any failed first dial, then fails with Reconnect attempts exhausted, carrying that error as its previous exception.

Fixed

  • -ERR handling (2.10.1 regression). An -ERR the server keeps the connection open for fails the read again and leaves the connection open. A fatal -ERR read together with one still ends it.
  • Reconnect replay. A reconnect whose replayed SUB the server rejects for the subscription limit now reports the rejection and completes. It used to fail every attempt until the reconnect gave up and closed the connection. That bug is long-standing and was in 2.10.0 as well. 2.10.1 made it easy to hit, since the rejection itself forced a reconnect.
  • Drain flushes. The flushes of drain() and drainSubscription() no longer lose the messages behind an -ERR that keeps the connection open. That loss is long-standing and was in 2.10.0 as well.
  • Flushes waited out their budget. A drain() waited out its whole budget (requestTimeoutMs, 10 s by default) against a healthy server when another fiber's read took its PONG. That happened with an application's processIncoming() loop, a service's run(), or a request() issued just before it in the same tick. flush(), rtt(), drainSubscription() and a service's drain() had a narrower race of the same kind, and rtt() then reported the wait as the round trip.
  • Request reply inbox.
    • The inbox is recorded before its SUB is written, so a permissions rejection read during a reconnect's replay trips the latch from 2.7.1 (#167).
    • An inbox the server rejects for the subscription limit is dropped and subscribed again, instead of staying dead. A PING sent behind the new SUB confirms it.
    • After a rejection, a request is not sent until the server has taken a new inbox. A request is also not sent once a drain() has unsubscribed its inbox.
  • Reconnect announcements. An operation called from a Reconnected listener, or from a Connected listener after a failed first connect(), could hang or wait out its timeout when the new connection was already gone, and the connection stayed Open on the dead socket. It now reconnects again.
    • A drain() no longer waits for a slow listener.
    • A supervisor that reconnects on the Closed of a drain() called from that listener now opens a new connection. It used to fail with Recovery was aborted before the connection opened.
    • A close the listener makes takes over what the reconnect read.
    • A publish() whose failed write ran the reconnect is retried according to the state it then finds.
  • Services. run() reported nothing when another subscription's handler threw, left the rest of that read queued, and stopped silently on a handler's CancelledException. It now reports the failure and delivers the rest.
    • A service's drain() reads on to its PONG.
    • A throwing request validator is answered like a throwing handler.
    • A request header named 1 no longer breaks an endpoint with a TypeError.
  • INFO. An async INFO that is valid JSON but not an object, such as INFO 1 or INFO [1], is reported as malformed and the last server info is kept. It used to fail the read, or to replace the server info with defaults.
  • Heartbeat with reconnect off. When the heartbeat gives up on a connection, the Closed event now says why (#172).

Quality gates

  • PHPStan level 8.
  • 2533 unit tests (with data sets), 149 live integration tests and 48 Behat scenarios.
  • All 45 runnable examples executed against a live server.
  • Combined statement coverage 99.24% (97% floor enforced in CI).
  • Infection covered MSI 93.88% over 7617 mutants (90% floor enforced in CI).
  • Every fix has tests that fail on the old code.
  • The branch was reviewed by independent reviewers, one track at a time and then as a whole. Every finding was fixed or documented, except two pre-existing issues left for a later release.
  • symfony-nats-messenger's suites pass against this release: 528 unit tests, 52 functional scenarios against live NATS, and 5 examples.

The full record is in the CHANGELOG.