Repository navigation
v2.10.0
Minor release. Two fixes for what happens when a socket dies or a close interrupts a reconnect, both found while reviewing symfony-nats-messenger's reconnect support: a subscribe(), flush() or rtt() whose write finds the socket dead now reconnects instead of throwing the socket's raw error, and a close now stops a reconnect's dial and waits for what it stopped, so a connect() after it dials afresh.
Upgrading is recommended for anyone running with reconnects on (the default), anyone consuming JetStream with fetches (a fetch subscribes its inbox first, so it was the most common victim of the first bug), and anyone who closes and reopens a connection. Read the upgrade notes first: some operations throw different exceptions after a dead socket.
Upgrade notes
- A
subscribe(),flush()orrtt()whose write finds the socket dead no longer throws the socket's own error (for exampleAmp\ByteStream\StreamException): it starts the reconnect and waits for it within its own timeout, as it waits for a reconnect already in flight.subscribe()then returns;flush()andrtt()throw aConnectionException(Connection lost before the server answered the PING). A wait that runs out throws aTimeoutException, and with reconnect off they throw aConnectionException(Reconnect is disabled). Code that caught the transport's exception there should catch those instead. unsubscribe()no longer throws when its write finds the socket dead.disconnect()anddrain()now return only once the reconnect orconnect()they stopped has ended. With the built-in transports that is at once; a custom transport that implements onlyTransportInterfacecan hold them up toconnectTimeoutMs. ImplementCancellableDialTransportInterfaceto let a close stop its dial.
Added
[feature]CancellableDialTransportInterface, a transport whose dial can be stopped: itsconnect()takes a cancellation that fires when the application closes the connection while a connect or a reconnect is dialling. Both built-in transports implement it.
Fixed
- A
subscribe(),flush()orrtt()whose write found the socket dead failed with the socket's raw error and left the connection Open on it, so every operation that wrote a control frame failed the same way until the heartbeat noticed. After a server restart the application had not seen yet, a JetStream fetch threw "Broken pipe", and a queue consumer built on it exited. Such a write is now a connection failure, as a failed publish write already was. disconnect()ordrain()during a reconnect that was dialling left that reconnect running after they returned: they cut its backoff short but not its dial, which with Amp's retry pauses takes some 6 s against a refused port. Aconnect()issued after the close failed withRecovery was aborted before the connection opened, and in a synchronous application, which never gave the reconnect the event-loop time to end, every laterconnect()did. A close now stops the dial, Amp's retry pauses included, and waits for what it stopped before it returns. Aconnect()racing the close still fails at once.- A reconnect that a close stopped delivered the messages queued on the connection on its way out: after
disconnect(), which discards them, or besidedrain()'s own delivery. It now leaves them to the close.
Quality gates
PHPStan level 8, 2332 unit tests (with data sets), 149 live integration tests, 48 Behat scenarios and 45 runnable examples executed against a live server, 99.11% unit and 99.23% combined statement coverage (95% and 97% floors enforced in CI) and 94.2% Infection covered MSI over 7353 mutants (90% floor enforced in CI). Every fix has a test that fails when the fix alone is reverted, and mutation testing on the changed lines left only equivalent mutants alive (listed in the message of the commit that pins the rest). Both fixes were also verified through symfony-nats-messenger against a live server. The full record is in the CHANGELOG.