v0.2.8 — a second joiner gets in, and a room stops going silent
v0.2.8 — a second joiner gets in, and a room stops going silent
Update with vox update. There's no compatibility step. v0.2.7 was tagged but never published: its release gate caught the sync defect below, and everything it contained ships here.
Fixed
- A second person joining right after the first is no longer locked out for 30 seconds. The node used to finish one connection handshake before it would start the next, and
vox connectexits as soon as it has joined, leaving the host mid-handshake. Handshakes now run concurrently, up to 64 at once. A flood of spoofed connection attempts can't use up those slots: an unverified source address must prove it can receive before anything is allocated to it. Measured: a host with two back-to-back joiners now lets the second in 5 times out of 5, against 0 out of 5 before. - Two members of a room could stop exchanging messages for good. When a member had both an anchor and another member to sync with, a failing sync with the anchor could take the room every round, so the direct sync was skipped every round too. A skipped sync is now owed and retried a second later. This had been failing CI intermittently since before v0.2.5. Measured in the configuration that breaks it: messages crossed in about 2s, where before they never did.
- The two ends of a connection could keep different copies. When a node had two connections to the same peer, each end chose which one to keep by arrival order, so they could disagree, and one end would go on using a connection the other had closed. Both ends now choose with the same order-independent rule. Duplicate pairs where the ends disagreed: 2 of 5 before, 0 of 28 after.
- A board no longer refuses a record it already holds. Re-announcing an unchanged record is now a no-op rather than a reported refusal. A changed record (your routable address, a new prekey bundle) is accepted immediately, where before it was refused for up to 60 seconds after startup.
- A member's node no longer freezes for 30–80 seconds. A peer's board, coordination and relay streams were served one at a time, so each waited behind the last, 20 seconds at a time. They're now served concurrently. Measured: freezes of 60s, 60s and 21s went to a single 20s (the last is being fixed in v0.2.9).
- A daemon started just as another vox finishes with the profile waits for it instead of failing with "another vox already has this profile open".
- Opening a stream can no longer wait forever. It's bounded at 20 seconds, and a dead connection now reports
Unreachable, notInternal. - A rate-limited member is told so (F14), with when it clears, instead of a generic failure.
- An urgent message from another node now interrupts its addressee (F15). Before, only messages written on the same node did.
service_rehearsal_proof, a real sshd reached through the overlay by a stranger who joins, is back in the blocking release gate after being excluded for the 30-second lockout.
Next, in v0.2.9
The remaining known defects, each tracked on the release-hardening board (epic #29): a node can still stall for up to 20s when a sync session holds a room's lock while it publishes; two members who both joined (neither created the room) cannot yet read each other (F12); a non-member can be served a room's sync; and the posting quota is being removed.