main-360fa77
·
49 commits
to main
since this release
fix(client): bound the control heartbeat write (P-9)
P-4 gave public tunnels a control heartbeat. Every server built before P-7
never reads the public control substream, and clients upgrade independently
of the servers they dial, so a new client against an old server fills the
yamux stream's 256 KiB flow-control credit and blocks forever inside the
`control.send()` that lives in the listen loop's `select!` arm. The tunnel
stays registered and stops serving — the worst failure shape there is.
Measured on the staging server still running the pre-P-7 build, with
BORE_CTRL_HEARTBEAT_MS=2 so that days of uptime compress into seconds:
served at t+5s and t+15s, then NO RESPONSE at t+30s, t+60s and t+90s, with
no further "new connection" in the client log.
`client::beat_once` now bounds the write with
`secret::ctrl_heartbeat_send_timeout()` (10s, BORE_CTRL_HEARTBEAT_SEND_TIMEOUT_MS)
and returns CtrlBeat::{Sent,Closed,PeerNotReading}. On PeerNotReading the
loop warns once, naming the cause and the remedy, and stands the heartbeat
down for the session: the client degrades to the legacy heartbeat-free path
instead of wedging. That is correct against both peers — an old server has no
reaper to trip, and a current server reaping a control path this broken is the
right outcome. Cancelling SinkExt::send at the deadline is safe because Framed
advances its write buffer only by the bytes the socket accepted.
Re-measured against the same old server with the guarded client: 8 probes out
of 8 served across 120s, the stand-down warning logged exactly once, and the
next accept 126us after it. The residual cost is one deadline for one proxied
connection, never again for the life of the session.
Gates at the io-trait mock level, because the real wedge needs ~256 KiB of
frames: beat_once_stands_down_when_the_peer_stops_reading (red-checked —
removing the bound makes the test HANG rather than fail, which is exactly the
production symptom) and beat_once_sends_immediately_to_a_reading_peer.
Also in this change: the public-tunnel measurement harness fixes found while
taking the evidence — the admin view field names are public_port/active, the
samplers now share one remote prefix and source lib.sh, srv/verify.sh reports
the running server build, and srv/redeploy.sh refuses to report success unless
the version string actually changed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BqfehFhjLR3Yj1sgaf5KCi