Skip to content

Reconnection

Gerasimos (Makis) Maropoulos edited this page Oct 6, 2026 · 2 revisions

When to use this page

  • The network between your client and server is unreliable (mobile, hotel WiFi, VPN drops).
  • You want clients to automatically rejoin their namespaces and rooms after a drop.
  • You need to know on the server side whether a connecting client is a reconnect attempt.

How reconnection works

Reconnection is client-driven. The server doesn't track sessions for offline clients, it simply accepts the next handshake and assigns a fresh Conn.ID. The client carries enough state (previously connected namespaces and joined rooms) to rejoin transparently.

The Go client has no built-in reconnect. neffos.Dial dials once; Conn.ReconnectTries and Conn.WasReconnected() exist so a caller that redials by hand (its own Dial call in a loop) can tell the server about it, but neffos itself does not retry. A new Dial (or, on the server side, a new Upgrade) is required to reconnect.

The JavaScript client (neffos.js 0.3.0) does reconnect on its own when you ask for it:

  1. The connection drops (server crash, network failure, sleep/wake). onclose fires with a CloseInfo.
  2. If Options.reconnect is set and shouldReconnect(info) allows it, the client waits with exponential backoff, then redials.
  3. The retry dial carries X-Websocket-Reconnect: N (N is the attempt number in the current cycle, starting at 1), as a real header with the ws package or, in a browser or any other runtime, as the URL parameter X-Websocket-Header-X-Websocket-Reconnect=N (the server reads it back into a header through neffos.URLParamAsHeaderPrefix).
  4. The server's Upgrade sets Conn.ReconnectTries = N on the new connection.
  5. The client reuses the same Conn object (wasReconnected() becomes true on it) and restores every previously connected namespace and joined room, in order, before handing control back to your handlers.

Enabling reconnection (JavaScript client)

A number sets just the first retry delay in milliseconds:

const conn = await neffos.dial("ws://localhost:8080/echo", handlers, {
    reconnect: 5000,
});

Or pass a ReconnectOptions object for full control:

const conn = await neffos.dial("ws://localhost:8080/echo", handlers, {
    reconnect: {
        initialDelay: 1000, // default
        maxDelay: 30000,    // default
        factor: 2,          // default: each retry waits twice as long as the one before
        jitter: 0.2,        // default: randomly shortens a delay by up to 20%, never lengthens it
        maxRetries: 10,     // default: no limit
        shouldReconnect: (info) => info.code !== 4001, // default: always reconnect
    },
});

probe (default false) brings back the 0.2.0 behavior of sending an HTTP HEAD request to the endpoint before each retry and redialing only once it answers. Call conn.close(), or abort the Options.signal you passed to dial, to stop a reconnect cycle in progress.

Detecting a reconnect on the client

if (conn.wasReconnected()) {
    console.log("reconnected after", conn.reconnectTries, "attempts");
}

conn.isClosed() stays false across a reconnect, since the same Conn is reused. It only becomes true once reconnection gives up (maxRetries reached, shouldReconnect returns false, or close() was called).

Detecting a reconnect on the server

Two places to look:

// 1. From inside Upgrade / OnConnect (or any event):
if c.WasReconnected() {
    log.Printf("client %s reconnected after %d tries", c.ID(), c.ReconnectTries)
}

// 2. If you call Server.Upgrade directly and want to distinguish "this was a
//    reconnect probe, not a real failure":
if neffos.IsTryingToReconnect(err) {
    // The HEAD probe, only sent when the client sets ReconnectOptions.probe. Respond and
    // let the client dial again.
    return
}

IsTryingToReconnect returns true for the HTTP HEAD probe a client sends while waiting for the endpoint to come back online, if it opted into probe: true (the neffos.js 0.3.0 default is no probe). The server returns 302 Found automatically; you do not have to handle this case unless you call Upgrade from a custom handler.

Re-joining rooms

On a drop, the client records which namespaces were connected and which rooms were joined in each, fires the local forced-leave and disconnect events, then redials. Once the new socket is up, it restores each namespace in order (reusing the same NSConn object your application already holds a reference to), then that namespace's rooms, one at a time. If the connection drops again before the restore finishes, the still-unrestored namespaces and rooms carry over into the next retry instead of being lost.

If you need different rejoin behavior (for example, user confirmation before rejoining), skip Options.reconnect and handle it yourself: listen for the OnNamespaceConnected event and check wasReconnected().

When NOT to reconnect

A few cases where the client does not attempt a reconnect, or stops trying:

  • Options.reconnect was never set. No reconnect logic runs at all; onclose just closes the connection.
  • shouldReconnect(info) returns false for this drop. You get the CloseInfo (code, reason, wasClean), so you can, for example, give up on an application-level close code the server sent on purpose.
  • maxRetries was reached. The client reports ERR_RECONNECT through onError (or console.warn without one) and closes the Conn for good.
  • The client called conn.close() itself, or the Options.signal you passed to dial was aborted. Either stops a reconnect cycle in progress.

The server also makes this easier to get right on purpose: Server.Shutdown and Server.Close now send close code 1001 (CloseGoingAway) to every connection, so a shouldReconnect that checks info.code can tell a deliberate server shutdown apart from a dropped network. In every case above, Conn.isClosed() is true once the client has given up.

Architecture diagram

sequenceDiagram
    autonumber
    participant C as Client
    participant S as Server
    C->>S: WebSocket handshake (initial)
    S-->>C: 101 Switching Protocols
    Note over C,S: connection healthy
    Note over C: network drops, onclose fires with CloseInfo
    Note over C: wait with exponential backoff (1s, 2s, 4s, ... capped at maxDelay)
    C->>S: GET /echo + X-Websocket-Reconnect: 1
    S-->>C: 101 Switching Protocols
    Note over S: Upgrade sets c.ReconnectTries = 1
    Note over C: same Conn reused, wasReconnected() becomes true
    C->>S: connect("default")
    S-->>C: ack
    C->>S: joinRoom("lobby")
    S-->>C: ack
    Note over C,S: state restored
Loading

With ReconnectOptions.probe: true, each retry first sends HEAD /echo and waits for a 302 Found before the GET above, matching the 0.2.0 behavior.

Verification

A simple manual test: launch a server, dial from a browser with reconnect: 2000, stop the server, restart it. Watch the browser DevTools: you should see the WebSocket close, a retry after about 2 seconds (longer on each subsequent attempt), then a new WebSocket handshake, then OnNamespaceConnected re-firing for each previously connected namespace.

See also

Clone this wiki locally