Repository navigation
1.2.0
Cold start is now measured, gated, and reported — and the warm pool that was supposed to hide it
actually works. Everything here came out of running the live 15-server setup and asking why the
pool was full for a server that starts in 200ms and empty for the one that takes 20 seconds.
Cold start is measured
The bridge spawns every child, so it is the only thing positioned to know what starting a server
costs — and it was discarding that. /api/status and the dashboard now report p50/p95/max per
server, timed from spawn() to the server answering initialize. A process existing tells you
nothing about the package-manager work that follows it.
Without this there was no way to tell a server worth pre-warming from one that is not, and no way
to tell the user either. On the machine this was built on, seven of fifteen servers turned out to
cost more than two seconds per start.
Concurrent cold starts are gated
Servers launched through npx, uvx or dnx share one package cache, and starting several
together can corrupt it. Seventeen simultaneous npx invocations here produced
npm error code ECOMPROMISED / Lock compromised and failed every one of those sessions.
No server definition can fix that — it is a property of the concurrency, not of any one command —
so the bridge queues cold starts instead, two at a time by default
(MCP_MAX_CONCURRENT_SPAWNS). The slot is held until the child first speaks, which is when the
expensive part is over. Queuing a cold start behind another is strictly better than both of them
corrupting the cache.
Pooling is recommended, never enabled
Once a server's measured cold start passes MCP_POOL_ADVICE_MS, doctor and the dashboard print
the number and the exact config that would hide it, sized from observed peak concurrency, and say
what it costs.
The bridge does not apply it. A warm pool spends a resident process per slot on somebody's
machine, and that is their call to make. Cold start is otherwise invisible to the person paying
for it, who just experiences a slow tool.
Pools also size themselves once opted into: sharing: "pool" with no minWarm targets peak
concurrent sessions over the last hour, capped by MCP_WARM_MAX. An explicit minWarm still
wins. Defaulting to a single warm child is how a burst of seventeen sessions ended up paying
sixteen cold starts on a server that was configured to pool.
Warm pool fixes
- A pool never refilled after a child died on its own. The exit handler removed the corpse and
stopped; refill only ran on boot, on take, on recycle and on reload. A pooled server that went
quiet drained to empty and stayed there, degrading to a cold spawn per session — the exact
failure pooling exists to prevent, and invisible, because an empty pool looks like one nobody
has asked for anything yet. One server sat atwarm: 0for three hours while another with
identical config stayed full, purely because it was busy enough that every take triggered a
refill. Refill now happens on an unattended exit, with backoff so a child that dies instantly
cannot become a spawn loop. - A pooled server counted every failure twice. A warm child adopted into a session kept the
pool's exit handler and gained the session's, and both reported the exit — so an identical
server lost health twice as fast for having been pooled. - A pool refilled one child per take, so a burst that emptied it recovered long after the
burst was over. Refill now runs toward the target in one pass, bounded by the spawn gate.
Health fixes
- A 401 or 403 only counts against health when the bridge holds the credential. For a server
withauth: {type: none}the client authenticates, so an opening 401 is the handshake
working. Counting it flipped a healthy server here tofailingevery 60 seconds for hours. The
rule is now the one that matters: failing means the tool actually failed. - The bridge's own 503 counts as a failure. Refusing a session at the concurrency cap is the
most client-visible failure the bridge produces, and it was invisible to health — an operator
saw a healthy server while the bridge was the thing saying no. - A static
authorizationheader did not mark a server as bridge-authenticated, so its 401 read
asunknownforever and its challenge was relayed to a client that could not act on it. - The health probe no longer calls servers the bridge does not authenticate.
Restart reporting
/api/status carries how long ago the bridge restarted and how many clients held a session
before it and have not come back; status warns and the dashboard shows a banner. The bridge
re-establishes those sessions correctly, but some clients treat one connection failure as
permanent and need an MCP reload, and nothing said which ones.
Other
cappedSessionsnext tosessions, so a server's count matches the cap actually enforced. The
cap counts Streamable HTTP sessions while the display summed those and classic SSE, which is how
a server showed33/32.- SSE streams are closed on shutdown instead of dropped. A clean close reads as "reconnect" to a
client; a dropped connection reads as a transport fault, and some clients latch on that
permanently. - A terminated session still answers 404. An earlier revision of this work returned 410 for
sessions known to be gone, which reads better to a human and is wrong: the transport spec
requires 404 and makes re-initializing on 404 a client MUST, so 410 strands exactly the
clients that do the right thing. The distinction moved to the response body and the log, where
a client is not allowed to care about it. - A failed request names the credential source it actually used — the audience,
auth.command,
or the configuredauthorizationheader — instead of a generic message.
75 tests, up from 65. Every new test was checked by reverting the fix it guards and confirming it
fails, and the suite ran clean five consecutive times.
Full changelog
https://github.com/girishkvs/mcp-pacemaker/blob/main/CHANGELOG.md