Repository navigation
Releases: girishkvs/mcp-pacemaker
Release list
2.0.1
Npm distribution safeguards, same-major update guidance, and controlled service replacement.
Added
- Exact-version update guidance that distinguishes the installed CLI from the running backend and does not silently cross major versions.
- Managed service stop/restart controls, persistent stop holds, and refusal of temporary npx roots for unattended service.
- Bundled third-party notices, fresh-consumer and exact-package replacement gates, and manual protected npm preparation/staging tooling.
- Bounded, opt-in pooling phase diagnostics for investigating deadline failures without logging configuration contents.
Fixed
- Canonicalize native log-watcher paths to avoid Windows short-path crashes.
- Correct cross-platform scanner path handling and owned temporary fixtures.
- Separate worker-startup expiry from native-helper timeout coverage without changing production deadlines.
Operational notes
- Preserves the 2.x queued-save contract and whole-batch Cancel/Undo. Resolve pending transactions before downgrading to 1.x.
- Windows automatic writes require .NET Framework 4.6.2 or later. The helper is unsigned.
- The earlier nine-second pooling stall remains unexplained; diagnostics do not establish a production fix.
- Npm publication is a separate step; this GitHub release does not upgrade a running installation.
Full changelog
1.3.1
Legacy 1.x distribution safeguards and same-major update guidance.
Security
- Require
smol-toml1.8.0 or newer within its major to address CVE-2026-85730.
Added
- Exact-version
legacychannel guidance, running-backend identity, and managed service stop/restart controls. - Bundled third-party notices, exact-package replacement checks, and manual npm preparation/staging tooling.
Fixed
- Canonicalize native log-watcher paths to avoid Windows short-path crashes.
- Correct cross-platform scanner path handling and owned temporary fixtures.
- Separate worker-startup expiry from native-helper timeout coverage without changing production deadlines.
Operational notes
- Retains the 1.x immediate-save contract; does not include 2.x batching or Windows save changes.
- Security and compatibility maintenance applies to the newest 1.3.x patch through March 31, 2027.
- Windows automatic writes require .NET Framework 4.6.2 or later and readable audit policy. The helper is unsigned.
- Npm publication is a separate step; this GitHub release does not upgrade a running installation.
Full changelog
2.0.0
Batched configuration saves, whole-batch Cancel/Undo, and ordinary-account Windows saves.
Breaking
- Version 2.0 changes the pooling save contract from immediate application to queued batches.
API clients must sendx-mcp-pooling-batch: 1and distinguish HTTP 202 acceptance from
activation. Older clients are rejected without staging a change. - The
prewarmCLI returns after staging, not after activation. Scripts must observe the batch
result before relying on the new settings. Undo restores the whole batch, not one server. - Update the CLI and reload open dashboard pages when upgrading the bridge from 1.3.0.
Before downgrading, use 2.0 to resolve any pending or interrupted configuration transaction.
Changed
- Pooling changes stage a complete configuration copy and reload five seconds after the last
accepted change. The UI and CLI distinguish Pending, Applying and completed activation;
Reload now flushes a batch, and Cancel/Undo explicitly applies to the whole batch. - Batch-capable clients identify the asynchronous save protocol. Older clients are rejected
before staging instead of reporting an accepted change as already enabled. - Windows configuration saves use directory-inherited auditing without requiring audit-read
privilege. A non-blocking notice explains that custom per-file audit rules may not carry
forward; ordinary access and supported integrity protections remain required. - Windows saves retain an existing owner that matches the caller user or token default owner,
preserving supported legacy writes after downgrade. Other owners still use checked transfer,
and source write access remains required.
Added
- Permanent real-version compatibility gates for 1.3.0 and 2.0.0 only, using pinned
published 1.3.0 code and an actual packed 2.0.0 candidate. CI covers CLI/API pairs on
Windows/Linux/macOS with Node 20/22 and built dashboard pairs in the Linux Chromium job:
immediate legacy saves/Undo, fresh-nonce protocol refusal, actual old-tab restart/refresh,
queued save/Cancel/Undo, settled upgrade/downgrade configuration bytes, and same-authority
legacy write/Undo capability after downgrade (or the unchanged known Windows audit refusal).
Fixed
- Preserve recovery checkpoints after uncertain post-placement failures without replaying
writes or accepting corrupt journals. - Avoid dispatching process-tree termination for children already known to have exited,
including failed session resumes.
Security
- Raise the
smol-tomlminimum and locked version from 1.7.0 to 1.8.0 to address
CVE-2026-85730.
Malformed TOML ending with a comment in an unclosed array or inline table now produces a
parse error instead of hanging the CLI while reading Codex configuration.
Platform notes
- Windows automatic configuration edits require .NET Framework 4.6.2 or later. Windows ARM64 execution has not been validated.
Full changelog
1.3.0
Pre-warming controls, opt-in shared sessions, cumulative process counters, and durable log commands.
Added
- Cumulative per-server process launch, failed launch, warm adoption, new session and resume
counters. The status snapshot identifies the bridge instance and interval start so process
churn can be compared without mixing time windows. Latency samples remain capped separately. - One-click pre-warming controls and conflict-safe Undo, with the requested warm count and
resident-process cost shown before activation. Pooling policy edits preserve active children. - A consolidated pre-warming dashboard tab and
mcp-pacemaker prewarmCLI view, both using
the same snapshot data. Explicit CLI actions enable, disable or undo one server's settings. mcp-pacemaker logsreads the durable log and retained rotation with--follow,--since,
--serverand literal--grep. Multiline context, partial UTF-8 writes and rotation are
handled without treating continuation lines as independent server records.- Opt-in shared stdio children for compatible stateless tools sessions. One real initialization
serves virtual sessions with owner-scoped IDs, progress, cancellation and cursors, bounded
queues, idle retention and draining. Unsupported features fail explicitly; uncertain tool
calls are never replayed. UI and CLI distinguish shared children from uninitialized pools.
Fixed
- Reserve session capacity before asynchronous startup, preventing queued cold starts from
overbooking the configured cap. Concurrent resumes of one session share one handshake. - Bind session lookup and deletion to the requested server. A stale exit or deleted resume
cannot remove or resurrect a replacement session. - Distinguish bridge-generated transport errors from identical numeric codes returned by an
upstream, so a legitimate server response is not mislabeled as a bridge timeout. - Guard shared JSON before serialization in both directions, including on Node 20, and bound
combined response sizes. Excessive depth or output produces controlled errors rather than
crashing the bridge or silently replaying tool calls. - Identify changed metadata fields when a pooling write is refused by the conflict guard.
- Avoid false Windows pooling conflicts from unrelated timestamp updates while retaining
audit, integrity-label, permission and file-identity checks. Refuse automatic writes when
audit policy cannot be verified or preserved. - Run configuration writes in a bounded, serialized worker so Windows permission inspection
does not block the bridge's request processing. - Replace per-write PowerShell startup and C# compilation with a bundled, reproducibly built
Windows security helper while retaining the same permission-preservation checks. - Bound configuration-write queueing and execution with shared cancellation. Report uncertain
or late-committed replacements honestly and reload saved state after a worker exits, even
with config watching disabled. Never replay a timed-out operation. - Update the UI build's
browserslistdependency to 4.28.9 to address its published memory-growth
and custom-stats advisories.
Operational notes
- Pooling and shared mode remain off unless explicitly enabled.
- Shared mode is for compatible stateless tools with one common credential context, not separate users or accounts.
- Windows automatic config edits require .NET Framework 4.6.2 or later and readable audit policy. Windows ARM64 execution has not been validated.
- A file replacement already in progress may finish after timeout. The bridge reports the outcome and reloads saved state; uncertain operations are never replayed automatically.
Full changelog
1.2.0
Cold start is now measured, gated, and reported — and the warm pool that was supposed to hide it
actually works. Everything here came out of running the live 15-server setup and asking why the
pool was full for a server that starts in 200ms and empty for the one that takes 20 seconds.
Cold start is measured
The bridge spawns every child, so it is the only thing positioned to know what starting a server
costs — and it was discarding that. /api/status and the dashboard now report p50/p95/max per
server, timed from spawn() to the server answering initialize. A process existing tells you
nothing about the package-manager work that follows it.
Without this there was no way to tell a server worth pre-warming from one that is not, and no way
to tell the user either. On the machine this was built on, seven of fifteen servers turned out to
cost more than two seconds per start.
Concurrent cold starts are gated
Servers launched through npx, uvx or dnx share one package cache, and starting several
together can corrupt it. Seventeen simultaneous npx invocations here produced
npm error code ECOMPROMISED / Lock compromised and failed every one of those sessions.
No server definition can fix that — it is a property of the concurrency, not of any one command —
so the bridge queues cold starts instead, two at a time by default
(MCP_MAX_CONCURRENT_SPAWNS). The slot is held until the child first speaks, which is when the
expensive part is over. Queuing a cold start behind another is strictly better than both of them
corrupting the cache.
Pooling is recommended, never enabled
Once a server's measured cold start passes MCP_POOL_ADVICE_MS, doctor and the dashboard print
the number and the exact config that would hide it, sized from observed peak concurrency, and say
what it costs.
The bridge does not apply it. A warm pool spends a resident process per slot on somebody's
machine, and that is their call to make. Cold start is otherwise invisible to the person paying
for it, who just experiences a slow tool.
Pools also size themselves once opted into: sharing: "pool" with no minWarm targets peak
concurrent sessions over the last hour, capped by MCP_WARM_MAX. An explicit minWarm still
wins. Defaulting to a single warm child is how a burst of seventeen sessions ended up paying
sixteen cold starts on a server that was configured to pool.
Warm pool fixes
- A pool never refilled after a child died on its own. The exit handler removed the corpse and
stopped; refill only ran on boot, on take, on recycle and on reload. A pooled server that went
quiet drained to empty and stayed there, degrading to a cold spawn per session — the exact
failure pooling exists to prevent, and invisible, because an empty pool looks like one nobody
has asked for anything yet. One server sat atwarm: 0for three hours while another with
identical config stayed full, purely because it was busy enough that every take triggered a
refill. Refill now happens on an unattended exit, with backoff so a child that dies instantly
cannot become a spawn loop. - A pooled server counted every failure twice. A warm child adopted into a session kept the
pool's exit handler and gained the session's, and both reported the exit — so an identical
server lost health twice as fast for having been pooled. - A pool refilled one child per take, so a burst that emptied it recovered long after the
burst was over. Refill now runs toward the target in one pass, bounded by the spawn gate.
Health fixes
- A 401 or 403 only counts against health when the bridge holds the credential. For a server
withauth: {type: none}the client authenticates, so an opening 401 is the handshake
working. Counting it flipped a healthy server here tofailingevery 60 seconds for hours. The
rule is now the one that matters: failing means the tool actually failed. - The bridge's own 503 counts as a failure. Refusing a session at the concurrency cap is the
most client-visible failure the bridge produces, and it was invisible to health — an operator
saw a healthy server while the bridge was the thing saying no. - A static
authorizationheader did not mark a server as bridge-authenticated, so its 401 read
asunknownforever and its challenge was relayed to a client that could not act on it. - The health probe no longer calls servers the bridge does not authenticate.
Restart reporting
/api/status carries how long ago the bridge restarted and how many clients held a session
before it and have not come back; status warns and the dashboard shows a banner. The bridge
re-establishes those sessions correctly, but some clients treat one connection failure as
permanent and need an MCP reload, and nothing said which ones.
Other
cappedSessionsnext tosessions, so a server's count matches the cap actually enforced. The
cap counts Streamable HTTP sessions while the display summed those and classic SSE, which is how
a server showed33/32.- SSE streams are closed on shutdown instead of dropped. A clean close reads as "reconnect" to a
client; a dropped connection reads as a transport fault, and some clients latch on that
permanently. - A terminated session still answers 404. An earlier revision of this work returned 410 for
sessions known to be gone, which reads better to a human and is wrong: the transport spec
requires 404 and makes re-initializing on 404 a client MUST, so 410 strands exactly the
clients that do the right thing. The distinction moved to the response body and the log, where
a client is not allowed to care about it. - A failed request names the credential source it actually used — the audience,
auth.command,
or the configuredauthorizationheader — instead of a generic message.
75 tests, up from 65. Every new test was checked by reverting the fix it guards and confirming it
fails, and the suite ran clean five consecutive times.
Full changelog
https://github.com/girishkvs/mcp-pacemaker/blob/main/CHANGELOG.md
1.1.0
Two features, and the defects that shook out of running a real 15-server setup daily.
Edit servers.json without restarting
Save the file and the bridge picks it up, or run mcp-pacemaker reload (also POST /admin/reload).
The reload is a diff: servers whose definition did not change keep their processes and their
live sessions. Adding one server no longer costs an outage on the fourteen that were already
working.
A file that does not parse, or a server with neither command nor url, is rejected whole —
the bridge logs why and keeps serving the last good config, so a half-written editor save cannot
take anything down. MCP_CONFIG_WATCH=0 requires an explicit reload instead of watching.
$ mcp-pacemaker reload
✓ reloaded ~/.mcp-pacemaker/servers.json on :8850
· added: postgres
· changed: github
· 2 session(s) restarted; untouched servers kept theirs
Per-server health
lastError was the only signal, and it is sticky: a server that failed once last week looked
exactly like one failing right now, and a recovered server still showed the old error. One server
here returned 401 on every call for hours while the dashboard reported nothing wrong.
Health is now a running verdict — ok, failing with a count of consecutive failures, or
unknown — in /api/status, doctor, top and the dashboard. unknown is deliberately not
ok: a server nobody has called is not claimed to be working.
It is derived from traffic already being proxied, so it costs nothing. A JSON-RPC error from a
server counts as healthy — it answered; a 401/403, a 5xx, a spawn failure or a non-zero exit
counts as failing.
MCP_HEALTH_INTERVAL_MS adds an optional probe for HTTP servers, so an idle one gets a verdict
before a client needs it rather than after a credential quietly expires overnight. Off by default,
since a probe is a real request to somebody else's service.
Auth fixes
- A credential is cached for no longer than it is actually valid.
refreshMinuteswas a flat
window, but an auth command usually hands back a JWT from the provider's own cache that may
already be most of the way through its life. The bridge then served a dead token for the rest of
the window while reporting it healthy, and every request 401'd. The JWTexpclaim now bounds
the cache entry. - An upstream 401 is no longer turned into a login prompt for a server the bridge
authenticates. The client cannot win that login — the bridge overwrites whatever token it
returns — so the effect was a browser sign-in loop. It is now reported as the bridge-side
failure it is, and the rejected credential is discarded rather than served out. - A bare
audienceis expanded into its token command in exactly one place. That expansion
doubles as the cache key, so two copies drifting apart would have silently disabled the cache.
Other fixes
- Session retention is enforced while the bridge runs, not only at startup. A bridge up for 28
hours was still honouring 45-hour-old records under a 24-hour window, andsessions.jsongrew
without bound. - A server named
api,admin,ui,statusor.well-knownis now reported. The bridge
answers those routes before consulting the server table, so such a server was unreachable and
nothing said so.doctorand the Health page now flag it. - A crash keeps its own error message. When a server exited and the in-flight request then
timed out, the generic timeout overwrote the exit code and stderr that explained why — the
symptom replacing the cause. - A scheduled recycle added by a config edit now takes effect. The interval was computed once
at startup, so a server givenrecycleMinuteslater would never be recycled.
Added
- A durable
bridge.lognext to the config, rolled atMCP_LOG_MAX_BYTES(5 MB default). The
in-memory buffer holds a few hundred lines, which a client reconnecting every 30 seconds fills in
about two minutes — so nothing that happened overnight could be explained the next morning.
Full changelog
https://github.com/girishkvs/mcp-pacemaker/blob/main/CHANGELOG.md
1.0.0
Keeps your MCP servers alive across editor restarts, reboots, and token expiry.MCP hosts run stdio servers as child processes, so when the host restarts, crashes, or exits — which every one-shot CLI run does — those servers die. mcp-pacemaker runs them as children of a small, long-lived bridge and exposes each over HTTP+SSE and Streamable HTTP, so hosts reconnect instead of respawning, and several hosts share the same warm, authenticated servers.## Getting started```bashnpm i -g github:girishkvs/mcp-pacemakermcp-pacemaker init````initdetects the MCP hosts you have installed, imports the servers already configured in them, rewires each host to the bridge, registers OS auto-start, and starts it. Every config file is backed up first,plan` shows the changes before anything is written, and `uninstall` restores them.## Highlights- Zero dependencies in the bridge. The long-running process pulls in nothing; the setup CLI is separate.- Seven hosts supported — VS Code, Cursor, Claude Desktop, Claude Code, Copilot CLI, Codex CLI, Gemini CLI — each written in its own config format.- Resumable sessions. A session id survives the bridge restarting, a recycle, or an idle reap: the bridge replays the client's `initialize` against a fresh server, so the client keeps the id it already holds. This is what makes idle reaping and scheduled recycle safe to run by default.- Pluggable auth. `command` runs anything that prints a token (az, gcloud, Vault, 1Password, a shell script) and caches it with proactive refresh; `env`, `static` and `none` are also supported. For servers where the client authenticates, the bridge relays OAuth discovery and challenges.- Auto-start and supervision on Windows, macOS and Linux, with a supervisor that restarts the bridge if it exits.- Dashboards and an HTTP API — a web UI at `/ui`, a terminal UI via `top`, and JSON/SSE endpoints for scripting.- Warm pooling, session caps with queueing, idle reaping, and scheduled recycle for servers holding a credential they can only refresh interactively.## A note on 1.0.0This starts at 1.0.0 rather than 0.x because the bridge, its config format and its HTTP surface were shaken out against a live 15-server setup used daily, not only a test suite. The `Hardening` section of the CHANGELOG lists the defects that came out of that — process-tree teardown on Windows, keep-alive socket lifetimes, empty token caching, and several others that are traps for anyone building something similar.Tested on Linux, macOS and Windows against Node 20 and 22.Full notes: CHANGELOG.md
Full changelog
https://github.com/girishkvs/mcp-pacemaker/blob/main/CHANGELOG.md