Skip to content

Releases: girishkvs/mcp-pacemaker

2.0.1

Choose a tag to compare

@girishkvs girishkvs released this 14 Sep 10:20
c8f2d72

Npm distribution safeguards, same-major update guidance, and controlled service replacement.

Added

  • Exact-version update guidance that distinguishes the installed CLI from the running backend and does not silently cross major versions.
  • Managed service stop/restart controls, persistent stop holds, and refusal of temporary npx roots for unattended service.
  • Bundled third-party notices, fresh-consumer and exact-package replacement gates, and manual protected npm preparation/staging tooling.
  • Bounded, opt-in pooling phase diagnostics for investigating deadline failures without logging configuration contents.

Fixed

  • Canonicalize native log-watcher paths to avoid Windows short-path crashes.
  • Correct cross-platform scanner path handling and owned temporary fixtures.
  • Separate worker-startup expiry from native-helper timeout coverage without changing production deadlines.

Operational notes

  • Preserves the 2.x queued-save contract and whole-batch Cancel/Undo. Resolve pending transactions before downgrading to 1.x.
  • Windows automatic writes require .NET Framework 4.6.2 or later. The helper is unsigned.
  • The earlier nine-second pooling stall remains unexplained; diagnostics do not establish a production fix.
  • Npm publication is a separate step; this GitHub release does not upgrade a running installation.

Full changelog

v2.0.0...v2.0.1

1.3.1

Choose a tag to compare

@girishkvs girishkvs released this 14 Sep 10:20
979b1f0

Legacy 1.x distribution safeguards and same-major update guidance.

Security

  • Require smol-toml 1.8.0 or newer within its major to address CVE-2026-85730.

Added

  • Exact-version legacy channel guidance, running-backend identity, and managed service stop/restart controls.
  • Bundled third-party notices, exact-package replacement checks, and manual npm preparation/staging tooling.

Fixed

  • Canonicalize native log-watcher paths to avoid Windows short-path crashes.
  • Correct cross-platform scanner path handling and owned temporary fixtures.
  • Separate worker-startup expiry from native-helper timeout coverage without changing production deadlines.

Operational notes

  • Retains the 1.x immediate-save contract; does not include 2.x batching or Windows save changes.
  • Security and compatibility maintenance applies to the newest 1.3.x patch through March 31, 2027.
  • Windows automatic writes require .NET Framework 4.6.2 or later and readable audit policy. The helper is unsigned.
  • Npm publication is a separate step; this GitHub release does not upgrade a running installation.

Full changelog

v1.3.0...v1.3.1

2.0.0

Choose a tag to compare

@girishkvs girishkvs released this 13 Sep 11:42
db4812a

Batched configuration saves, whole-batch Cancel/Undo, and ordinary-account Windows saves.

Breaking

  • Version 2.0 changes the pooling save contract from immediate application to queued batches.
    API clients must send x-mcp-pooling-batch: 1 and distinguish HTTP 202 acceptance from
    activation. Older clients are rejected without staging a change.
  • The prewarm CLI returns after staging, not after activation. Scripts must observe the batch
    result before relying on the new settings. Undo restores the whole batch, not one server.
  • Update the CLI and reload open dashboard pages when upgrading the bridge from 1.3.0.
    Before downgrading, use 2.0 to resolve any pending or interrupted configuration transaction.

Changed

  • Pooling changes stage a complete configuration copy and reload five seconds after the last
    accepted change. The UI and CLI distinguish Pending, Applying and completed activation;
    Reload now flushes a batch, and Cancel/Undo explicitly applies to the whole batch.
  • Batch-capable clients identify the asynchronous save protocol. Older clients are rejected
    before staging instead of reporting an accepted change as already enabled.
  • Windows configuration saves use directory-inherited auditing without requiring audit-read
    privilege. A non-blocking notice explains that custom per-file audit rules may not carry
    forward; ordinary access and supported integrity protections remain required.
  • Windows saves retain an existing owner that matches the caller user or token default owner,
    preserving supported legacy writes after downgrade. Other owners still use checked transfer,
    and source write access remains required.

Added

  • Permanent real-version compatibility gates for 1.3.0 and 2.0.0 only, using pinned
    published 1.3.0 code and an actual packed 2.0.0 candidate. CI covers CLI/API pairs on
    Windows/Linux/macOS with Node 20/22 and built dashboard pairs in the Linux Chromium job:
    immediate legacy saves/Undo, fresh-nonce protocol refusal, actual old-tab restart/refresh,
    queued save/Cancel/Undo, settled upgrade/downgrade configuration bytes, and same-authority
    legacy write/Undo capability after downgrade (or the unchanged known Windows audit refusal).

Fixed

  • Preserve recovery checkpoints after uncertain post-placement failures without replaying
    writes or accepting corrupt journals.
  • Avoid dispatching process-tree termination for children already known to have exited,
    including failed session resumes.

Security

  • Raise the smol-toml minimum and locked version from 1.7.0 to 1.8.0 to address
    CVE-2026-85730.
    Malformed TOML ending with a comment in an unclosed array or inline table now produces a
    parse error instead of hanging the CLI while reading Codex configuration.

Platform notes

  • Windows automatic configuration edits require .NET Framework 4.6.2 or later. Windows ARM64 execution has not been validated.

Full changelog

v1.3.0...v2.0.0

1.3.0

Choose a tag to compare

@girishkvs girishkvs released this 09 Sep 18:53

Pre-warming controls, opt-in shared sessions, cumulative process counters, and durable log commands.

Added

  • Cumulative per-server process launch, failed launch, warm adoption, new session and resume
    counters. The status snapshot identifies the bridge instance and interval start so process
    churn can be compared without mixing time windows. Latency samples remain capped separately.
  • One-click pre-warming controls and conflict-safe Undo, with the requested warm count and
    resident-process cost shown before activation. Pooling policy edits preserve active children.
  • A consolidated pre-warming dashboard tab and mcp-pacemaker prewarm CLI view, both using
    the same snapshot data. Explicit CLI actions enable, disable or undo one server's settings.
  • mcp-pacemaker logs reads the durable log and retained rotation with --follow, --since,
    --server and literal --grep. Multiline context, partial UTF-8 writes and rotation are
    handled without treating continuation lines as independent server records.
  • Opt-in shared stdio children for compatible stateless tools sessions. One real initialization
    serves virtual sessions with owner-scoped IDs, progress, cancellation and cursors, bounded
    queues, idle retention and draining. Unsupported features fail explicitly; uncertain tool
    calls are never replayed. UI and CLI distinguish shared children from uninitialized pools.

Fixed

  • Reserve session capacity before asynchronous startup, preventing queued cold starts from
    overbooking the configured cap. Concurrent resumes of one session share one handshake.
  • Bind session lookup and deletion to the requested server. A stale exit or deleted resume
    cannot remove or resurrect a replacement session.
  • Distinguish bridge-generated transport errors from identical numeric codes returned by an
    upstream, so a legitimate server response is not mislabeled as a bridge timeout.
  • Guard shared JSON before serialization in both directions, including on Node 20, and bound
    combined response sizes. Excessive depth or output produces controlled errors rather than
    crashing the bridge or silently replaying tool calls.
  • Identify changed metadata fields when a pooling write is refused by the conflict guard.
  • Avoid false Windows pooling conflicts from unrelated timestamp updates while retaining
    audit, integrity-label, permission and file-identity checks. Refuse automatic writes when
    audit policy cannot be verified or preserved.
  • Run configuration writes in a bounded, serialized worker so Windows permission inspection
    does not block the bridge's request processing.
  • Replace per-write PowerShell startup and C# compilation with a bundled, reproducibly built
    Windows security helper while retaining the same permission-preservation checks.
  • Bound configuration-write queueing and execution with shared cancellation. Report uncertain
    or late-committed replacements honestly and reload saved state after a worker exits, even
    with config watching disabled. Never replay a timed-out operation.
  • Update the UI build's browserslist dependency to 4.28.9 to address its published memory-growth
    and custom-stats advisories.

Operational notes

  • Pooling and shared mode remain off unless explicitly enabled.
  • Shared mode is for compatible stateless tools with one common credential context, not separate users or accounts.
  • Windows automatic config edits require .NET Framework 4.6.2 or later and readable audit policy. Windows ARM64 execution has not been validated.
  • A file replacement already in progress may finish after timeout. The bridge reports the outcome and reloads saved state; uncertain operations are never replayed automatically.

Full changelog

v1.2.0...v1.3.0

1.2.0

Choose a tag to compare

@girishkvs girishkvs released this 07 Sep 23:31

Cold start is now measured, gated, and reported — and the warm pool that was supposed to hide it
actually works. Everything here came out of running the live 15-server setup and asking why the
pool was full for a server that starts in 200ms and empty for the one that takes 20 seconds.

Cold start is measured

The bridge spawns every child, so it is the only thing positioned to know what starting a server
costs — and it was discarding that. /api/status and the dashboard now report p50/p95/max per
server, timed from spawn() to the server answering initialize. A process existing tells you
nothing about the package-manager work that follows it.

Without this there was no way to tell a server worth pre-warming from one that is not, and no way
to tell the user either. On the machine this was built on, seven of fifteen servers turned out to
cost more than two seconds per start.

Concurrent cold starts are gated

Servers launched through npx, uvx or dnx share one package cache, and starting several
together can corrupt it. Seventeen simultaneous npx invocations here produced
npm error code ECOMPROMISED / Lock compromised and failed every one of those sessions.

No server definition can fix that — it is a property of the concurrency, not of any one command —
so the bridge queues cold starts instead, two at a time by default
(MCP_MAX_CONCURRENT_SPAWNS). The slot is held until the child first speaks, which is when the
expensive part is over. Queuing a cold start behind another is strictly better than both of them
corrupting the cache.

Pooling is recommended, never enabled

Once a server's measured cold start passes MCP_POOL_ADVICE_MS, doctor and the dashboard print
the number and the exact config that would hide it, sized from observed peak concurrency, and say
what it costs.

The bridge does not apply it. A warm pool spends a resident process per slot on somebody's
machine, and that is their call to make. Cold start is otherwise invisible to the person paying
for it, who just experiences a slow tool.

Pools also size themselves once opted into: sharing: "pool" with no minWarm targets peak
concurrent sessions over the last hour, capped by MCP_WARM_MAX. An explicit minWarm still
wins. Defaulting to a single warm child is how a burst of seventeen sessions ended up paying
sixteen cold starts on a server that was configured to pool.

Warm pool fixes

  • A pool never refilled after a child died on its own. The exit handler removed the corpse and
    stopped; refill only ran on boot, on take, on recycle and on reload. A pooled server that went
    quiet drained to empty and stayed there, degrading to a cold spawn per session — the exact
    failure pooling exists to prevent, and invisible, because an empty pool looks like one nobody
    has asked for anything yet. One server sat at warm: 0 for three hours while another with
    identical config stayed full, purely because it was busy enough that every take triggered a
    refill. Refill now happens on an unattended exit, with backoff so a child that dies instantly
    cannot become a spawn loop.
  • A pooled server counted every failure twice. A warm child adopted into a session kept the
    pool's exit handler and gained the session's, and both reported the exit — so an identical
    server lost health twice as fast for having been pooled.
  • A pool refilled one child per take, so a burst that emptied it recovered long after the
    burst was over. Refill now runs toward the target in one pass, bounded by the spawn gate.

Health fixes

  • A 401 or 403 only counts against health when the bridge holds the credential. For a server
    with auth: {type: none} the client authenticates, so an opening 401 is the handshake
    working. Counting it flipped a healthy server here to failing every 60 seconds for hours. The
    rule is now the one that matters: failing means the tool actually failed.
  • The bridge's own 503 counts as a failure. Refusing a session at the concurrency cap is the
    most client-visible failure the bridge produces, and it was invisible to health — an operator
    saw a healthy server while the bridge was the thing saying no.
  • A static authorization header did not mark a server as bridge-authenticated, so its 401 read
    as unknown forever and its challenge was relayed to a client that could not act on it.
  • The health probe no longer calls servers the bridge does not authenticate.

Restart reporting

/api/status carries how long ago the bridge restarted and how many clients held a session
before it and have not come back; status warns and the dashboard shows a banner. The bridge
re-establishes those sessions correctly, but some clients treat one connection failure as
permanent and need an MCP reload, and nothing said which ones.

Other

  • cappedSessions next to sessions, so a server's count matches the cap actually enforced. The
    cap counts Streamable HTTP sessions while the display summed those and classic SSE, which is how
    a server showed 33/32.
  • SSE streams are closed on shutdown instead of dropped. A clean close reads as "reconnect" to a
    client; a dropped connection reads as a transport fault, and some clients latch on that
    permanently.
  • A terminated session still answers 404. An earlier revision of this work returned 410 for
    sessions known to be gone, which reads better to a human and is wrong: the transport spec
    requires 404 and makes re-initializing on 404 a client MUST, so 410 strands exactly the
    clients that do the right thing. The distinction moved to the response body and the log, where
    a client is not allowed to care about it.
  • A failed request names the credential source it actually used — the audience, auth.command,
    or the configured authorization header — instead of a generic message.

75 tests, up from 65. Every new test was checked by reverting the fix it guards and confirming it
fails, and the suite ran clean five consecutive times.

Full changelog

https://github.com/girishkvs/mcp-pacemaker/blob/main/CHANGELOG.md

1.1.0

Choose a tag to compare

@girishkvs girishkvs released this 03 Sep 04:46

Two features, and the defects that shook out of running a real 15-server setup daily.

Edit servers.json without restarting

Save the file and the bridge picks it up, or run mcp-pacemaker reload (also POST /admin/reload).

The reload is a diff: servers whose definition did not change keep their processes and their
live sessions. Adding one server no longer costs an outage on the fourteen that were already
working.

A file that does not parse, or a server with neither command nor url, is rejected whole —
the bridge logs why and keeps serving the last good config, so a half-written editor save cannot
take anything down. MCP_CONFIG_WATCH=0 requires an explicit reload instead of watching.

$ mcp-pacemaker reload
✓ reloaded ~/.mcp-pacemaker/servers.json on :8850
·   added:   postgres
·   changed: github
·   2 session(s) restarted; untouched servers kept theirs

Per-server health

lastError was the only signal, and it is sticky: a server that failed once last week looked
exactly like one failing right now, and a recovered server still showed the old error. One server
here returned 401 on every call for hours while the dashboard reported nothing wrong.

Health is now a running verdict — ok, failing with a count of consecutive failures, or
unknown — in /api/status, doctor, top and the dashboard. unknown is deliberately not
ok: a server nobody has called is not claimed to be working.

It is derived from traffic already being proxied, so it costs nothing. A JSON-RPC error from a
server counts as healthy — it answered; a 401/403, a 5xx, a spawn failure or a non-zero exit
counts as failing.

MCP_HEALTH_INTERVAL_MS adds an optional probe for HTTP servers, so an idle one gets a verdict
before a client needs it rather than after a credential quietly expires overnight. Off by default,
since a probe is a real request to somebody else's service.

Auth fixes

  • A credential is cached for no longer than it is actually valid. refreshMinutes was a flat
    window, but an auth command usually hands back a JWT from the provider's own cache that may
    already be most of the way through its life. The bridge then served a dead token for the rest of
    the window while reporting it healthy, and every request 401'd. The JWT exp claim now bounds
    the cache entry.
  • An upstream 401 is no longer turned into a login prompt for a server the bridge
    authenticates.
    The client cannot win that login — the bridge overwrites whatever token it
    returns — so the effect was a browser sign-in loop. It is now reported as the bridge-side
    failure it is, and the rejected credential is discarded rather than served out.
  • A bare audience is expanded into its token command in exactly one place. That expansion
    doubles as the cache key, so two copies drifting apart would have silently disabled the cache.

Other fixes

  • Session retention is enforced while the bridge runs, not only at startup. A bridge up for 28
    hours was still honouring 45-hour-old records under a 24-hour window, and sessions.json grew
    without bound.
  • A server named api, admin, ui, status or .well-known is now reported. The bridge
    answers those routes before consulting the server table, so such a server was unreachable and
    nothing said so. doctor and the Health page now flag it.
  • A crash keeps its own error message. When a server exited and the in-flight request then
    timed out, the generic timeout overwrote the exit code and stderr that explained why — the
    symptom replacing the cause.
  • A scheduled recycle added by a config edit now takes effect. The interval was computed once
    at startup, so a server given recycleMinutes later would never be recycled.

Added

  • A durable bridge.log next to the config, rolled at MCP_LOG_MAX_BYTES (5 MB default). The
    in-memory buffer holds a few hundred lines, which a client reconnecting every 30 seconds fills in
    about two minutes — so nothing that happened overnight could be explained the next morning.

Full changelog

https://github.com/girishkvs/mcp-pacemaker/blob/main/CHANGELOG.md

1.0.0

Choose a tag to compare

@girishkvs girishkvs released this 31 Aug 07:22

Keeps your MCP servers alive across editor restarts, reboots, and token expiry.MCP hosts run stdio servers as child processes, so when the host restarts, crashes, or exits — which every one-shot CLI run does — those servers die. mcp-pacemaker runs them as children of a small, long-lived bridge and exposes each over HTTP+SSE and Streamable HTTP, so hosts reconnect instead of respawning, and several hosts share the same warm, authenticated servers.## Getting started```bashnpm i -g github:girishkvs/mcp-pacemakermcp-pacemaker init````initdetects the MCP hosts you have installed, imports the servers already configured in them, rewires each host to the bridge, registers OS auto-start, and starts it. Every config file is backed up first,plan` shows the changes before anything is written, and `uninstall` restores them.## Highlights- Zero dependencies in the bridge. The long-running process pulls in nothing; the setup CLI is separate.- Seven hosts supported — VS Code, Cursor, Claude Desktop, Claude Code, Copilot CLI, Codex CLI, Gemini CLI — each written in its own config format.- Resumable sessions. A session id survives the bridge restarting, a recycle, or an idle reap: the bridge replays the client's `initialize` against a fresh server, so the client keeps the id it already holds. This is what makes idle reaping and scheduled recycle safe to run by default.- Pluggable auth. `command` runs anything that prints a token (az, gcloud, Vault, 1Password, a shell script) and caches it with proactive refresh; `env`, `static` and `none` are also supported. For servers where the client authenticates, the bridge relays OAuth discovery and challenges.- Auto-start and supervision on Windows, macOS and Linux, with a supervisor that restarts the bridge if it exits.- Dashboards and an HTTP API — a web UI at `/ui`, a terminal UI via `top`, and JSON/SSE endpoints for scripting.- Warm pooling, session caps with queueing, idle reaping, and scheduled recycle for servers holding a credential they can only refresh interactively.## A note on 1.0.0This starts at 1.0.0 rather than 0.x because the bridge, its config format and its HTTP surface were shaken out against a live 15-server setup used daily, not only a test suite. The `Hardening` section of the CHANGELOG lists the defects that came out of that — process-tree teardown on Windows, keep-alive socket lifetimes, empty token caching, and several others that are traps for anyone building something similar.Tested on Linux, macOS and Windows against Node 20 and 22.Full notes: CHANGELOG.md

Full changelog

https://github.com/girishkvs/mcp-pacemaker/blob/main/CHANGELOG.md