Skip to content

v0.11.0 — a process that says why it is leaving, You.com and DC Hub

Choose a tag to compare

@keysersoft keysersoft released this 22 Sep 10:23
· 100 commits to main since this release
e5089ad

The process now tells you it is in trouble, and gets out of the way

v0.10.0 shipped 66 connectors. This release is mostly about what happens when the thing is running at three in the morning.

Between 20 and 22 September the cloud backend hit Node's default heap ceiling four times in one day, and each time it took the whole container down — frontend, every tenant's MCP endpoint, the lot. We found out because a third-party directory emailed to say our connector looked unhealthy. Nothing in the stack had noticed.

Six changes, all of which a self-hoster gets:

  • One vitals line every 30 seconds — heap, RSS, both limits, event-loop delay, active handles by type, and live gauges: MCP requests in flight, unclosed handlers, tool calls in flight, aborted streams, registered tools. The first thing it showed us in production was a customer connector failing every call in 20 ms, in a loop, invisible until then.
  • /health reads the heap. It answers 503 above HEALTH_HEAP_PERCENT (default 90) and puts the numbers in the JSON even when it is up. Before this, /health returned 200 through eight-second GC cycles with average mu = 0.001 — a process already serving nobody.
  • A graceful exit instead of the death spiral. After HEAP_EXIT_CONSECUTIVE ticks (3) above HEAP_EXIT_PERCENT (90) of the heap limit, or above HEAP_RSS_EXIT_PERCENT (90) of the cgroup limit, the process logs why and exits so a supervisor can restart it. The alternatives are minutes of stop-the-world GC, then a hard abort — or a kernel SIGKILL with no log line at all.
  • An optional heap snapshot, off by default (HEAP_SNAPSHOT_PERCENT=0), written to HEAP_SNAPSHOT_DIR.
  • A bounded shutdown, SHUTDOWN_TIMEOUT_MS, so a stuck handler cannot hold a restart open.
  • Backend and frontend as separate containers in the cloud compose, with memory limits. A backend crash restarts the backend. The single-container self-host path is unchanged and still supported — but start.sh is rewritten to serve both shapes, so read it if you have patched it.

Turn the snapshot on deliberately, and know what it costs. Ours was set to fire at 60 % of a 4 GB heap inside a 5 GB container. V8 builds a snapshot entirely in memory first — roughly another heap — so RSS went through the cgroup limit and the kernel killed the process before a byte reached disk. Nineteen times in two and a half hours, each one leaving a 0-byte file. It turned one crash an hour into one every five minutes. The guard now watches RSS against the cgroup limit as well as the heap against V8's, and refuses a snapshot unless rss + heap fits in 85 % of the limit.

The leak itself is not fixed. #659 is open and honest about what we know: it is not the volume, not aborted SSE streams (three independent confirmations, the last from a live experiment), not tool result size, not slow calls. The newest evidence says it arrives as a step — 583 MB to 2.3 GB in sixty seconds, then a floor that does not move for sixteen hours — rather than as a drift. What this release ships is survival, not a cure: limits that are ours instead of Node's, a process that says why it is leaving, and the counters to catch the next one.

Two connectors — 255 → 257

  • You.com (#638, from the community) — youcom_search and youcom_news. The first web-search connector in the catalog.
  • DC Hub (#653) — European data centre and grid intelligence: facilities, capacity, power, grid connection queues, energy prices, site scores. 16 tools.

The DC Hub adapter shipped describing its own API wrongly, and the corrections are worth naming because the wrong shape is the dangerous one. We re-ran all sixteen tools against the live API:

  • Four tools require a key, not two. dchub_pipeline and dchub_energy_prices answer 403 without one; dchub_site_score needs a paid plan.
  • dchub_facility_detail does not fail — it answers 200 with less. Contacts missing, coordinates rounded, and only _gated: true in the payload to say so. A model reading that record quotes it as complete. The tool now says which fields were withheld and how.
  • The anonymous allowance is ten calls a day, not five.
  • Their 403 body is addressed to the model, with relay_verbatim, marketing copy and a payment link. We did not filter it — the same field carries the useful "which plan do you need" — but the adapter now tells the model that half of that response is advertising, not instruction. If you write adapters for third-party APIs, this is a category worth watching for.

The live spec now walks all sixteen tools. It used to walk the three that worked, which is why the drift survived review.

Smaller, but you will meet them

  • A bodyTemplate that will not parse now says which placeholder was empty (#651), instead of a JSON parse error pointing at a column number.
  • The admin UI and the marketplace filters work on a phone (#660, #680).
  • German, Simplified Chinese and Japanese READMEs (#609, from the community), with a language switcher. They are wired into scripts/adapter-count.mjs --check with an expected match count per file, so a reworded sentence fails the guard rather than silently dropping out of it.
  • The adapter count no longer lives under .github/ (#658). The Docker Hub description is computed at publish time, and a guard now rejects any count quoted in a workflow or issue template. The point was not tidiness: a count in a workflow file gave every adapter PR — including every fork's — a legitimate reason to edit .github/, which is exactly the noise a malicious change hides in.
  • Registry publishing confirms by isLatest, not by list order (#649). The registry returns every version you have ever published, unordered; reading "the last one" happened to work for v0.9.0 and would have failed every release after it.
  • Eleven dependency PRs, including @rekog/mcp-nest and @rekog/mcp-nest-auth 2.0.6, which are runtime code.

Upgrading

Nothing to do. No migrations, no config changes, and every new knob has a default that matches the old behaviour except the two that are the point: /health can now answer 503 under heap pressure, and the process can exit on purpose. If your orchestrator treats either as a failure to page a human, that is the intended reading.

variable default what it does
VITALS_INTERVAL_SEC 30 vitals cadence; 0 disables the line and the guard
HEAP_WARN_PERCENT 75 warn once above this
HEAP_EXIT_PERCENT 90 heap share that starts the exit countdown
HEAP_RSS_EXIT_PERCENT 90 same, for RSS against the cgroup limit
HEAP_EXIT_CONSECUTIVE 3 ticks above the line before exiting
HEAP_SNAPSHOT_PERCENT 0 (off) write one heap snapshot above this — read the warning above first
HEAP_SNAPSHOT_DIR /tmp/anythingmcp-diagnostics where it lands; mount a volume if you want to keep it
HEALTH_HEAP_PERCENT 90 heap share at which /health reports 503
SHUTDOWN_TIMEOUT_MS 20000 cap on graceful shutdown, in ms

257 adapters, 2,385 tools, 20 of which still need no API key.

Full changelog: v0.10.0...v0.11.0