Releases: ryanameier/kimi-swarm-bridge
Release list
v0.5.1
- Per-worker research time budget: workers label their
web_searchandread_pagecalls with
their item; after 120s the results ask that worker to write up what it has, and after 180s new
lookups are refused (KIMI_WORKER_SOFT_BUDGET_S,KIMI_WORKER_HARD_BUDGET_S, 0 = off). In a
live 20-worker run, 19 workers finished within about 2.5 minutes and one took about 2 more. - Prompt: if the coordinator needs the workers' details, it reads all section files in one shell
command instead of one at a time (seen live: about 20 separate reads, 1.5–2 minutes). - Prompt: after the workers finish, the coordinator writes the summary and recommendations itself
instead of launching one more agent to parse and finish the report (seen in a live run: 20
workers done in about 2 minutes, then a single finishing agent ran alone for 5+ minutes). skills/kimi-swarm: a Claude skill that guides Claude to offer Kimi for big, independent parts of a
request in every Claude app, including ones that load connector tools on demand and never show
the connector's instructions (the desktop app, Claude Code). Install it once per user.- Removed
glama.jsonand the last Glama mention; the project is no longer listed on Glama. - The "offer Kimi for part of a request" guidance is also in the
kimi_delegate_tasktool
description, for clients that show tool descriptions but not MCP server instructions.
Kimi Swarm skill: download kimi-swarm.zip below. Team/Enterprise owners can add it for everyone under Organization settings → Plugins & skills; individual users upload it under their own skills. See docs/cloudflare-deploy.md.
v0.5.0
- Claude offers Kimi for independent parts of a request: when part of a request is substantial
and independent of the rest, Claude asks whether to hand it to Kimi so both parts run at once
(server instructions). Per-user preferenceofferKimiinkimi_swarm_settings:ask
(default),autooroff, changeable from chat. - Stalled ai& model calls are cut off: no response headers within 60s, or no data for 90s in a
streaming response, errors the call so Kimi retries it instead of a worker hanging (seen once:
a 15-worker swarm stuck on one worker for over 4 minutes). Logged asoutcomein timing events. - Prompt: independent research items get one worker each when they fit under the ceiling.
- Organization-wide ai& concurrency limit: every employee's model calls pass through one
AiandGateDurable Object that keeps requests in flight underAIAND_CONCURRENCY_LIMIT
(setupKIMI_AIAND_CONCURRENCY, default 100 = ai&'s starting limit, 0 = off). Extra requests
queue instead of getting HTTP 429.GET/POST /admin/aiand-limitshows usage and changes the
limit live. - Timing logs (
{"event":"timing"}) for every model request (queue wait, time to first byte,
total, ai& inference time), search, browser render and, in log/allowlist mode, page fetch. - The coordinator writes the report frame and runs
kimi-assemblein one command. - Faster swarms: the coordinator is shown the exact AgentSwarm call shape (its first launch was
often rejected and retried), workers get a soft tool-call budget by depth (5 / 8 / 14) so one
worker can't hold up the swarm, andread_pagestops waiting for a slow page 8s after half the
batch is done (reader model timeout 20s, browser 15s). kimi-assemble: joins the workers' section files into one document (one contiguous table,
sections in order) without a model. In a 15-worker run the coordinator had spent 4 minutes
repairing a hand-assembled report.- Cloudflare defaults for new deployments: 20 agents per task (Kimi uses fewer when it can), 20
running at once, andEGRESS_MODE=log. Existing deployments keep their settings on re-run. - Setup writes
cloudflare/employee-guide.<worker>.md, the employee guide with the connector URL
and domain filled in. The guide gained copy-paste starter prompts and the one extra step for
personal Claude Pro/Max accounts. - README cleanup: removed the Glama-hosted Claude Desktop walkthrough and pilot wording. The
Claude Desktop wrapper inscripts/claude-desktop/now needsKIMI_MCP_URLand reads the token
from the Keychain itemkimi-swarm-mcp(waskimi-swarm-glama, with a Glama URL hardcoded). - README: benchmarks of Kimi Swarm vs Claude on web research briefs (time, completeness, cost) and
what limits scaling past them. kimi_model_settings: users switch the ai& models Kimi uses from chat, separately for the
coordinator and the AgentSwarm workers (for example a cheaper worker model). The tool lists the
ai& models with live prices; admins can narrow the list withKIMI_ALLOWED_MODELS. Choices are
per user, persist, and are applied through Kimi Code's config (hot reloaded, no restart).- Task results report token usage per agent and an estimated USD cost from ai& prices.
- Kimi is told to use the fewest workers that do the task well, overriding Kimi Code's default
guidance to maximize agents; the per-user ceiling stays an upper bound. - The admin self-test reports which bridge build a container runs.
- Default model is
zai-org/glm-5.3(deployment settingAIAND_MODEL, setupKIMI_MODEL);
model capabilities (for example image input) come from the ai& catalog. New deployments start
at 4 agents with a user-adjustable cap of 20. - Failed tasks report why (
failureReason, from Kimi's turn record), for example exhausted model
credits. - Web research replaces Firecrawl:
web_search(Brave Search API,BRAVE_API_KEY) and
read_page, which fetches a page and returns only the requested facts via a small reader model
(KIMI_READER_MODEL), with Cloudflare Browser Rendering for JavaScript pages. The Worker
attaches the Brave key and retries Brave rate limits. Firecrawl andFIRECRAWL_API_KEYare
removed. - Lower token use per step: small web tool definitions, Kimi compacts context at a 128k window
(KIMI_CONTEXT_WINDOW), and workers keep notes and return concise summaries. docs/using-kimi-swarm.md: a one-page guide for employees.
See the README benchmarks and docs/cloudflare-deploy.md for setup and the new admin settings.
v0.4.0: organization deployment on Cloudflare
Organization deployment on Cloudflare: every employee gets Kimi Swarm in Claude with their own
isolated workspace, signing in through the organization's identity provider. Admin guide:
docs/cloudflare-deploy.md.
Cloudflare edition (cloudflare/)
- One Sandbox container per employee behind Cloudflare Access OIDC sign-in
(workers-oauth-provider). The Access policy decides who gets Kimi Swarm. npm run setup: one-command, re-runnable setup (KV, R2, secrets, regions, limits, deploy,
health check).npm run smoke: post-deploy checks.- Persistence:
/workspaceand Kimi state are backed up to R2 and restored when a container
starts. Job records are backed up before a delegated task is acknowledged. Dependencies and
caches are excluded, and/workspaceaboveBACKUP_MAX_MBis skipped (reported to admins). - Keys stay out of containers. The Worker intercepts requests to ai& and Firecrawl and attaches
the real keys; containers only hold placeholders. - Outbound policy
EGRESS_MODE:open(default),log(every outbound HTTP(S) request is
logged), orallowlist(EGRESS_ALLOWLIST). - Per-employee daily ai& request budget (
AIAND_DAILY_REQUEST_LIMIT, default 3000). Account-level
ai& or Firecrawl errors (for example exhausted credits) are logged asegress-upstream-error. - The Worker answers the MCP handshake and tool list from a snapshot, so connecting does not
wake a sleeping container and sessions survive container restarts. - Containers stay awake while Kimi works (up to 6 hours unattended) and sleep after 30 minutes
idle. - Admin endpoints: list employees, status and usage, backup, restart, in-container self-test,
offboarding (revokes sign-ins, destroys the container, deletes its state and backups), and
backup listing/deletion. - Backups are named after their sandbox; each directory keeps its two newest backups, and setup
adds a 90-day R2 expiry rule for employees who stop using Kimi Swarm. - Container workbench: python3, pypdf, reportlab, python-docx, openpyxl, pillow, poppler, git,
zip, ripgrep, jq;firecrawl-mcpfor web search and scraping.
Bridge
- Files: signed single-use upload links and download links (
kimi_create_upload_links,
kimi_create_download_links,kimi_list_files; the singular names remain as aliases).
MCP server instructions tell Claude to move chat attachments to/workspace/inputsand to
bring deliverables from/workspace/outputsback into the conversation. kimi_file_panel: an in-chat upload/download panel (MCP App) for hosts other than Claude.kimi_swarm_settings: each user sets their AgentSwarm ceiling from chat, up to the
deployment cap. Kimi still decides how many workers a task needs;SWARM_CONCURRENCY
bounds how many run at once.- Waits are clamped by
KIMI_MAX_WAIT_MSso hosted clients get a timeout status instead of a
dropped call. Prompts can name the coordinator (KIMI_COORDINATOR_NAME). - Stateless Streamable HTTP mode for fronting proxies (
x-kimi-mcp-mode: stateless).
Fixes
/authorizereturns 400 instead of 500 for unknown OAuth clients.- The Codex plugin bundle includes
kimi_swarm_settings.
Project
- CI runs typecheck and tests for the bridge and the Cloudflare Worker.
v0.3.4 — Durable jobs and long-task recovery
v0.3.4 — Durable jobs and long-task recovery
This release adds persistent connector-owned job recovery for long-running Kimi and native AgentSwarm work.
Highlights
- Adds a durable SQLite job registry for hosted deployments.
- Adds
kimi_recent_jobsfor recovery after client timeout, disconnect, reconnect, or bridge restart. - Returns a durable
jobIdfrom delegation when durable jobs are configured. - Persists the Kimi
sessionId, latestpromptId, job status, cached result/error data, workspace, swarm mode, and timestamps. - Enforces connector ownership for session-oriented operations when durable jobs are enabled.
- Synchronizes durable job state across wait, continuation, abort, handoff, and review-package flows.
- Reconciles durable status from Kimi during direct
kimi_get_handoffrecovery. - Preserves the durable registry across process restarts and supports recovery with a fresh handler/runtime.
- Updates MCP/Glama tool descriptions so long-running work prefers the asynchronous recovery path:
kimi_delegate_task→kimi_wait_until_idle→kimi_get_handoff. - Keeps
kimi_delegate_and_waitfor shorter work, while making timeout/disconnect recovery explicit and duplicate-safe. - Updates the generated plugin bundle and MCP server version to
0.3.4.
Durable job configuration
Durable jobs are enabled only when both are configured:
KIMI_ORGANIZATION_ID=<customer-or-organization-id>
KIMI_CONNECTOR_INSTANCE_ID=<stable-connector-instance-id>
The default registry database is:
/data/kimi-swarm-bridge/jobs.sqlite
KIMI_JOB_DB_PATH can override that path.
Hosted deployments should place the database on persistent storage.
Recovery behavior
For long-running work:
- Call
kimi_delegate_task. - Keep the returned
jobIdandsessionId. - If the client disconnects or times out, call
kimi_recent_jobs. - Recover the existing bound
sessionId. - Call
kimi_wait_until_idle. - When idle, call
kimi_get_handofforkimi_review_package.
A wait timeout does not abort the underlying Kimi session.
The durable registry is an ownership and recovery index; Kimi remains the source of truth for the live session and transcript.
Ownership behavior
When durable jobs are configured, session-oriented tools reject sessions that are not owned by the configured connector identity.
Job IDs and Kimi session IDs are not treated as authorization credentials.
The current deployment model remains one isolated hosted connector/runtime per customer organization. Hardened shared multi-tenant identity and workspace isolation remain subsequent deployment work.
Native AgentSwarm evidence
kimi_get_handoff continues to return a fresh structured swarmEvidence snapshot so final native AgentSwarm execution can be verified from Kimi wire/session evidence rather than model prose.
Validation
Release gate completed with:
- 21 test files passed
- 309 tests passed
- 1 test skipped
- TypeScript typecheck passed
- generated plugin build passed
git diff --checkpassed- package, MCP server, and generated plugin all report version
0.3.4
v0.3.3 — Claude Desktop bridge and final swarm evidence
What's changed
Claude Desktop support
- Added a validated local stdio → Streamable HTTP bridge for Claude Desktop.
- Added macOS Keychain-based Glama token retrieval so credentials do not need to live in
claude_desktop_config.json. - Added reusable setup scripts under
scripts/claude-desktop/. - Added README instructions covering:
- remote MCP verification
- Keychain setup
- Claude Desktop configuration
/workspaceusage for hosted jobs- managed ai& model handling
- long-running AgentSwarm timeout recovery
Final AgentSwarm evidence retrieval
kimi_get_handoffnow returns a fresh structuredswarmEvidencesnapshot.- This allows long-running swarm jobs to be verified after:
kimi_delegate_and_waittimes outkimi_wait_until_idlereachesidlekimi_get_handoffretrieves final wire evidence
- No second prompt or second AgentSwarm invocation is required.
Validation
- 20 test files passed
- 280 tests passed, 1 skipped
- TypeScript typecheck passed
- Production build passed
- Generated plugin bundle updated
v0.3.2 — MCP tool quality metadata
MCP tool quality improvements
- Adds substantive titles and descriptions for all 11 public MCP tools.
- Adds parameter-level descriptions and usage guidance.
- Adds MCP behavior annotations for read-only, destructive, and idempotent behavior.
- Migrates tool registration to
registerTool(...)metadata. - Adds
glama.jsonmaintainer metadata. - Preserves the existing Kimi AgentSwarm execution and structured swarm-evidence behavior.
- Updates package, MCP server, and Codex plugin metadata to version 0.3.2.
Validation:
- TypeScript type-check passes.
- 19 test files passed.
- 277 tests passed, 1 skipped, 0 failed.
v0.3.1 — Native swarm execution evidence
Adds structured native Kimi AgentSwarm execution evidence to the bridge.
- Detects actual
AgentSwarmtool calls from Kimi wire records - Reports requested, observed, and completed worker counts
- Reports coordinator provider/model telemetry
- Reports per-worker provider/model telemetry
- Fails safely when wire evidence is unavailable
- Includes regression tests for real-shaped Kimi wire events
Validated against Kimi Code 0.42.0 with:
- 1 native AgentSwarm call
- 4 requested workers
- 4 completed workers
- coordinator on
openai/moonshotai/kimi-k3 - all 4 workers on
openai/moonshotai/kimi-k3