Skip to content

Releases: ryanameier/kimi-swarm-bridge

v0.5.1

Choose a tag to compare

@ryanameier ryanameier released this 26 Sep 00:04
ea20ebb
  • Per-worker research time budget: workers label their web_search and read_page calls with
    their item; after 120s the results ask that worker to write up what it has, and after 180s new
    lookups are refused (KIMI_WORKER_SOFT_BUDGET_S, KIMI_WORKER_HARD_BUDGET_S, 0 = off). In a
    live 20-worker run, 19 workers finished within about 2.5 minutes and one took about 2 more.
  • Prompt: if the coordinator needs the workers' details, it reads all section files in one shell
    command instead of one at a time (seen live: about 20 separate reads, 1.5–2 minutes).
  • Prompt: after the workers finish, the coordinator writes the summary and recommendations itself
    instead of launching one more agent to parse and finish the report (seen in a live run: 20
    workers done in about 2 minutes, then a single finishing agent ran alone for 5+ minutes).
  • skills/kimi-swarm: a Claude skill that guides Claude to offer Kimi for big, independent parts of a
    request in every Claude app, including ones that load connector tools on demand and never show
    the connector's instructions (the desktop app, Claude Code). Install it once per user.
  • Removed glama.json and the last Glama mention; the project is no longer listed on Glama.
  • The "offer Kimi for part of a request" guidance is also in the kimi_delegate_task tool
    description, for clients that show tool descriptions but not MCP server instructions.

Kimi Swarm skill: download kimi-swarm.zip below. Team/Enterprise owners can add it for everyone under Organization settings → Plugins & skills; individual users upload it under their own skills. See docs/cloudflare-deploy.md.

v0.5.0

Choose a tag to compare

@ryanameier ryanameier released this 25 Sep 21:50
de99cad
  • Claude offers Kimi for independent parts of a request: when part of a request is substantial
    and independent of the rest, Claude asks whether to hand it to Kimi so both parts run at once
    (server instructions). Per-user preference offerKimi in kimi_swarm_settings: ask
    (default), auto or off, changeable from chat.
  • Stalled ai& model calls are cut off: no response headers within 60s, or no data for 90s in a
    streaming response, errors the call so Kimi retries it instead of a worker hanging (seen once:
    a 15-worker swarm stuck on one worker for over 4 minutes). Logged as outcome in timing events.
  • Prompt: independent research items get one worker each when they fit under the ceiling.
  • Organization-wide ai& concurrency limit: every employee's model calls pass through one
    AiandGate Durable Object that keeps requests in flight under AIAND_CONCURRENCY_LIMIT
    (setup KIMI_AIAND_CONCURRENCY, default 100 = ai&'s starting limit, 0 = off). Extra requests
    queue instead of getting HTTP 429. GET/POST /admin/aiand-limit shows usage and changes the
    limit live.
  • Timing logs ({"event":"timing"}) for every model request (queue wait, time to first byte,
    total, ai& inference time), search, browser render and, in log/allowlist mode, page fetch.
  • The coordinator writes the report frame and runs kimi-assemble in one command.
  • Faster swarms: the coordinator is shown the exact AgentSwarm call shape (its first launch was
    often rejected and retried), workers get a soft tool-call budget by depth (5 / 8 / 14) so one
    worker can't hold up the swarm, and read_page stops waiting for a slow page 8s after half the
    batch is done (reader model timeout 20s, browser 15s).
  • kimi-assemble: joins the workers' section files into one document (one contiguous table,
    sections in order) without a model. In a 15-worker run the coordinator had spent 4 minutes
    repairing a hand-assembled report.
  • Cloudflare defaults for new deployments: 20 agents per task (Kimi uses fewer when it can), 20
    running at once, and EGRESS_MODE=log. Existing deployments keep their settings on re-run.
  • Setup writes cloudflare/employee-guide.<worker>.md, the employee guide with the connector URL
    and domain filled in. The guide gained copy-paste starter prompts and the one extra step for
    personal Claude Pro/Max accounts.
  • README cleanup: removed the Glama-hosted Claude Desktop walkthrough and pilot wording. The
    Claude Desktop wrapper in scripts/claude-desktop/ now needs KIMI_MCP_URL and reads the token
    from the Keychain item kimi-swarm-mcp (was kimi-swarm-glama, with a Glama URL hardcoded).
  • README: benchmarks of Kimi Swarm vs Claude on web research briefs (time, completeness, cost) and
    what limits scaling past them.
  • kimi_model_settings: users switch the ai& models Kimi uses from chat, separately for the
    coordinator and the AgentSwarm workers (for example a cheaper worker model). The tool lists the
    ai& models with live prices; admins can narrow the list with KIMI_ALLOWED_MODELS. Choices are
    per user, persist, and are applied through Kimi Code's config (hot reloaded, no restart).
  • Task results report token usage per agent and an estimated USD cost from ai& prices.
  • Kimi is told to use the fewest workers that do the task well, overriding Kimi Code's default
    guidance to maximize agents; the per-user ceiling stays an upper bound.
  • The admin self-test reports which bridge build a container runs.
  • Default model is zai-org/glm-5.3 (deployment setting AIAND_MODEL, setup KIMI_MODEL);
    model capabilities (for example image input) come from the ai& catalog. New deployments start
    at 4 agents with a user-adjustable cap of 20.
  • Failed tasks report why (failureReason, from Kimi's turn record), for example exhausted model
    credits.
  • Web research replaces Firecrawl: web_search (Brave Search API, BRAVE_API_KEY) and
    read_page, which fetches a page and returns only the requested facts via a small reader model
    (KIMI_READER_MODEL), with Cloudflare Browser Rendering for JavaScript pages. The Worker
    attaches the Brave key and retries Brave rate limits. Firecrawl and FIRECRAWL_API_KEY are
    removed.
  • Lower token use per step: small web tool definitions, Kimi compacts context at a 128k window
    (KIMI_CONTEXT_WINDOW), and workers keep notes and return concise summaries.
  • docs/using-kimi-swarm.md: a one-page guide for employees.

See the README benchmarks and docs/cloudflare-deploy.md for setup and the new admin settings.

v0.4.0: organization deployment on Cloudflare

Choose a tag to compare

@ryanameier ryanameier released this 25 Sep 14:01
cafa976

Organization deployment on Cloudflare: every employee gets Kimi Swarm in Claude with their own
isolated workspace, signing in through the organization's identity provider. Admin guide:
docs/cloudflare-deploy.md.

Cloudflare edition (cloudflare/)

  • One Sandbox container per employee behind Cloudflare Access OIDC sign-in
    (workers-oauth-provider). The Access policy decides who gets Kimi Swarm.
  • npm run setup: one-command, re-runnable setup (KV, R2, secrets, regions, limits, deploy,
    health check). npm run smoke: post-deploy checks.
  • Persistence: /workspace and Kimi state are backed up to R2 and restored when a container
    starts. Job records are backed up before a delegated task is acknowledged. Dependencies and
    caches are excluded, and /workspace above BACKUP_MAX_MB is skipped (reported to admins).
  • Keys stay out of containers. The Worker intercepts requests to ai& and Firecrawl and attaches
    the real keys; containers only hold placeholders.
  • Outbound policy EGRESS_MODE: open (default), log (every outbound HTTP(S) request is
    logged), or allowlist (EGRESS_ALLOWLIST).
  • Per-employee daily ai& request budget (AIAND_DAILY_REQUEST_LIMIT, default 3000). Account-level
    ai& or Firecrawl errors (for example exhausted credits) are logged as egress-upstream-error.
  • The Worker answers the MCP handshake and tool list from a snapshot, so connecting does not
    wake a sleeping container and sessions survive container restarts.
  • Containers stay awake while Kimi works (up to 6 hours unattended) and sleep after 30 minutes
    idle.
  • Admin endpoints: list employees, status and usage, backup, restart, in-container self-test,
    offboarding (revokes sign-ins, destroys the container, deletes its state and backups), and
    backup listing/deletion.
  • Backups are named after their sandbox; each directory keeps its two newest backups, and setup
    adds a 90-day R2 expiry rule for employees who stop using Kimi Swarm.
  • Container workbench: python3, pypdf, reportlab, python-docx, openpyxl, pillow, poppler, git,
    zip, ripgrep, jq; firecrawl-mcp for web search and scraping.

Bridge

  • Files: signed single-use upload links and download links (kimi_create_upload_links,
    kimi_create_download_links, kimi_list_files; the singular names remain as aliases).
    MCP server instructions tell Claude to move chat attachments to /workspace/inputs and to
    bring deliverables from /workspace/outputs back into the conversation.
  • kimi_file_panel: an in-chat upload/download panel (MCP App) for hosts other than Claude.
  • kimi_swarm_settings: each user sets their AgentSwarm ceiling from chat, up to the
    deployment cap. Kimi still decides how many workers a task needs; SWARM_CONCURRENCY
    bounds how many run at once.
  • Waits are clamped by KIMI_MAX_WAIT_MS so hosted clients get a timeout status instead of a
    dropped call. Prompts can name the coordinator (KIMI_COORDINATOR_NAME).
  • Stateless Streamable HTTP mode for fronting proxies (x-kimi-mcp-mode: stateless).

Fixes

  • /authorize returns 400 instead of 500 for unknown OAuth clients.
  • The Codex plugin bundle includes kimi_swarm_settings.

Project

  • CI runs typecheck and tests for the bridge and the Cloudflare Worker.

v0.3.4 — Durable jobs and long-task recovery

Choose a tag to compare

@ryanameier ryanameier released this 24 Sep 00:06

v0.3.4 — Durable jobs and long-task recovery

This release adds persistent connector-owned job recovery for long-running Kimi and native AgentSwarm work.

Highlights

  • Adds a durable SQLite job registry for hosted deployments.
  • Adds kimi_recent_jobs for recovery after client timeout, disconnect, reconnect, or bridge restart.
  • Returns a durable jobId from delegation when durable jobs are configured.
  • Persists the Kimi sessionId, latest promptId, job status, cached result/error data, workspace, swarm mode, and timestamps.
  • Enforces connector ownership for session-oriented operations when durable jobs are enabled.
  • Synchronizes durable job state across wait, continuation, abort, handoff, and review-package flows.
  • Reconciles durable status from Kimi during direct kimi_get_handoff recovery.
  • Preserves the durable registry across process restarts and supports recovery with a fresh handler/runtime.
  • Updates MCP/Glama tool descriptions so long-running work prefers the asynchronous recovery path:
    kimi_delegate_task → kimi_wait_until_idle → kimi_get_handoff.
  • Keeps kimi_delegate_and_wait for shorter work, while making timeout/disconnect recovery explicit and duplicate-safe.
  • Updates the generated plugin bundle and MCP server version to 0.3.4.

Durable job configuration

Durable jobs are enabled only when both are configured:

KIMI_ORGANIZATION_ID=<customer-or-organization-id>
KIMI_CONNECTOR_INSTANCE_ID=<stable-connector-instance-id>

The default registry database is:

/data/kimi-swarm-bridge/jobs.sqlite

KIMI_JOB_DB_PATH can override that path.

Hosted deployments should place the database on persistent storage.

Recovery behavior

For long-running work:

  1. Call kimi_delegate_task.
  2. Keep the returned jobId and sessionId.
  3. If the client disconnects or times out, call kimi_recent_jobs.
  4. Recover the existing bound sessionId.
  5. Call kimi_wait_until_idle.
  6. When idle, call kimi_get_handoff or kimi_review_package.

A wait timeout does not abort the underlying Kimi session.

The durable registry is an ownership and recovery index; Kimi remains the source of truth for the live session and transcript.

Ownership behavior

When durable jobs are configured, session-oriented tools reject sessions that are not owned by the configured connector identity.

Job IDs and Kimi session IDs are not treated as authorization credentials.

The current deployment model remains one isolated hosted connector/runtime per customer organization. Hardened shared multi-tenant identity and workspace isolation remain subsequent deployment work.

Native AgentSwarm evidence

kimi_get_handoff continues to return a fresh structured swarmEvidence snapshot so final native AgentSwarm execution can be verified from Kimi wire/session evidence rather than model prose.

Validation

Release gate completed with:

  • 21 test files passed
  • 309 tests passed
  • 1 test skipped
  • TypeScript typecheck passed
  • generated plugin build passed
  • git diff --check passed
  • package, MCP server, and generated plugin all report version 0.3.4

v0.3.3 — Claude Desktop bridge and final swarm evidence

Choose a tag to compare

@ryanameier ryanameier released this 20 Sep 22:47

What's changed

Claude Desktop support

  • Added a validated local stdio → Streamable HTTP bridge for Claude Desktop.
  • Added macOS Keychain-based Glama token retrieval so credentials do not need to live in claude_desktop_config.json.
  • Added reusable setup scripts under scripts/claude-desktop/.
  • Added README instructions covering:
    • remote MCP verification
    • Keychain setup
    • Claude Desktop configuration
    • /workspace usage for hosted jobs
    • managed ai& model handling
    • long-running AgentSwarm timeout recovery

Final AgentSwarm evidence retrieval

  • kimi_get_handoff now returns a fresh structured swarmEvidence snapshot.
  • This allows long-running swarm jobs to be verified after:
    1. kimi_delegate_and_wait times out
    2. kimi_wait_until_idle reaches idle
    3. kimi_get_handoff retrieves final wire evidence
  • No second prompt or second AgentSwarm invocation is required.

Validation

  • 20 test files passed
  • 280 tests passed, 1 skipped
  • TypeScript typecheck passed
  • Production build passed
  • Generated plugin bundle updated

v0.3.2 — MCP tool quality metadata

Choose a tag to compare

@ryanameier ryanameier released this 13 Sep 22:58

MCP tool quality improvements

  • Adds substantive titles and descriptions for all 11 public MCP tools.
  • Adds parameter-level descriptions and usage guidance.
  • Adds MCP behavior annotations for read-only, destructive, and idempotent behavior.
  • Migrates tool registration to registerTool(...) metadata.
  • Adds glama.json maintainer metadata.
  • Preserves the existing Kimi AgentSwarm execution and structured swarm-evidence behavior.
  • Updates package, MCP server, and Codex plugin metadata to version 0.3.2.

Validation:

  • TypeScript type-check passes.
  • 19 test files passed.
  • 277 tests passed, 1 skipped, 0 failed.

v0.3.1 — Native swarm execution evidence

Choose a tag to compare

@ryanameier ryanameier released this 13 Sep 20:52

Adds structured native Kimi AgentSwarm execution evidence to the bridge.

  • Detects actual AgentSwarm tool calls from Kimi wire records
  • Reports requested, observed, and completed worker counts
  • Reports coordinator provider/model telemetry
  • Reports per-worker provider/model telemetry
  • Fails safely when wire evidence is unavailable
  • Includes regression tests for real-shaped Kimi wire events

Validated against Kimi Code 0.42.0 with:

  • 1 native AgentSwarm call
  • 4 requested workers
  • 4 completed workers
  • coordinator on openai / moonshotai/kimi-k3
  • all 4 workers on openai / moonshotai/kimi-k3