Skip to content

Releases: parisbs/codex-subagent-mcp

v0.3.0

Choose a tag to compare

@parisbs parisbs released this 18 Sep 01:32
fce7a63

Defaults that survive a fresh install, and results that mean what they say. Most of this came from spending a day measuring real delegations against the installed CLI: three places were asking Codex a question in the wrong directory, and two numbers in every result meant something other than what they looked like.

Upgrade notes

This is a minor release with breaking changes:

  • The sandbox ceiling now defaults to workspace-write. An unconfigured server used to let a call request danger-full-access, which removes the sandbox entirely, network included. Set CODEX_SUBAGENT_MAX_SANDBOX=danger-full-access if you need it.
  • The token line in a result changed shape, separating tokens this turn from tokens thread so far. A follow-up used to report the thread's running total as that call's cost.
  • working_dir: "" is refused instead of falling back to the server's own directory.

Highlights

  • A result now tells you what Codex actually applied, not only what was requested: the metadata line ends with applied=confirmed, differs or unconfirmed, and a sandbox recorded as wider than the one requested fails the delegation.
  • New CODEX_SUBAGENT_DEFAULT_SANDBOX, so someone who delegates edits all day sets it once instead of repeating sandbox: "workspace-write" on every call. read-only stays the built-in default.
  • The recursion guard, the preflight, the model catalog and codex_recommend all run in the delegation's working directory. Codex resolves configuration against the directory it runs in, so each of them was answering about somewhere else — and for the recursion guard, that was the one place a repository could have registered this server.
  • A thread survives a restart. A follow-up on a thread this server has no record of recovers the model, effort and directory from Codex's own session file, validated through the usual policy.
  • A read-only run is told what it cannot verify, after a measured run reported 80 failing tests that were sandbox artefacts, and that its shell has no network while the web-search tool works.
  • Results report the true command count, the effective working directory and uncached input.
  • Writing a delegation: what a delegation costs and how to shape one, measured rather than guessed — including an explicit list of what target_files, working_dir and read-only do not mean.
  • Releases are built and staged by GitHub Actions with npm provenance through trusted publishing, and reach users only when a maintainer approves the staged release with two-factor authentication.

codex_review was planned for this release and dropped after measurement: wrapping codex exec review cost more than a plain delegation on the same commit, found less, hides its usage in a subagent thread, and has no sandbox flag. The issue stays open with the numbers attached.

Verified against

  • Codex CLI 0.154.0 on macOS.
  • Release smoke test (real CLI, gpt-5.6-luna): 23 of 23 checks passed — preflight, catalog, recommendation, five refusals including the new ceiling and empty-directory ones, a read-only delegation with applied-settings confirmation, effort clamping, a follow-up that restated its model and separated this turn's tokens from the thread total, prompt-injection inertness, and a background job.
  • Build, tests and startup on Linux (Node 22, 24, 26), Windows (Node 22, 24) and macOS (Node 24). A real delegation has still never been run on Windows or Linux.

Full details in the changelog.

v0.2.0 — hardening, configuration awareness and delegation guardrails

Choose a tag to compare

@parisbs parisbs released this 14 Sep 21:31
bf88742

Hardening after the first review of 0.1.0, plus what real use and a close look at Codex's own configuration turned up. Every change touching the Codex CLI was checked against the real binary.

Upgrade notes

This is a minor release with breaking changes:

  • Node 22 or newer is required. Node 20 reached end-of-life on 2026-04-30.
  • Follow-ups restate the thread's model, effort and working directory. A resumed Codex session does not keep its model; without it, Codex picks one from the configuration of the directory it resumes in. For a thread this server has no record of (another server process, or after a restart), pass model explicitly or set CODEX_SUBAGENT_DEFAULT_MODEL.
  • auto_approve: true on codex_follow_up is refused. It was accepted and silently dropped.
  • More runs are reported as failed: a turn Codex marks as failed, and a clean exit with no answer.
  • codex_delegate refuses to run on the static fallback catalog.

Security

Two advisories are fixed in this release. Upgrade if you run 0.1.0.

  • GHSA-m9wq-wr2p-3rc4 — argument injection through thread_id in codex_follow_up (high).
  • GHSA-6946-2h8r-6372 — configured model and effort ceilings not enforced on codex_follow_up (medium).

Highlights

  • web_search: true finally works; it had failed argument parsing since 0.1.0.
  • A broken Codex config.toml is reported as config-error with Codex's own message, not as "not signed in".
  • Usage limits and other failed turns say why they failed instead of showing only an exit code.
  • Configuration warnings from Codex appear once, under "Codex notices", instead of as errors.
  • Recommendations respect CODEX_SUBAGENT_MAX_EFFORT.
  • Output retained per run is bounded, a malformed event no longer crashes the server, and the timeout holds even when a child's descendant keeps its pipes open.
  • Guardrails for use beyond programming chats: tool descriptions state that delegating sends content to OpenAI and spends your Codex usage, a delegation without a model asks Claude to confirm the model with you, results are framed as information rather than instructions, and a delegated run can no longer call this server again through Codex's own MCP configuration.
  • A new guide, Staying in control of delegation: client permissions, server ceilings, version ranges and your own rules for Claude.
  • Guidance on keeping this server's ceilings outside the working tree, and on how Codex's own configuration interacts with delegations (SECURITY.md).

Verified against

  • Codex CLI 0.154.0 on macOS.
  • Release smoke test (real CLI, gpt-5.6-luna at low): 14 of 14 checks passed — preflight, catalog, recommendations, refusals, a read-only delegation, a follow-up that restated its model and effort, effort adjustment, timeout, prompt-injection inertness and a background job.
  • Build, tests and startup on Linux (Node 22, 24, 26), Windows (Node 22, 24) and macOS (Node 24). A real delegation has not yet been run on Windows or Linux.

Full details in the changelog.

v0.1.0 — first public release

Choose a tag to compare

@parisbs parisbs released this 12 Sep 00:07
2f6e6ab

Available on npm: https://www.npmjs.com/package/codex-subagent-mcp

claude mcp add codex-subagent -- npx -y codex-subagent-mcp

First public release.

Added

  • Eight MCP tools: codex_doctor, list_codex_models, codex_recommend, codex_delegate,
    codex_follow_up, codex_job_status, codex_job_result and codex_job_cancel.
  • Model and reasoning effort as two independent axes: -m for raw capability, and
    model_reasoning_effort for how long the model deliberates. Conflating them was the original
    prototype's core bug.
  • A model catalog read from codex debug models at runtime and cached for ten minutes, so new
    models appear without a release here. The requested reasoning effort is clamped to what the chosen
    model actually supports.
  • A preflight that runs on every tool call and reports whether the CLI is installed, recent enough
    and signed in, with the installation commands for the detected platform. It never installs
    anything, and it never degrades to a plausible-looking answer when the CLI is unavailable.
  • read-only as the default sandbox. Writing is opt-in and explicit.
  • Two execution modes: blocking with MCP progress notifications, and background with a job id.
  • Delegation policy through five environment variables — CODEX_SUBAGENT_DEFAULT_MODEL,
    CODEX_SUBAGENT_DEFAULT_EFFORT, CODEX_SUBAGENT_ALLOWED_MODELS, CODEX_SUBAGENT_MAX_SANDBOX
    and CODEX_SUBAGENT_MAX_EFFORT. Configuration may only restrict; no setting makes delegations
    more permissive.
  • A quality contract injected into every prompt, rather than writing an AGENTS.md into somebody
    else's repository.
  • CODEX_BIN for a Codex executable that is not called codex or is not on PATH.

Security

  • The CLI is always invoked with spawn and an argv array, never a shell command string, and the
    prompt is written to the child's stdin. A prompt is attacker-influenced text; string interpolation
    into a shell would be a command-injection hole.
  • On Windows, a .cmd shim from a global npm install is reported as unsupported-shim rather than
    silently run through a shell, which would reintroduce that hole.

Verified against

  • Codex CLI 0.154.0 on macOS, including a real delegation end to end.
  • Build, tests and startup on Linux (Node 20, 24 and 26), Windows (Node 20 and 24) and macOS
    (Node 24). A real delegation has never been run on Windows or Linux; see
    docs/VERSIONING.md for what 1.0 requires.

Known limitations

  • The Codex sandbox restricts writes and network access, not reads. A delegation can read files
    outside the working directory, credentials included. The README explains the mitigation.
  • Background jobs live in memory and do not survive a server restart.