Skip to content

v2.16.1

Choose a tag to compare

@nokhodian nokhodian released this 24 Sep 11:01
· 668 commits to main since this release

[2.16.1] — 2026-09-24

Added

  • org_task and org_plan_graph nodes take a brief. The creator's instructions for a task — scope, acceptance criteria, paths, what failed last time, up to 4000 characters — are stored on the task and delivered in the same mailbox message as its title, every time it is dispatched: at creation, later when its dependencies complete, and again after a refused close or a checkpoint resume. Split children inherit it, and per-task skill suggestions still follow it. Until now org_task took only a title, so a coordinator sent the details in a follow-up org_send, which joined the dispatch only if it landed inside the 500 ms coalescing window; on the 2.16.0 release run it often did not — the publisher asked for the details twice, and the maintainer finished tasks before their briefs arrived and was then woken four more times (about 3M tokens).
  • org_tasks takes an optional taskId. It then returns just that task — status, result and latest evidence — instead of the whole DAG. On a long run the full listing is large enough to be spilled to a file, and on the 2.16.0 release run the coordinator read single tasks back out of those files with dd and python.
  • {{home}} and {{org_root}} work in a role's policy paths. policy.fileRead, policy.fileWrite, policy.sandbox.allowWrite and policy.sandbox.denyWrite now take the same placeholders as responsibilities, expanded when the daemon loads the org, so the file-tool roots and the OS sandbox see absolute paths. Before, only responsibilities were expanded, so a tracked config couldn't grant a role a scratch dir under the operator's home without hard-coding it. On the 2.16.0 release run that produced 17 "path escapes every root" denials for files under $HOME/monomind-release/<ver>/logs and $HOME/mrg-tmp. org validate reports an unknown placeholder in these fields as an error, naming the field.
  • Tool-call traces now say which turn a call belongs to, and org_send carries the sender's chain. (Fixes #327) Every call a role made carried the same chain_id and hop across all its turns, and a role's native org_send mail carried no trace at all, so mono-agent's loop control could not tell ten calls answering one request from a loop between roles — a loop running only through org_send was stopped after 64 calls. _meta.trace now has turn, a per-role count that is the same for every call in one turn and goes up by one each turn (checkpointed, so a resumed role continues it). A role's org_send, in the same org or across orgs, now starts with [trace <chain_id> hop=<hop+1>] for the sender's chain, replacing any trace line already in the body; the receiver adopts it, so A → B → A shows up as one chain with hops 1, 2, 3. Human, operator and dispatch mail is not stamped. The _meta.trace shape is documented in doc/concepts/org-runtime.md.
  • run_config.max_tool_rounds and a role's own max_tool_rounds set the tool-call round cap. (Fixes #326) Fence-protocol runtimes (every runtime but claude and vercel) stopped after 10 tool-call rounds per message, fixed. A role with a longer job, like the 13 automation calls in mono-agent's C-46 live gate, could not finish it in one turn. The default stays 10. Values must be a positive integer up to 200, and org validate rejects anything else.

Changed

  • A sandboxed org role can run node $SCRIPT again. The git policy's Bash classifier denies any command it can't read (node $X, python3 $F, $BIN …, env -i …, eval) for every role below policy.git: 'push', because without an OS sandbox such a command could run git at any level. It did this even when the role's Bash ran inside the SDK sandbox, which already enforces the level where git runs: 27 denials on the 2.16.0 release run. When a role's session actually got the sandbox (the runtime decision, never the config alone), those commands are now allowed. Every git call written out literally is still checked per level, including one after an unreadable part of the same command. A visible git call with a hidden subcommand (sh -c "git …", git $SUB, an alias, a guard override) still fails closed. Unsandboxed roles, mode: 'off', a host without the sandbox, and non-Claude runtimes behave exactly as before.
  • Org roles no longer see Claude Code's own scheduling, task-list, question and plan-mode tools. On the 2.16.0 release run a builder called ScheduleWakeup and AskUserQuestion and the captain created a TaskCreate("placeholder"). None of them does anything useful in a headless org session: nobody is there to answer the question, and a harness task or wakeup is invisible to the other roles, the org's task DAG and the daemon's watchdogs. Every claude-runtime role, at every policy.git level, now runs with AskUserQuestion, ScheduleWakeup, TaskCreate, TaskUpdate, TaskList, TaskGet, CronCreate, CronDelete, CronList, EnterPlanMode and ExitPlanMode in the SDK's disallowedTools. Their org equivalents are ask_human, org_task/org_tasks and org_task_block.
  • monomind init now sets up the memory database by default. Memory was only initialized as part of --start-all, so init --no-start-all (and init wizard) left monomind doctor reporting "Memory Database: Not initialized" until you ran monomind memory init yourself. Init now creates the same .swarm/memory.db (copied to .claude/memory.db) in every mode, --minimal included, whether or not services auto-start. An existing database is kept as it is, never recreated, so re-running init leaves stored entries in place. Pass --no-memory to opt out; --only-claude skips it because that mode writes no runtime state, and --skip-claude skips only the .claude/ copy. If the database can't be created, init still succeeds and prints a warning telling you to run monomind memory init. Doctor's fix hint for a missing database now says monomind memory init instead of memory configure, which never created one.
  • A role that hits the tool-call round cap is now told so. (#326) Its pending calls used to be dropped with only a bus notice, so the role stopped mid-task without knowing why. Now each of those calls comes back as a tool result saying the round cap was reached. The role then gets one wrap-up round whose calls run, to report what it finished and ask to be continued (for example with org_send). Calls after that round are dropped with the bus notice as before, and agent exec still reports stop_reason: "tool_round_cap".
  • The prompt hook can wait up to 10 s for the Jev decision model. Hosted Jev sometimes takes 1.5–9 s to answer, and MONOMIND_JEV_HOOK_TIMEOUT_MS was capped at 4000 because every hook process force-exits at 5 s. The cap is now 10000 (the default stays 1500), and while a Jev provider is configured the route hook's force-exit moves to that limit plus 1.5 s so the route is still recorded after a slow pick; every other hook, including the pre-bash/pre-write security gates, keeps 5 s. The trade-off: a slow model delays each prompt by up to the configured limit, and a timed-out pick still makes the hook skip Jev for 5 minutes.

Fixed

  • monomind init wrote one full copy of every shared skill per platform. .agents/skills is a single directory that codex, kimi, opencode, gemini, cursor and several other platforms all declare as their skill root, and each adapter wrapped the same body in its own skills:<platform>:<name> block — a default init left .agents/skills/mastermind-org/SKILL.md with three stacked copies (49 lines instead of ~16), and mastermind-idea/SKILL.md grew by 2,281 lines. A shared root now carries one co-owned skills:agents:<name> block. The next init, init upgrade or platforms install/upgrade folds existing per-platform blocks into that one block where the first sat, leaving text outside the blocks byte-for-byte and backing the file up to .monomind/backups/ first. Which platforms installed into the shared root is recorded in .monomind/platforms/shared-skills.json, so platforms uninstall of one platform keeps the block while another platform still uses it. Platform-specific roots such as .claude/skills keep their per-platform markers.
  • org report marked every role EXHAUSTED on a well-cached run. The per-role budget percentage compared each role's total tokens — cache reads included — against budget_tokens, which the runtime enforces on input+output only, so on the 2.16.0 release run the captain read "8391% of 500000 — EXHAUSTED" while it had really used about 13% of its budget. The percentage is now computed on the basis the policy enforces (input+output, or the billable total when run_config.budget_tokens_basis is "billable"), the basis is named on the line (13% of 500000 in+out), and cache tokens are still shown, labelled separately (41955000 tokens (65000 in+out, 41890000 cache)). --format json role rows gain uncachedTokens.
  • runtime.json still said every role was running after the org had stopped. The stop checkpoint is captured before sessions drain — deliberately, so unconsumed mail survives into a resume — which froze each live role's status at running in checkpoint.roleState, after org stop, org_complete and a normal exit alike. A role that was live at that moment is now recorded as stopped; resuming from the checkpoint still brings it back as running.
  • A detached org run log just stopped, with nothing saying how the run ended. org run now prints one final line on every exit path — org_complete, org stop, Ctrl-C/SIGTERM, the idle watchdog, a budget, or a crash — naming the outcome (complete, stopped, budget or error, with the cause), the wall time and the run's total cost from its usage events: org release run run-… ended — outcome: complete (achieved), wall time 1h57m12s, cost $63.63. The exit code is unchanged.
  • Org events went to a dashboard built from a deleted worktree, and heals left dashboards running on every port. The org event forwarder accepted whatever dashboard .monomind/control.json named as long as its pid was alive; on the 2.16.0 release run that was a server.mjs from a worktree that no longer existed. It now also requires the server script to still exist on disk (recorded as server in control.json, or read from the process's command line) and, when control.json records a version, that it is this CLI's. A stale dashboard it can prove is this project's own is stopped before it is replaced, a live dashboard already serving this project on ports 4242–4251 is reused instead of starting another, and a dashboard it starts records its server path and version. Events to a secondary dashboard now carry that server's dashboard-token-<port> credential.
  • A role could silently lose its shell to a sandbox that failed to start. A Bash result starting with bwrap: is the OS sandbox's own error, not the command's; on the 2.16.0 release run there were 31 of them and a QA role spent about 7 minutes and 27 tool calls without a working shell. Each one now raises a sandbox-fault audit event, and two in a row end the role's process and resume the same session in a new one (a new process builds a new sandbox), with a message telling the role to re-run the command. It is bounded to two restarts per task session; after that the role's coordinator is told its shell is not running.
  • Evidence refusals misnamed two common slips. A typo'd headSha was refused with "the tree moved after those checks ran", sending the role to re-run checks that were fine; org_task_done now asks git whether the sha is a commit at all and, when it is not, says "unknown commit (typo?)". A worktree left as a template — a literal <…> or {{…}}, or a path such as …/monomind/SRC that does not exist — is now refused as a placeholder with a hint to pin the real worktree path, instead of as an unknown worktree. Both were seen on the 2.16.0 release run.
  • org_complete read as a session error, and the last turn's usage went unrecorded. Ending the run aborts every live session, and each one announced "session ended with an error" before the daemon's own "stopped with the org" status. A session aborted by the org's stop or completion now reports a session-stopped status, and a genuine crash keeps the error. The turn in flight when a session is cut off never received its result, so its tokens reached the role's budget but no usage event; it now gets one (subtype: "aborted", no cost, since the SDK reports cost only on result).
  • Org cost and tokens were under-counted after a resumed session. The SDK's total_cost_usd and modelUsage are running totals for the CLI process, and the runtime turned them into per-turn deltas keyed by session id alone. With run_config.session_scope: "task" every resume runs in a new process whose total usually starts again from zero, so the first turn after a resume was billed as max(0, small − previous) = 0 — on the 2.16.0 release run nine turns of more than 100k tokens each (14.6M in all) were recorded at about $0, the run reported $63.63 against roughly $72.70 re-priced, and budget_usd caps tripped late because maxUsd enforcement saw the same numbers. Deltas are now taken per process: the first result of a new process counts in full when its total is lower than what the previous process last reported for that session, and only the increase when Claude Code carried the old total over.
  • A sandboxed org role could lose its Bash tool to bwrap: Can't create file …: Read-only file system. (Fixes #323) The Claude Agent SDK's sandbox binds /dev/null over "dangerous files" that don't exist yet — .gitconfig, .bashrc, .mcp.json in the role's cwd, ide, local and friends in ~/.claude — and bubblewrap has to create a 0-byte mount-point file for each. monomind made the parent directories read-only (policy.sandbox.denyWrite: ["."] resolves to the org root, which is or holds every role's cwd, and ~/.claude is always denied), so every Bash call failed unless another sandboxed process happened to hold the stubs. On the 2.16.0 release run that was 31 failed calls, ~7 minutes of a QA role's time and three mandatory checks skipped until the next QA round. Such a directory now goes to the SDK as its existing children, each still read-only, and the directory itself becomes a writable bind mount (so it can't be renamed away). The OS layer now allows only one new thing: creating new entries directly inside it. The Claude Code and git names that matter there are the SDK's own denies, and the file tools still refuse writes anywhere in a denyWrite directory. Linux only; macOS's seatbelt needs no mount points and is unchanged.
  • monomind search X and search X --type code missed code symbols that monograph search -q X found. (Fixes #322) Three things stacked up. The code capability only switched on when the saved capability fingerprint counted code files, and that fingerprint is trusted for 24 hours, so one taken before the code existed (init writes one up front) left code search off even after monograph build. It now also switches on whenever .monomind/monograph.db exists. Code hits were scored with the raw FTS5 bm25 rank, about 1e-6 on a small index, so they sorted below every document, media or data hit (scored 0.5–1) and fell out of the top --limit. They are now scored by rank position on the same scale as documents. And --type filtered the merged list after the limit, so any other type that filled the top 20 left --type code with "No results found."; the filter now picks which capabilities are searched before the limit is applied.
  • monograph wiki never created Section nodes for markdown headings. (Fixes #321) The command promised "headings → Section nodes", but the phase that splits documents into sections — and the PDF, co-occurrence and --llm phases that read those sections — were never registered with the build pipeline, so each .md file produced one Document node, no heading was searchable on its own, and --llm had nothing to enrich. A build now creates one Section per heading (linked to its parent heading and its file) plus one for a heading-less .txt/.rst/.md file, and monograph search -q "<subheading>" finds it. A rebuild replaces a file's sections instead of leaving the old ones behind when a heading moves. @monoes/monograph 1.6.9.
  • doc/agent-exec-protocol.md said "rev 8" but its revision history stopped at rev 7. (Fixes #324) The rev 8 change was real: on 2026-09-17 the §2 handshake example gained the org-tool-providers, org-endpoint-roles, org-federation and org-decision-attribution capabilities, but no history entry was written for it. The history now has one.
  • A tool provider's tool never received arguments its inputSchema didn't list. (Fixes #325) The runtime built its argument validator from properties alone, and every runner's zod object stripped the other keys, so a tool advertising {"type":"object","additionalProperties":true} got {} from tools/call whatever the role passed. mono-agent's granted-automation tools hit this in a live org run. When the top-level schema allows additional properties, which JSON Schema does when the keyword is absent, unlisted keys are now kept and passed through on every runtime: the Claude SDK's MCP server, the fence runners and the Vercel runner. Listed properties are still type-checked, a schema-valued additionalProperties checks the extra keys, and additionalProperties: false still strips them. The fence protocol lists such a tool's parameters with ...other keys.

npm: @monoes/monograph 1.6.9, @monoes/monomindcli 2.16.1, monomind 2.16.1