Skip to content

cmagent v0.3.16-rc2

Choose a tag to compare

@coremail-cyt coremail-cyt released this 07 Sep 01:42

0.3.16-rc2 -- 2026-09-06

After the first rc2 tag

  • An imported bundle cannot write outside the record it names, on
    Windows either.
    The check that kept a bundle's file paths inside
    their own directory was Unix-shaped: .., an absolute path, a leading
    /. None of those describe C:\escape or the drive-relative
    \escape, and is_absolute() answers false for both on a Linux
    build -- so a bundle crafted for a Windows machine had a way out that
    the tests, all running on Linux, could not see. Every component of
    every path must now be an ordinary name, which is a rule each platform
    decides the same way and each platform's tests can check.

  • And it cannot be led out by a link. Nothing in a bundle creates a
    symbolic link, but a directory it writes into may already contain one,
    and the write followed it. Each level from the record's root down is
    checked before any directory is created or any file opened.

  • Two processes cannot run one session. A turn that is mid-flight
    and a turn that was interrupted look identical on disk -- both are a
    session whose last message is still marked running -- so a gateway
    starting while a cmagent tui was an hour into a long task would read
    that mark as an interrupted turn and resume it, leaving two agents
    writing one conversation. The session now carries an OS file lock that
    the scan and the agent both take: a session someone else is running is
    not an interrupted one. A lock rather than a pid file because the
    kernel releases it when the process dies, however it dies.

  • A finished turn forgives its retries immediately. The recovery
    allowance was cleared on the next startup that saw the turn finished;
    it is now cleared when the turn finishes, so a session that needed one
    resume months ago starts its two attempts over rather than carrying
    one spent.

  • install-local.sh handles a path with a space in it. The list of
    desktop copies to replace was a space-separated string, so
    /Applications/Review App/ split into two entries that named nothing,
    and a [ in a directory name was read as a glob. It is an array now.

Moving a configuration between machines

  • Settings can be exported to a file and imported on another
    machine.
    Under Settings -> Import / Export in the desktop app: tick
    the providers, agents, presets, skills, MCP servers and global
    defaults that should travel, and the answer is one JSON file. The
    import side reads the file first and shows what it would do -- what is
    already here, which records would be replaced, and which reference
    something neither the file nor this machine has -- before anything is
    written.

    Three rules the format keeps, because a configuration file that moves
    is also a credential file that moves:

    • Channels do not travel. An account is a login on one platform
      and its token is not another machine's to hold. .env mixes
      provider keys with channel tokens and the gateway's own
      CMAGENT_TOKEN_* logins, so the export does not read that file: it
      asks each exported record which variable it reads, and a name a
      channel or the gateway owns is refused even if a provider file
      claims it.
    • Keys leave only sealed. Carrying them is a separate tick and it
      requires a passphrase; the keys are encrypted with it
      (ChaCha20-Poly1305 over a PBKDF2-SHA256 key) while the rest of the
      file stays readable. Nothing in this path can write a key in the
      clear. A provider that arrives without one is marked "needs a key"
      in the plan, with a box to type it -- a provider imported without
      its key is a provider that cannot answer.
    • Consent does not travel. An imported skill arrives with
      Installed trust and none of the tool or command grants it had on
      the other machine, and an imported MCP server arrives disabled.
      Both are code this machine would otherwise start running because a
      file was opened.

    Every id in a bundle becomes a path (agents/<id>/,
    providers/<id>.toml), and a bundle is a file somebody was handed, so
    ids are checked when the file is READ -- one segment, no separators,
    no .. -- rather than when it is half written.

    Not carried, deliberately: channels, gateway users and tokens,
    workspace registrations and browser paths (absolute paths that mean
    nothing elsewhere), an agent's brain.db, and everything under
    data/. Those are records or this machine's own, not configuration.

A long task survives the gateway stopping

  • A running turn writes itself to disk, and is picked up again on the
    next start.
    The context reached disk only when a turn ENDED, so a
    gateway that stopped mid-turn -- a crash, an update, a reboot -- lost
    everything since the previous turn, including the message that started
    this one: the conversation on disk ended one turn earlier, and an hour
    of tool calls was gone.

    Now a turn is marked and saved before the first model call
    (TurnStatus::InputOnly -- the status that was defined for exactly
    this and that nothing wrote), and checkpointed as it runs, at most
    once every ten seconds, so the most a stop can cost is ten seconds of
    work.

    On the way up, the gateway looks -- across the workspace it was
    started in AND every workspace the user has registered -- for sessions
    whose last turn never finished, and continues them, telling the agent
    what happened so it checks its own work before repeating a step that
    writes. Four limits, because starting to work by itself is a serious
    thing for a program to do: only turns interrupted in the last 24
    hours; at most two attempts per session, counted on disk beside it and
    cleared the moment a turn finishes normally -- otherwise a turn that
    is what crashes the gateway would be resumed forever by a machine that
    restarts on boot; never a channel conversation, whose turn ends by
    SENDING, so that a restart cannot turn into a message arriving on
    somebody's phone about something they moved on from an hour ago; and
    every decision is logged, including the ones where it declines.

A long task is no longer killed for being long

  • A session whose agent is working is never evicted. The gateway
    releases sessions that have been idle for 30 minutes, and "idle" was
    measured by last_activity -- which moves when the USER sends a
    message. An agent an hour into a long refactor has not been sent one,
    so the reaper called it idle and agent_handle.abort() killed it
    mid-tool-call.

    It was also silent. Dropping the session drops the broadcast channel
    the page was listening to, so the browser saw a plain connection
    error, kept agentBusy set, and waited for a done nobody would
    send. From there every keystroke went to /steer and every Stop to
    /cancel -- both of which look the session up -- so the tab answered
    HTTP 404: session not found to everything and could not be used
    again without a reload.

    Now eviction asks what the agent is doing, not when the user last
    typed. A working session is kept for as long as it works, and its
    clock reset, so it also gets a full window to be read once it
    finishes. A session blocked on a permission prompt is a third case: it
    is kept four windows (two hours at the default) and then let go --
    nobody may ever answer, and this sweeper is what used to clean that up
    -- rather than pinned for the life of the process. A session that
    really is doing nothing is still released after 30 minutes, quietly:
    nothing was in flight, the conversation is on disk, and the next
    message resumes it with nobody noticing. The release goes to the log,
    which is where an operator looks; how long a session has been quiet is
    already on its row in the session list. The one ending worth telling a
    person about is a turn of THEIRS that stopped, and the page finds that
    out by asking and says it once.

    An open tab is also no longer mistaken for an idle session. The page's
    event stream registers a placeholder -- a channel with no agent behind
    it -- and the sweeper reaped it, ending the stream; the browser
    reconnected, registering the next placeholder for the next sweep. A
    tab left open churned a connection every half hour, about a session
    that was never running.

  • And when a turn does end without saying so, the page recovers. A
    connection-level SSE error while a turn is running now asks the
    gateway whether that turn still exists; if it does not, the page
    unlocks and says what happened. Input typed at a turn that is gone is
    no longer refused: the session is resumed from disk and the words are
    sent as an ordinary message. Stop on a session the gateway no longer
    holds answers "nothing to stop" rather than 404. And a steer that
    fails gives the typed text back to the composer instead of eating it.

Smaller things

  • The desktop's diagnostics report has a repair button. It was
    something to read: it named what was wrong and told you to go and run
    cmagent doctor --fix in a terminal. collect_report says fixing is
    not something an HTTP request sets in motion, which is the right rule
    for a gateway reached over a network and the wrong reading of the
    desktop app, where the request comes from the person at the machine --
    the same door already lets the settings screens write provider files,
    store API keys and choose the OS sandbox.

    The button promises a number, and the number is the report's own
    auto_fixable: how many findings a pass would settle without asking
    anyone anything. Both it and cmagent doctor --fix press the same
    function, so the two cannot come to mean different things. What it
    does NOT cover stays in the terminal and keeps saying so: repairing a
    stale context window or stripping an invalid shell-allowlist entry is
    decided inside the scan loop, and deciding it a second time in the
    button would be a second copy of the rule; repointing a broken
    endpoint needs someone to choose a provider, which a button cannot
    do for them.

  • The trace follows its own tail. Records land in the trace AS a
    turn runs -- a tool result per wave -- so the one view whose whole
    subject is what the agent is doing needed a manual reload to say what
    the agent is doing. It now polls the newest page while a turn is
    running and appends what is new: someone reading the tail keeps
    reading the tail, someone who has scrolled back is left where they
    are, and a gap too big for one page reloads rather than drawing a
    ledger with a silent hole in it. The polling stops after the turn's
    last records arrive, so an idle session with the tab open asks for
    nothing.

  • Slack's "reply in a thread" is a setting you can see. The adapter
    has read thread_replies from the account file since it was written,
    defaulting to on; nothing offered it, so turning it off meant editing
    the TOML by hand. It is now in the channel form, shown for Slack only,
    with a test that the value still reaches the adapter.

  • A session let go while it was still asking something says so. It
    used to report "released after N minutes idle" -- the wrong story for
    a session that spent those minutes waiting for an answer, and the
    reading that makes an unseen permission prompt look like it did not
    matter.

  • The import/export panel, having now been looked at in a browser:
    agent descriptions are one line each rather than a paragraph apiece,
    every list is in one order rather than the filesystem's, and the
    Export button is pinned to the bottom of the panel instead of sitting
    below a list as long as the machine has records.

  • scripts/dev-shot.sh fails instead of hanging. A CDP call that
    never comes back used to wait forever, and every message the driver
    prints comes after the step that hung -- so a leftover headless shell
    still holding the debugging port produced a silent, output-free hang.
    Calls now have a deadline that names the method and the likely cause,
    and the script refuses to start when something else is on the port.

Fixes

  • The provider catalog knows this summer's models. Grok 4.5 and
    4.6 (xAI's current best; 500k context, $2/$6 below 200k tokens),
    Gemini 3.6, 3.7 and 3.8 Flash and 3.5 Flash-Lite (all 1M in, 64k out,
    at the introductory $0.75/$3.75 the pricing page lists through the end
    of 2026), and Claude Fable 5.1 (released 2026-09-01) with Opus 4.8
    joining the legacy list. The picker's default moves to Grok 4.6 and
    Gemini 3.8 Flash; Anthropic's stays on Sonnet 5. Verified against the
    vendors' model pages on 2026-09-03, and for xAI against the live
    /v1/language-models listing.

  • And this month's OpenAI models. GPT-6 Astra, the new flagship
    (1.05M context, $10/$50), and GPT-5.5 Pro, which the docs list and the
    catalog had missed. GPT-5.6 Sol's price follows OpenAI's cut to
    $4/$20. The default stays Sol: Astra lists at 2.5x its input price, so
    it is offered rather than chosen for you. Note that above 272k input
    tokens OpenAI reprices the whole request; the catalog carries one
    price per direction, so a cost estimate for a very long request reads
    low.

After the second rc2 push

  • The font-size control in the top bar says what it controls. It
    showed only a bare percentage ("100%"), which told you a number had
    been chosen but not what it scaled. An "Aa" now sits beside it.

  • Closing the desktop window no longer quits the app. It minimizes
    to a tray icon instead -- Show/Quit menu, left-click to restore -- so
    a stray click on the X no longer kills a long task running in the
    background. The web UI gets its own Exit button next to Settings for
    when a real quit is wanted, since the window's close button no longer
    means that.

  • A WeChat scan login could leave the account unreachable by its own
    owner.
    The scan flow saves the account directly, skipping the
    record form that would otherwise default dm_policy to "open" -- so
    a fresh login had no policy set at all, which the shared allow-list
    code reads as "allowlist, empty," admitting no one. Fixed to fill in
    "open" only when the key is absent: a later relogin never overwrites
    an owner's own choice, but does repair an account saved before this
    fix existed.

  • Lunkr's contact search stopped losing real matches. Two of its
    three sources could produce a hit that never made it back to the
    caller. The pinyin/prefix fuzzy source never carries a uid, and a hit
    with no uid was rightly dropped downstream -- Lunkr's own DM/group
    gate matches by uid -- so every match found only through it silently
    disappeared; it is now resolved via a follow-up lookup. And searching
    your own email or name found nothing at all, since you are not your
    own "recent contact" and the keyword search can answer with no
    results for a plain prefix match; it is now checked directly against
    the logged-in account's own identity.

  • A channel session says which platform, and which account, even
    after its name has changed.
    A Lunkr/Telegram/WeChat/... session
    used to say only "channel" once the auto-title from the first real
    message overwrote the "channel:<platform>-<id>" name it
    started with -- so on a platform like Lunkr that can run more than
    one login, nothing in the session list said which account a
    conversation came in on. It is checked against the session id now,
    which never changes, and Lunkr's own account label follows along --
    in the web session list, the TUI session picker, and the workspace
    browser.

  • Channel settings no longer offer fields a platform never reads.
    Auditing for the same class of drift last release's Lunkr
    bot_discussions fix found elsewhere too: WeChat had no group chats to
    require a mention in, and a DM policy / allow-list / notify field it
    could never usefully be given an id for; Lunkr's stream_mode
    offered a choice its adapter's API refuses every shape of; and the
    token fields showed on platforms whose login is a QR scan, not a
    typed key. Each now shows only where the adapter actually reads it.
    Along the way: a settings group whose fields are all hidden for the
    current record no longer shows as an empty heading with nothing
    behind it, and its field count reflects only what's visible for this
    platform; and a lone checkbox now sits inside its own label instead
    of on a line by itself.

  • Switching platforms mid-setup in the channel wizard no longer keeps
    the old screen up.
    Lunkr and WeChat share one wizard host in the
    settings dialog; switching the Platform dropdown between them while
    it was open reused the existing instance, so picking WeChat right
    after Lunkr kept showing Lunkr's role/method/email/password screen. A
    platform change now rebuilds the wizard from scratch.

  • Opening a session showed the wrong project's git branch. The
    sidebar's initial git status was sampled from the gateway's default
    workspace rather than the session's own, so every workspace other
    than the default one initially showed that project's branch until
    the next per-turn update corrected it. It is now seeded from the
    session's actual workspace, and an empty/non-repository workspace now
    sends an empty branch too, so switching to one also clears whatever
    the browser was showing before.