Skip to content

v0.6.7

Latest

Choose a tag to compare

@github-actions github-actions released this 07 Oct 14:46

Context compaction

  • New /compact command. It summarizes the older part of the conversation and the model reads that summary instead of the old messages. /compact show prints the saved summary.
  • Compaction also runs on its own when the conversation fills 90% of the room the model window leaves for it, or when the oldest tasks are about to drop out of agent.conversationMaxPairs. It shrinks the history to about 65% of that room. Before, older messages were cut off.
  • The full session transcript stays on disk unchanged. The opening request of the current task and your latest message are always kept word for word.
  • The summary is written by the model the session already uses, with no tools. The TUI shows the status, the elapsed time and which part is being summarized, and Esc cancels it.
  • If summarizing fails, the agent falls back to the old trimming and does not try again on its own until the next turn. If the provider refuses a request as too long for its window, the agent compacts and retries that request once.
  • Settings live under agent.compaction in config.json: auto, triggerRatio, targetRatio, summaryMaxTokens, timeoutMs and maxTotalTimeoutMs. The whole operation has a 10-minute limit by default.
  • Over HTTP: POST /api/sessions/{id}/compact and GET /api/sessions/{id}/compaction.

Windows: local models on NVIDIA GPUs

  • With an NVIDIA driver that supports CUDA 12.4 or newer, the agent now picks the CUDA 12.4 engine build, which ships its own CUDA runtime. Before, it could pick the CUDA 13.3 build, which did not include the runtime, so on a PC without the CUDA Toolkit the model silently ran on the CPU.
  • An existing install with that problem is replaced automatically the next time the local model starts, with no manual steps.
  • RTX 50-series cards get the Vulkan build, since the CUDA 12.4 build has no code for them.
  • If a CUDA download arrives without its runtime, the agent installs the Vulkan build of the same release instead and tries CUDA again only on a newer release.
  • cuda-13.3 is no longer picked automatically. You can still set it with localModels.managed.backendVariant, and it needs the CUDA Toolkit installed.

Approvals

  • When you decline a tool call, the model is told that you declined it, not "approval denied". The same call is not asked again in that turn.
  • When the agent sends several calls that need approval in one step, it now waits for each approval and runs them, instead of keeping only one call.

Memory

  • Memory no longer stores names you never typed. Small local models used to copy an example name into your profile.
  • One-off instructions such as "reply exactly OK" or "don't use tools" are no longer saved as lasting preferences. Short pings skip the memory step. Words like "remember", "always" or "from now on" still save what follows.
  • The same profile fact is not written twice.

Local models and Fusion

  • The managed llama-server now starts with an API key, so a web page open in your browser cannot call it. The key is generated once and kept in <models dir>/llama-server.key, or taken from ATOMIC_AGENT_LLAMA_API_KEY. Scripts or other apps that call the managed server directly must send Authorization: Bearer <key>. A server started by an older version runs without a key until you restart it with atomic-agent models stop and models start.
  • On Macs and other machines with unified memory, the automatic context size leaves the system at least 4 GiB or 25% of RAM, whichever is larger, and caps the KV cache at 1/16 of RAM.
  • A large localModels.completionMaxTokens no longer squeezes the conversation to almost nothing. The reply reserve is capped at half the model window, and the context panel shows the reserve and free space the window allows.
  • When the device probe fails, the managed server starts with a 32,768-token context instead of 16,384, which left no room for the prompt.
  • A local server that is stuck on a request is detected as frozen. A failed llama-server start no longer crashes the agent.
  • Each Fusion worker row in the TUI now shows how much context the worker holds, for example 12.3k ctx.
  • Local Fusion workers get time to read their brief, and their reply limit matches their time limit. A large completionMaxTokens no longer cuts the worker pool down to one worker.
  • models status works on a fresh install, before the server has ever started.

Providers and fallback

  • An account that cannot pay (402, or a 403/429 that says funds, balance or billing) ends the turn with the provider's own message, instead of waiting five minutes on the last fallback.
  • Provider errors now include what the provider said. A wrong Gemini key no longer shows only rejected the request (400).
  • When the main provider refuses its API key and the fallback chain runs out, the turn ends right away with the key's error instead of waiting for an outage to pass.
  • A fallback provider with no API key is skipped and marked no key, skipped in the Fallback pane, instead of adding a certain 401 to every turn.
  • A key with a non-ASCII character, or a missing key, now says what is wrong with it instead of rejected the API key (401).
  • A broken provider entry that is not active, such as one saved without a model, no longer stops the agent from starting.
  • /model switches to a local server such as LM Studio or Ollama without asking for a key.
  • Perplexity chat works. Every turn used to fail with 404.
  • models search --refresh lists Gemini models and works with the Anthropic preset.
  • AI/ML API offers mistral-small-2603 in place of the retired mistral-large-2512.
  • In the LLM tab, d on a cloud provider removes it after you confirm with y. The active provider cannot be removed, and the removed provider also leaves the fallback chain.

Windows

  • On Windows the agent is told that PowerShell cmdlets such as New-Item and Move-Item do not work in its cmd.exe shell, and is shown the mkdir and move forms to use instead.
  • os.fs.write is described as making files only, so models stop creating empty files where they meant to make folders.
  • cmd.exe gets each command line exactly as written.
  • HTTPS requests no longer fail when the certificate revocation server cannot be reached.
  • atomic-agent.exe is signed again. In v0.6.6 it shipped unsigned.

Tasks, Telegram, web search

  • atomic-agent task create --at now waits until that time. Before, the task ran on the next scheduler tick. --at takes Unix milliseconds or an ISO-8601 date or date-time, rejects anything else, and warns when the time is in the past.
  • A scheduled task's run now stops when its agent stops, when you cancel it, or when the process that owns it exits.
  • The Telegram bot keeps reconnecting after a dropped connection.
  • Web search skips Exa when no Exa key is set, instead of calling its keyless tier.

Replies and sessions

  • When a reply has a link on the same site as a link from this turn's tool results but with a different path, the agent is asked once to copy the link exactly or drop it.
  • Reading the same unchanged lines of a file again shows a pointer to the earlier read instead of a second copy.
  • Chat titles no longer include a thinking model's reasoning.
  • A turn left running by a crash or shutdown is marked as cancelled at the next start.
  • TUI: one progress bar per download, and a thread stays whole when you switch away mid-answer.

serve and HTTP API

  • Idle streams on /v1/chat/completions and /api/events get a : keepalive comment every 15 seconds, so a turn waiting on an approval is no longer dropped by the client's timeout.
  • /api/events no longer replays approval requests that were already cancelled.
  • serve notices when a client disconnects and stops that turn.
  • Each tool call's result is sent on the stream when it finishes. /api/capabilities lists only the tools the agent is offered.
  • New routes: GET and POST /api/coding-mode, POST /api/context-preview, and the two compaction routes above. Session rows carry their title.
  • error and provider_waiting frames name the billing cause and the providers that failed before. A stopped turn's loop_failed frame reads as cancelled.
  • serve now writes its structured log.

Config, privacy, security

  • config.json is written readable by its owner only (0600). atomic-agent config set - < config.json reads the whole file from stdin, so an API key never has to go on the command line.
  • skill browse and skill search read ClawHub and the GitHub taps together.
  • A GITHUB_TOKEN or GH_TOKEN is sent only to GitHub's own hosts. Before, any download URL that contained "github.com" could receive it.
  • A subscription provider such as Claude Code no longer receives the agent's API keys in its environment, so the CLI uses your subscription instead of an API key it finds there.
  • Analytics events and crash reports now say whether they came from the TUI, the command line or the desktop app, and the terminal and desktop app share one install id per machine. When you turn analytics off, one last event records that it was turned off. Crash reports name only the kind of host a failed connection went to (localhost, private, known_cloud or other), never the host.
  • ATOMIC_AGENT_ANALYTICS=off turns off analytics and crash reports for that process, whatever the config says.
  • User guides moved to docs/user/ (tui.md, memory.md).

Full Changelog: v0.6.6...v0.6.7