Clio Coder 0.5.0
public launch and documentation
- Prepare the v0.5.0 public launch with a shorter product README, current contributor and security guidance, a roadmap that separates plans from shipped behavior, and consistent package/repository metadata. Earlier changelog entries record development history; their settings and commands are not a substitute for the current reference.
- Update and review the shipped Markdown against the implementation, with browser documentation rendered from the same versioned files. Keep the terminal as the primary coding interface and identify the wider graphical app as an early preview.
- Prefer current guides over historical proposals in ordinary
clio_docssearches, retain explicit historical lookup, align citation anchors with the browser renderer, and point follow-up queries to the current gateway capability.
engine, prompts, and long sessions
- Upgrade Pi agent, AI, and TUI packages to 0.86.1. Preserve the native system-message transcript through replay, continuation, compaction, prewarming, and worker requests; adapt custom transports at the engine boundary. Keep only the documented TUI compatibility patch.
- Compose prompts from a stable identity/constitutional prefix, conditional capability and role guidance, and changing context/scope layers. Shared typed turn constraints inform both admission and prompt guidance; ordinary English is not treated as an authorization parser. Record prompt layout version 3 while retaining older manifest readability.
- Count ready skills through an invalidation-aware snapshot and suppress empty skill reminders, including their unnecessary marketplace lookup.
- Let empty error/abort responses reach lifecycle observers. Run detached task-memory work at settled native tool-batch boundaries, deliver completed reminders as context updates, and retain context-budget and compaction checks. Abort cancels pending work and suppresses late reminders without discarding completed memory or late usage accounting.
- Keep internal dispatch-plan fields out of public memory descriptions while retaining exact call fingerprints and result provenance.
- Measure generation throughput across assistant generation spans instead of including tool, approval, and worker waits. Show tool-call preparation and retain per-call first-token timing.
- Clarify fork identities and tool-only message previews, show operator text in the session tree, and describe cache-affecting events separately from actual backend cache reuse.
- Recover a narrowly identified pre-output LiteLLM connection failure through the visible same-route retry path, without hiding SDK retries, changing routes, or retrying a partial answer.
command admission and context controls
- Preserve upstream command failure in shell pipelines with Bash
pipefail; explicitset +o pipefailretains intentional last-stage semantics. Track working-directory changes through command chains, and preserve repository-script safety checks through redirected compound commands. - Scope worker approval reuse to the same run, exact arguments, approval axis, action class, and policy identity; preserve timeout-denial attribution and explain approval reuse accurately. Deny skips the displayed invocation; Stop ends the turn.
- Wire the existing context-init options into slash parsing and completion, including preview and offline heuristic generation. Route init, refresh, and reset output through the TUI instead of raw terminal output; preview preserves initialization hints and writes no files.
- Fit context legend values to the actual overlay width and reflow them on terminal resize.
graphical application
- Let
clio-coder docs [topic]open a reusable local documentation server and return control to the terminal. Add--foregroundfor a private terminal-owned server and--stopfor the temporary docs server; an idle server exits after the last browser page closes. An installed background application remains a separate lifecycle. - Render long documentation with native expandable sections, heading links that reveal their destination, and an expand/collapse-all control. Preserve readable table columns and keyboard access to horizontal scrolling.
- Polish documentation navigation, search, page outlines, and Markdown links. Keep the expanded sidebar and add an icon-only rail with a collapse button, a keyboard shortcut, and a saved browser preference. Replace heavy native scrollbars with slim themed scrollbars.
- Bundle the local graphical application.
clio-coder guistarts a server bound to 127.0.0.1 and prints a private launch link;--openopens it in your browser. It reads the same configuration, sessions, traces, evidence and documentation as the CLI and the terminal interface, andclio-coder docs [topic]opens the shipped documentation in it. This is a first version: conversation, sessions, traces, fleet and dispatch history, evidence, evals, usage, the library, settings and targets, toolchain, system health and other coding agents. - Every inspector states whose data it shows and what it may do to it, distinguishes a store that was never created from one that is empty, says what a bounded list dropped, and closes by naming what stays on the machine. Opening a page never runs another coding agent's executable; a version probe is an explicit button.
clio-coder uninstallreports graphical background or desktop files it cannot verify, instead of failing, when run from a checkout whose application bundle has not been built.
library
-
Distinguish Clio-local/plugin skills from shared, other-agent, and explicit-path discoveries in model-facing inventory. Include source, scope, and discovered file paths; stop reporting the combined available count as Clio installations.
-
Fix
/skillinstallation through the active profile and add project-scope installation. Keep bundled packages available without automatic installation. Report damaged copies accurately across all package kinds and provide explicit repair guidance with recovery backups. Allow operator-approved canonical library commands through bash while preserving direct instruction-file protections and package integrity validation.
configuration and terminal
-
Detect consecutive past-EOF reads of an unchanged file even when offsets differ, then use the existing bounded synthesis lock. Clarify that shadow-agent routing does not authorize operator
/runinvocation and that undocumented integration rationale must not be presented as fact. -
Give interactive loop lockout an explicit final-synthesis instruction and one bounded recovery for a markup-only answer. Keep tools disabled, deliver the recovery exchange atomically, preserve usable prose, and respect cancellation. Repeated model calls can still exhaust recovery; no result is fabricated.
-
Give the composer bold, state-aware rails: teal at rest, a bounded teal/orange activity highlight, and an orange permission pulse. Mark full-auto explicitly with editor-only red text/endcaps. Hold motion while a draft is present and honor reduced-motion/screen-reader settings; reuse the existing presentation ticker.
-
Correct context reconciliation so growing conversation usage does not inflate fixed system-prompt or tool-definition estimates. Refresh per-call decomposition, separate tool results from definitions and messages, and label provider totals versus estimated category splits in the shared context views.
-
Answer live model/fleet configuration questions with a safe routing snapshot, including profile bindings and target runtime IDs. Rank partial settings-search matches instead of requiring every query word in one control; preserve secret omission.
-
Simplify normal editor rails; move target/model/thinking to the footer and rotate resolved composer shortcuts below it. Keep the compact footer at two lines, preserve operational-notice priority, and retain permission controls and turn-preparation states.
-
Render in-session
/doctoras a width-aware diagnostic report, with errors and warnings first, separate check labels and complete wrapped details. Preserve CLI output and diagnostic behavior. -
Enable interactive Demo guidance by default, with
--demo/--no-demosession overrides and a sharedinterface.demosetting. Add task-grounded capability suggestions, a bounded investigation reminder, follow-through on accepted offers, and expiring contextual footer tips; preserve the two-line footer, user keybindings, permissions, and quiet headless/worker prompts. Refine onboarding to answer greetings, request-writing, and self-contained examples without unnecessary skill discovery; suppress stale agent tips and follow live shortcut changes. -
Add live local-machine telemetry to Status: CPU/RAM meters, process RSS, Linux network and whole-disk rates, supported GPU sysfs counters, ready MCP connections, and active plugins. Distinguish WSL/local metrics, Clio cost ceilings, and unavailable provider quotas; sampling does not launch MCP servers or contact inference endpoints.
-
Redesign the compact footer and paged dashboard: live agent activity, a shared
/contextmulticolor occupancy grid, and detailed session/project/harness status. Keep measured usage separate from configured limits and retain responsive width and height bounds. -
Default LiteLLM onboarding to credential entry and return authentication failures to the credential step instead of presenting an empty catalog and requesting a manual model ID.
-
Reorganize
configureand/settingsaround shared Connections, Chat, Fleet, Context & Memory, Permissions & Limits, Appearance, Integrations, and Advanced sections. Guided connection setup skips unnecessary key/model questions, previews effective defaults, and preserves existing preferences when reconnecting. The complete control catalog supports search, custom numeric values, and validated collection editing. Version-2 settings remain compatible; no reset or new migration is required. -
Add read-only
context(scope="settings")awareness of effective session settings and their exact configuration destinations. Credentials, endpoint URLs, external-agent commands, and arbitrary configuration contents are omitted. Configured ceilings are distinguished from remaining budgets; inspection grants no authority to change permissions. -
Route internal diagnostics through the active terminal owner and notice area, retaining stderr for headless use. Deduplicate repeated runtime notices and keep routine provider, middleware, and listener diagnostics from writing over the live TUI.
-
Honor
--no-skillsin automatic skill/marketplace prompt guidance and first-turn reminders as well as discovery; explicitly supplied--skillpaths remain supported.
agent behavior
-
Present shadow/internal helpers as inline agent-to-agent cards, with compact footer counts and lifecycle notices instead of floating fleet cards. Respect compact, standard, and detailed output styles, preserve full inspection and replay, and label single-task dispatches with the agent name.
-
Include the current local OS account and hostname in the main session’s workspace prompt, reusing receipt identity detection. Treat them as execution facts rather than personal identity; keep them out of persistent project configuration.
-
Recognize “Find where” and “Look at” as evidence of an investigation when checking a read-only Scout assignment, so configuration topic words do not alone reject reconnaissance. Tool authority and declared-write checks remain unchanged.
-
Preserve required terminal formats through loop lockout and markup recovery. Workers receive their declared result shape immediately; a JSON report is not redirected into prose. Keep bounded recovery and result validation, prevent contract revisions from reopening tools after a repeated-call or hard-cap lock, and guide focused source reads instead of repeated small-page exploration. For workers, detect consecutive reads that only revisit already-delivered lines of unchanged files, while allowing new ranges and separate numbered citation reads. Simplify Scout’s report instructions and explicitly end tool-use guidance during locked synthesis.
-
Give eligible native shadow/internal helpers a terminal
clio_submit_resulthandoff using their existing result contract. Validate before acceptance, preserve work-tool lockout during bounded repairs, reject mixed handoff/work batches, and seal the typed payload alongside legacy JSON text. Dispatch and background collection present compact helper results without repeating receipt bookkeeping; conformance never upgrades unmeasured evidence. Artifact-producing helpers retain their existing workflow. -
Explain built-in tools, gateway discovery, skills, the library, shadow helpers, and fleets together in the session prompt. Encourage independent helpers to run with
detach:true, and show helper start/completion/failure in the TUI notice area, including internal runs without transcript islands.
run
- Seal unresolved headless blocks as failure: receipts retain the safety decisions, blocked attempts, and
noopflag. A block with no same-class recovery and no successful write exits 1 with outcomefailedand detailnoop, even without--fail-on-noop. A successfullimitationcall also fails with detaillimitation.--fail-on-noopadditionally rejects runs whose tools all failed without a block. Ordinary prose-only answers and recovered tool failures can still succeed. Report artifacts and repeated denials do not mask an unresolved no-op (#378). - Add
run --timeout <seconds>, which ends the run through the same coordinated shutdown as SIGTERM, seals atimed_outreceipt with statusfailed, and exits 124. A usage error exits 2 even when a timeout is set. - Add
run --cwd <dir>with the semantics ofacp --cwd, so an orchestrator can run Clio in a worktree without changing its own working directory.
safety
-
Repository test commands now run without confirmation at
auto-edit:npm test,pytest,python -m pytest,python -m unittest,cargo test,go test,ctest,make test,make check,ninja test,meson test,mvn test, andgradle test, with bare-word arguments only. They execute repository code.npm run lint,build,typecheck, andcistill ask, and the validation evidence vocabulary now matches this list (#377). -
Safety admission now follows symlinks at every path component, including dangling links and chains, so a write through a link that points outside the workspace asks for confirmation instead of publishing outside it. A path that cannot be resolved (a link loop or more than 40 links) counts as outside, and a call that needs confirmation with no permission listener registered is refused instead of waiting forever.
-
A
..after a symlink now resolves from the link's target, as the kernel resolves it, in safety admission and in the read, write, and edit tools, soecho x > data/link/../fand a write ofdata/link/../fare judged and published where the file actually lands;cdtargets must stay inside under both the logical and the physical reading. -
Bash admission now resolves each write target, cd, and new link from every directory an earlier
cdin the same command can leave the shell in, keeping the call's cwd. A relative write after a cd whose target is expanded at run time, or repeated by a loop or function, asks as an unknown base. A symbolic link to outside the workspace, or a hard link to a file there, created bylnorcp, asks for confirmation. Commands inside( )subshells, brace groups,ifand loop bodies, and$( ),<( ),>( )substitutions are now classified at all. -
Bash admission now asks before a write to a target the shell expands at run time (
> $HOME/.bashrc,> ${OUT}/x,tee {a,b}), a link to outside made bylink, bymvor a link-keepingcpof an existing outside link, or bylnbehindnice,timeout,xargs,find -exec, and similar wrappers, and acda here-document body hid from the scanner. The bash tool no longer passes an inheritedCDPATHto its child shell. -
read,ls,grep, andfindnow follow the autonomy level when the path they name resolves outside the workspace, through an absolute path, a.., or a link: the call asks atsuggestandauto-edit, runs atfull-auto, and is denied atread-onlyand by a headless run belowfull-auto. Before, only the zero-access list stood between these tools and any path on the machine, at every level. Zero-access paths stay refused everywhere, and installed skill, plugin, and extension trees, offload scratch files, and dispatch receipts stay readable without an ask. -
Safety admission now judges the path a file tool opens rather than the one the call spells. The tools drop a leading
@and fold Unicode spaces, sowrite @../xwas judged as a directory named@..inside the workspace and published outside it without asking; acwdargument the file tools never read could pull a../..path, a zero-access one included, back inside the workspace on paper; and a read could fall back to a curly-apostrophe, NFD, or macOS AM/PM spelling that named a different link than the one judged. All three are closed, for bashcwdas well. The Claude SDK and ACP seams now judge the directory a search names and a glob that leaves it.
grep
grepnow shows a match on a line that is not valid UTF-8. ripgrep reports such a line as base64 bytes, and the match used to be dropped without a notice; it now renders with U+FFFD in place of each invalid sequence, as the fallback already did.
eval
- Score a
clio-coder runtask whose receipt sealsnoop: trueas failed with failure classnoop, even when its verifier passes on the untouched workspace, and report the blocked tool calls as the reason ineval runoutput, the JUnit report, andartifacts.failureReason(#378). - Admit
custom.*keys on theclio-coder.eval.measure.v1grader channel as finite numbers or booleans, andcustom.digest.*keys as lowercase SHA-256 hex strings, stored per trial and usable by assertions and suite thresholds. A key the artifact redactor would rewrite fails the item instead of storing[redacted]. - Add the tool-bench harness under
evals/tool-bench/with seeded search and holdout suites for theedit,read, andwritetools that report latency, filesystem call counts, memory, CPU, and a behavior digest per scenario. - Add
grepsuites covering search behavior, holdouts, and full profiles. - Add
findsuites to the tool-bench harness (find.yaml,find.holdout.yaml, and the full profiles): 14 default scenarios covering name and extension globs, deep nesting, hidden and ignored paths, symlink loops, an outside link, and the error paths, with order-normalized digests.
providers
- Make ordinary
targets --probea metadata-only check.--probe --reasoningexplicitly requests reasoning generation, and--probe --toolsexplicitly checks streamed tool calls. Capability declarations and cached metadata are not evidence that a model passed those checks. - Harden Ollama’s native HTTP chat stream: propagate cancellation through residency checks and the request, preserve structured server errors, and reject malformed or incomplete NDJSON streams. Retain native
num_ctx, thinking, sampling, and keep-alive controls. Only managed loads Clio can own and release are pinned; pre-existing models, user-managed targets, and uncertain ownership keep the server’s normal lifetime policy. - Authenticate llama.cpp router residency requests with the same target credentials as inference and honor cancellation during load and polling. Preserve bounded recovery of displaced models when a replacement fails.
- Keep LM Studio chat on the shared OpenAI-compatible transport with REST management. Respect user-managed lifecycle even when load options are present; allow inference to proceed through unavailable metadata when explicit load control is not required. Evict Clio-owned neighbors only for recognized capacity failures, not configuration, authentication, or arbitrary HTTP errors; attempt restoration if replacement fails.
- Treat LiteLLM’s authenticated
/v1/modelslisting as the selectable alias catalog. Enrich matching aliases through/v1/model/infoor/model/info; restricted detail/liveness endpoints do not invalidate a successful inference listing. Missing capability metadata remains unknown, and an explicitly empty listing is not expanded using admin metadata. - Keep context declarations and one-run planning ceilings within observed server limits. Ollama slot discovery conservatively reports one because the client’s environment does not describe the daemon; operators can set
maxConcurrentRequestsexplicitly. Consolidate local sampling translation and preserve runtime-specific thinking controls without changing the target schema. - Add
targets --probe --tools, a live check that the chat or default model streams a schema-valid tool call through the engine path a turn uses. The result shows in the targets table and--json, and a failure warns at runtime resolution. A model the probe loaded on a local server is released before the command returns, and a model that was already resident is left alone. The tool probe has its own 120-second generation timeout, set with--tools-timeout <seconds>, so a cold local model does not fail as a timeout. Plain--probeis unchanged. - The tool probe now releases only the model it exercised on the server it probed, so a model a concurrent chat turn loads in the same process while the probe runs stays pinned for that turn.
- Flag llama.cpp idle-slot eviction in
targets --probeanddoctorwhen a router runs with--kv-unified, more than one slot, the host-RAM prompt cache, and idle-slot caching together, and name--no-cache-idle-slotsas the fix. - Read the context window a resident Ollama model is actually served at from
/api/ps, so an Ollama target is planned and compacted against the serving window instead of the assumed runtime default. Ollama commonly serves a model far below its own maximum, and the smaller number is the one a run has to respect. - Read an Ollama model's maximum context window from
/api/show(and from/api/tagswhere the server reports it there), so a model that is not resident is planned against the smaller of that maximum and 131072 tokens (Ollama opens a cold model at a window no API reports before load), andclio-coder targets --probeshows the serving window beside the maximum. - Report Ollama's HTTP 400
exceed_context_size_errorwith the server's own text instead of[object Object], so a prompt past the serving window is recognised as a context overflow and the session compacts and retries instead of failing the turn. - Add an opt-in
targets[].ollama.numCtxsetting that sendsoptions.num_ctxon every Ollama chat request and becomes the window Clio plans against. It is unset by default because a changednum_ctxmakes Ollama reload the model. - Restore the
degradedruntime notice on Ollama, llama.cpp, and LM Studio targets: a turn generating under 2 tokens per second after 30 seconds warns once with its rate and the models resident on the target (#381). - Release the Ollama models a Clio process loaded when it exits, including models its dispatched workers loaded and models an
acpsession loaded, so a finished run or session no longer leaves them pinned withkeep_alive: -1. Models that were already resident stay loaded, and a worker-loaded model stays warm for later dispatches until the orchestrator exits (#379). - Rename the Ollama runtime id to
ollama, soconfigure --runtime ollamaworks like every other local runtime.ollama-nativestays accepted with one deprecation warning until v0.7.0, asettings.yamlthat names it keeps loading, andclio-coder upgraderewrites it toollama(#376). - Name the closest registered runtime when
configure --runtimeortargets convert --runtimeis given an unknown id, as inunknown runtime id: olama (did you mean 'ollama'?)(#376). - Report a memory step served by the chat target as a
route-fallbackruntime notice instead of reusing thedegradedkind (#381).
fleet
fleet.concurrencynow defaults toautoin new settings and in asettings.yamlwithout the key. An explicit number orautoalready insettings.yamlis kept as written, and no migration rewrites it.fleet.concurrency: autosizes the local node from usable CPUs, available memory, and any cgroup memory limit, up to 8 workers, instead of a fixed 4. SSH nodes and the global pool are not clamped by the orchestrator host, and the binding limit shows in/settings, the footer worker chip, andclio-coder configure.- A worker's provider context-overflow error no longer trips the target's cooldown or retries onto a route of the same size, since the same prompt overflows on every route with that window.
- A bare HTTP 500, "internal server error",
ECONNREFUSED,ECONNRESET, or "fetch failed" from a worker is classifiedtarget-transientand retried on another target instead of being charged to the worker runtime. - The per-route target cooldown is now a half-open circuit breaker:
fleet.retry.breakerThreshold(default 1) sets how many consecutive target failures open a route, one probe run tests it when the cooldown expires while other dispatches still see it cooling, and each failed probe doubles the cooldown up to five minutes. - A retry that leaves a failed target now moves to the first configured route on another target that carries the request's required capabilities and whose breaker is closed, instead of the first other target in settings.
- A retry that moved to another target keeps the assignment's failover mode, so a later retry in the same chain can still leave a failing target instead of being pinned to it.
- A worker's context overflow is retried once onto an eligible route whose context window is strictly larger, including another model on the same target, and the assignment lineage records
context-overflow: <from> -> <to> on <target>/<model>. Without such a route, undernonefailover, or after an overflow retry, the assignment fails with a detail naming why. - The
/settingstargets rows show each route's dispatch breaker in the running session: an open route replaces the health cell with its remaining cooldown, a probe in flight showsprobing, and the detail line names each route's state and last failure.clio-coder targetsanddoctorshow none, since they never dispatch. - Recover
worktree: truetask worktrees after a crash. Each claim now carries a lease on the Clio process that created it; at the next start, a dead owner's worktree that holds no work is removed with its branch, and one with commits or uncommitted files is kept, marked abandoned, and named once on stderr. Nothing is merged or deleted on recovery, and a live owner is never touched.doctorlists every task worktree that outlived its run with its branch, age, and the git commands to inspect or drop it. - Slurm support is the clio-kit Slurm MCP server reached through the existing local stdio MCP gateway: a new guide (
docs/guide/slurm.md) gives themcp.yamlentry, the fivemcp_slurm__*capabilities, and how autonomy treats them (every call asks, because a user-scope MCP server is classunknown), and the newslurm-jobslibrary skill teaches the submit, poll, describe, fetch-output loop with the job id stamped into the final answer. - Add
fleet.worktrees.root(diskby default,tmpfs,auto, or an absolute path) for the working tree ofworktree: truetasks on local placement. Off-disk roots check free space first and fall back to disk with a notice, a fleet with nodes keeps worktrees under the project root, the ownership claim always stays in the project, restart recovery covers the configured root and prunes a tmpfs worktree lost at reboot, anddoctorreports the resolved root with its filesystem and free space. Compete candidates are unchanged.
doctor
doctorreports HPC toolchains:cc/gcc,c++/g++,clang,gfortran,mpicc,mpicxx,mpirun,nvcc,cmake,make,ninja,meson,python3, andsbatch, each with its resolved path and version line from a bounded, parallel--version, in text and--json. An absent tool is anINFOrow; it warns only when the workspace validation contract names it orruntime.kind: slurmneedssbatch. An installed tool whose--versionfails, such as an unconfigured Slurm client, warns.- Add
doctor --deep, which runs the normal checks plus the live tool-call probe on every configured target (bounded by--tools-timeout <seconds>, default 120) and a dry run of the workspace validation contract that resolves each validator command on PATH and reports whether the policy engine would run it without an approval ask at the configured autonomy. The dry run executes nothing, and--deepcomposes with--json. - Add
/doctorto the TUI, which renders the same findings in-session as one notice at the level of its worst row./doctor deepruns the deep checks against the session's targets and autonomy. - Report Slurm through the clio-kit MCP server: whether
clio-kitis on PATH and ships the Slurm server, whichmcp.yamldeclares it and its trust, and whethersbatchandsqueueexist. Informational only.