Skip to content

Releases: viktordanov/uah

uah v1.8.4

Choose a tag to compare

@viktordanov viktordanov released this 04 Oct 01:21
1eaf7db

uah v1.8.4

/goal: keep working until it's done

  • What it does: /goal <objective> sets a goal. Each time the session goes idle, uah starts another turn until the agent marks the goal complete after Codex's requirement-by-requirement self-audit.
  • Commands: /goal, /goal edit, pause, resume, clear and status, all usable while a turn runs. The footer shows the progress, for example "Pursuing goal (12.5K / 50K)", and automatic turns are marked in the transcript.
  • Limits:
    • from Codex: a token budget ([goals] max_goal_token_budget), and a goal is marked blocked after three turns in a row that make no progress or whose commands all fail, or after a run error;
    • added in uah: at most 50 automatic turns in a row ([goals] max_continuations; /goal resume allows 50 more), and esc always pauses.
  • Prompt cache: goal messages are appended to the conversation; the system prompt doesn't change.
  • Persistence: the goal survives resume, compaction and restarts. Subagents don't inherit it.

Auto-review works like Codex's

  • One review conversation per session. The reviewer keeps a single conversation for the session and sends only what happened since its last review. On long sessions that cut its uncached input by 34–38%.
  • Parallel approvals still review at the same time: the extra reviews run on throwaway copies.
  • The reviewer can run read-only commands before deciding (Codex's exec_command), in the read-only sandbox with no network.
  • Unchanged: it fails closed, and the circuit breaker is the same.
  • An MCP server can't speak for you. Resource text attached with @server:uri is hidden from the reviewer, so a server can't plant text that reads like your approval.

MCP

  • Resources. The agent gets Codex's list_mcp_resources, list_mcp_resource_templates and read_mcp_resource tools. You can attach a resource with @server:uri in the composer, and it shows in the @ menu.
  • Prompts as commands. Run an MCP prompt as /mcp__<server>__<prompt> args; its result becomes your message.
  • Restarts. A server that stops is restarted after 1, 2, 4, 8 and 16 s, at most 5 times in a row. Its tools keep their names in the meantime, so the prompt cache holds. /mcp shows the restart.
  • tools/list_changed. A changed tool list applies from the next run, never in the middle of one.
  • Logins. A server waiting for a login reconnects at the next run after uah mcp login in another terminal, with no /new.
  • Unsupported auth. uah mcp list, get and login no longer reject the whole configuration when one server has an unsupported auth mode.

Security and reliability fixes

A full code review found these. Each fix was then reviewed again with uah review until it came back clean.

  • Sandbox launchers. They now live in a folder that sandboxed commands can't change, in every mode, and an existing launcher is used only after it's verified, not just because it exists. Before, uah exec --ephemeral kept them in a folder sandboxed commands could write.
  • apply_patch and symlinks. Patches now write through directory handles, one path part at a time, and any symlink fails the patch. Before, a folder swapped for a symlink between the check and the write could send a patch outside the allowed folders. On macOS, protected names such as .git are now matched without regard to case, so .GIT/hooks is protected too.
  • Forbid rules for apply_patch. They are now checked first for every patch, yolo included.
  • MCP redirects. Configured headers and OAuth tokens now go only to the server's own address, and redirects to another address are refused. Before, a redirect to another site sent your MCP auth headers there.
  • Subagent settings. Subagents get the parent's current settings when they're created, resume_agent keeps a child's own settings, and forks inherit the parent's settings and are told their own $TMPDIR.
  • uah exec.
    • --stdin accepts lines of any length, and an input error exits 1 instead of dropping the rest silently.
    • Any failure now exits 1, and an older answer is never reported as success.
    • The disk-limit (3) and interrupt (130) exit codes are kept.
  • Large queues. Resuming a session with a very long queue no longer hangs.
  • --no-instructions. The prepared context now says that loading instruction files is turned off.
  • Crash cleanup. It kills a leftover process group only after checking it's the same process, by boot and start time, so a process ID reused after a reboot is never hit. This needs uah-core v0.9.0 and uagent v0.8.0, which this release uses.

uah v1.8.3

Choose a tag to compare

@viktordanov viktordanov released this 03 Oct 21:00
1eafbb5

uah v1.8.3

Reviews

  • New uah review: a headless review that mirrors codex review. Pick the target with --uncommitted, --base <branch> or --commit <sha> (with --title), and add custom instructions as an argument or through -.
    • It prints the review as Codex does. --json prints one line with the findings, verdict, model, effort, time and token usage, and -o writes the review to a file.
    • It exits 0 when the reviewer answered, whatever it found.
  • The reviewer is no longer told something untrue. Its system prompt leaves out your AGENTS.md files, but the prepared context said they were there, so it went looking for them. It is now told they were left out on purpose and which ones they are, and it reads them, as the review rubric expects.
  • The REVIEW line shows the reviewer's model, effort and tokens.

Fixes

  • Token counts are back in session lists. uah sessions and uah sessions --json showed 0 tokens for every session, because a library dropped the stats when it read each run's summary. uah now reads them itself. The session index rebuilds once.
  • The status line always shows the elapsed times. On a narrow window, a long file path or command shortens first (to ~, then …/file), then the "esc to interrupt" hint drops. The times are never cut. Paths in the status line are now shown relative to the workspace or ~.

Faster tests and CI

  • Time to green: CI now takes about 2.5 minutes instead of 5–5.5. The slow tests run in their own job alongside the others, and the agentbench dry run only runs when bench tasks change, plus nightly.
  • Local race tests: go test -race ./... takes about 31 s instead of 60 s. Independent tests now run in parallel, the waits are shorter but assert the same things, and duplicated tests are removed.

uah v1.8.2

Choose a tag to compare

@viktordanov viktordanov released this 03 Oct 19:06
1eaf629

uah v1.8.2

The agent can ask you, with options

  • New tool: the agent can end its turn with up to 3 questions, each with a few concrete options. It uses Codex's request_user_input, with the same name and schema. You answer in a picker, and the agent continues with your answers. There is no timeout and no asking mid-turn.
  • Picker keys: ↑↓ or a number chooses an option; n adds a note to the option; the last row, "Type your own answer", is for free text; tab moves to the next question; enter answers; esc interrupts.
  • Previews beside the options. When the options are concrete things to compare, such as code variants, layouts or configs, each option can carry a preview, shown next to the list and updating as you move. On narrow windows it stacks below the list.
  • Where it's offered:
    • only in the interactive TUI, to the main agent;
    • uah exec and subagents don't get it; the agent asks in its final message instead.
  • To turn it off, for example for an embedded or web terminal:
    [tools.experimental_request_user_input]
    enabled = false
    or set UAH_REQUEST_USER_INPUT=off. When it's off, the agent asks in its final message as before.

Thanks to @mwotton for asking for this in #4.

Approvals look the same

  • Approval prompts (commands outside the sandbox, patches, MCP tools) now use the same framed panel as questions.
  • Their keys and behaviour are unchanged.

uah v1.8.1

Choose a tag to compare

@viktordanov viktordanov released this 03 Oct 17:18
1eaf804

uah v1.8.1

Context preparation: documented, and no more pinned copies

  • New guide: docs/context-preparation.md, plus a section near the top of the README. It explains what a new session is told about its environment, and covers:
    • every module key, and how when matches;
    • the evaluation order;
    • the library of optional modules;
    • security;
    • recipes and troubleshooting.
  • uah prompts init no longer copies the context modules into ~/.uah/prompts/context/. There, every copy replaced its built-in for good, so later changes to the text never reached you.
    • It now writes them as a reference to ~/.uah/prompts/context.defaults/, which uah never reads.
    • Its output explains which files take effect and how: prompt files need config lines, while a module override works by its path. It also says how to undo.
  • New uah prompts status: lists your prompt files and module overrides, and flags copies identical to the built-in as pinned.
  • New uah prompts prune [--dry-run]: deletes those identical copies. If you ran uah prompts init before this release, run uah prompts prune once.
  • Your own library-style modules (enabled: false, turned on with [context] modules) work from ~/.uah/prompts/context.d/ without replacing any built-in.

A built-in skill that explains uah

  • uah-customization ships in the binary and is installed to ~/.uah/skills/.system/, as Codex does with its system skills. A skill of yours with the same name wins.
  • Ask the agent how uah works or how to change it: context modules, prompts, hooks, skills, configuration layers.
  • Tests keep the skill and the guide in step with the real module schema, so their explanation can't fall out of date.

uah v1.8.0

Choose a tag to compare

@viktordanov viktordanov released this 03 Oct 15:09
1eafb90

uah v1.8.0

Against v1.7.5, measured side by side in a real environment (fish, macOS, auto mode, high effort, adaptive 2-steps; 8 tasks × 4 runs, all passed):

  • −32% cost (API prices)
  • −18% wall time (median per run)
  • −20% requests
  • −47% uncached input
  • −25% output

Against Codex 0.159.3, on the 5 tasks both can run: −48% wall time and −55% cost. uah passed 20/20 and Codex 19/20.

Context preparation

Every new session, and every subagent, now starts with one developer message that describes where it runs, so the model stops tripping over its environment:

  • The shell and its traps. For example, in fish heredocs, for … do and x=1 fail; zsh, nu, PowerShell and others are covered too.
  • OS tool differences. For example, macOS has BSD sed -i '', stat -f and date -v.
  • The sandbox: what's writable, a private writable $TMPDIR per session in every sandbox mode (read-only included), and what needs escalation.
  • The git state.
  • The instruction files, with "these are all of them". The model no longer searches for more AGENTS.md files.
  • Guidance on output size limits.

Measured alone against v1.7.5: −13% wall time and −39% failed commands. Fish errors, heredoc failures and AGENTS.md hunting dropped to zero.

  • @path lines in AGENTS.md (global or project) are expanded in place in the system prompt: relative to the file, with ~, loop- and size-guarded.
  • The text lives in Markdown modules.
    • uah prompts show context/environment/fish prints one; your copy in ~/.uah/prompts/context/ overrides it.
    • Add your own in ~/.uah/prompts/context.d/ or a project's .uah/context.d/.
    • Each module's front matter says when it applies (shell, OS, sandbox, agent, files, or a check command).
    • Checks run without a shell, in the read-only sandbox with no network.
    • Project modules need uah context trust.
    • A library of optional modules ships turned off: go, python-venv, node, rust, docker, git-lfs. Enable them with [context] modules.
  • uah context shows which modules apply here and why; --show prints the exact block.
  • On by default. Turn it off with context_preparation = false, --no-context-preparation or UAH_CONTEXT_PREPARATION=off.

Effort updates: adaptive effort without cache misses

The prompt cache is kept per reasoning effort, so every switch of adaptive effort used to re-bill the conversation uncached. uah now keeps each request's effort fixed and changes the effort with a configuration_update item in the conversation, as Codex does. The cache survives every switch, including /effort and alt+e.

Measured alone: −30% cost, −47% uncached input, −8% wall time.

  • It applies to models whose catalog supports it (gpt-6.1-sol, gpt-6-sol, gpt-6-astra, gpt-6-luna); other models switch effort per request as before.
  • If the backend ever rejects the item, uah retries without it, says so once, and switches per request for the rest of that session.
  • UAH_EFFORT_UPDATES=off turns it off.

Also new

  • Verbosity, as Codex sends it. uah now sends text.verbosity from the model catalog (low for gpt-6.1-sol). Answers are about 28% shorter. Override with model_verbosity, --model-verbosity or UAH_MODEL_VERBOSITY.
  • Prompt cache misses by cause. /usage, /status and uah sessions show list the misses: cold start, effort switch, idle gap, compaction, model change. The line reads like "prompt cache 94% · missed 41k: idle 28k, cold start 9k · ≈6% of usage", and --json gives each request.

Breaking

  • Runs on uagent v0.7.0 and uah-core v0.8.0, which add a developer message role. Sessions from earlier versions resume normally.

uah v1.7.5

Choose a tag to compare

@viktordanov viktordanov released this 02 Oct 22:39
1eaf475

uah v1.7.5

Network commands escalate from the first try

  • When the sandbox has no network (read-only, or workspace-write without network_access), the Bash tool's sandbox note now tells the model that a command needing the network, localhost included, fails there, so it asks for escalation from the first try. Before, the model ran such commands in the sandbox first, failed, and only then escalated. Codex's prompt still says to try in the sandbox first.
  • Measured over 264 benchmark runs, all passing: no sandbox network attempts in 56 runs on network tasks (87 in the control's 40), one request fewer, and 26% faster on a task whose prompt doesn't mention the sandbox. No escalations on tasks without network, and no slowdown.

Benchmark

  • New agentbench task curl-parallel-nohint: four slow endpoints with no hint about the sandbox in the prompt.

uah v1.7.4

Choose a tag to compare

@viktordanov viktordanov released this 02 Oct 21:07
1eafdc3

uah v1.7.4

Breaking

  • The model variables are now UAH_LLM_* (formerly UNREAL_HARNESS_LLM_*); the old names are no longer read. uah now uses uagent v0.6.0 and uah-core v0.7.0, and uagent keeps its state in ~/.local/state/uagent (an existing unreal-agent state folder is moved there once).

Fixes

  • Queued and steered messages reach the model in the order you typed them; two quick presses could arrive reversed.
  • A forbid rule now also refuses commands uah can't split into plain words (substitutions, redirects, VAR= prefixes, a $CMD name) when they match or may match the rule, in every mode including yolo. Such a command is refused with a request to write it out plainly.
  • Trust for a project hook that runs no script now holds only in the workspace where you granted it.
  • A PreToolUse hook that answers "allow" now approves the call (no auto-review, PermissionRequest hook, or prompt); forbid rules still apply first, and deny and ask still refuse.
  • MCP: closing a session is final (a late startup can no longer restart servers); a login that expires mid-session marks the server as needing uah mcp login <name>; a server with an unsupported auth mode fails on its own instead of failing every server.
  • apply_patch is all or nothing: if a write fails partway, every file is put back as it was.
  • Tests never read your real home (~/.codex/AGENTS.md, skills).

Docs and benchmark

  • New: Adaptive effort costs. Each reasoning effort has its own prompt cache, so adaptive effort keeps two; over 6–7-message chats it still uses 7% (1-step) to 25% (2-steps) less and is 24% to 40% faster, with every run passing. It includes a three-term model that matches the measurements within about 2 points, projections by session length, and how certain each claim is.
  • agentbench measures multi-message sessions per turn (-turns), with five new chat tasks, and its CI dry-run no longer flakes on git lock files or a shared service port.

uah v1.7.3

Choose a tag to compare

@viktordanov viktordanov released this 02 Oct 15:46
1eaf874

uah v1.7.3

Approvals for parallel tool calls run at once

  • When the model sends several tool calls in one response, their checks (PreToolUse hooks, the sandbox auto-review, or your prompt) now run at the same time instead of one after another, on macOS and Linux. With four escalated calls in one response, the wait went from 11.9 s to 3.5 s and the task from 65 s to 49 s.
  • An interrupt now cancels the reviews and prompts in progress at once, instead of waiting for each review.
  • Several approval prompts can be open together; they show one at a time in the order they arrived. "Don't ask again" on one also closes the other open prompts its new rule covers, as approved.
  • The auto-review circuit breaker counts each review as it finishes, so if it trips partway through a response, the response's other reviews still finish.

Benchmark

  • New agentbench task curl-parallel-endpoints: four slow, independent endpoints behind escalation.
  • Task scripts run with git's background maintenance off, which fixes a flaky dry-run on CI.

uah v1.7.2

Choose a tag to compare

@viktordanov viktordanov released this 02 Oct 15:20
1eafff0

uah v1.7.2

Fixes

  • Adaptive effort now shows when a request runs at the lowered effort. While the agent works, a crowded footer cuts the folder path and mode first, so high→medium stays visible. Before, the footer showed plain high on most window widths, so adaptive effort looked like it never switched. The lowered requests themselves were already working.
  • Changing settings during a run (alt+e, /adaptive, /config) no longer stops at the first setting that can't be applied live. Every other changed setting, adaptive effort included, still reaches the running turn.

uah v1.7.1

Choose a tag to compare

@viktordanov viktordanov released this 02 Oct 14:21
1eaf85b

uah v1.7.1

uah runs on its own runtime: uah-core

  • uah's runtime is now uah-core, uah's own runner. It began as a fork of unreal-agent (MIT, © 2026 Unreal Labs, credited in the notices) and carries uah's work: freeform tools, the wake policy, and the request and resume performance fixes. The runner binary is uah-core-runner; uagent v0.5.0 looks for it.
  • The model's preamble no longer names another product.

Fixes

  • A failed save at the end of a run (the session file's final write or a held checkpoint) now fails the run with "failed to save the session: …" instead of passing silently.
  • Paths are shortened to ~ only when they are your home folder or inside it; a sibling folder such as /Users/you-backup is no longer shown as ~-backup.
  • /adaptive completes its values with what each does at your effort ("follow-ups after tool results at low") and marks the current one; /adaptive 2, two, 1, one and 0 work too.

Docs

  • The architecture rules are rewritten for today's code: every package, the import directions, naming, and the documented exceptions.
  • Wrong statements on the first pages are fixed (ctrl+enter sends now and cuts off the answer under way; when a model name is checked; usage), superseded design records are labelled as history, and the docs index starts with a reading path.