Skip to content

Releases: Vulkgryph/Forge

Forge 0.6.0

Choose a tag to compare

@windingcreek windingcreek released this 30 Sep 00:20

Forge 0.6.0

A minor release rather than a patch, because it changes which headers leave your
machine and adds three configuration keys.

macOS is the supported platform for this release. Windows and Linux install
from source — see Platforms.

Security

A crawl could be aimed at the machine Forge was running on. There was no
restriction on which addresses web_search and web_fetch could reach:
localhost, the RFC1918 ranges, and 169.254.169.254 — where a cloud instance
serves its own credentials — were all reachable.

The delivery route needed no cooperation from the model. robots.txt was
checked for the URL that was requested; the fetcher then followed up to five
redirects on its own, and whatever came back was indexed without the
destination ever being checked. One hostile page in an ordinary crawl could
redirect to 127.0.0.1:2375 and have a local service read back to the model as
a page from the web.

Two exposure windows, checked against the tags rather than estimated:
web_fetch has followed redirects since 0.1.0, so a single fetch could be
aimed at a local address in every release there has been; the crawler arrived in
0.5.0, so the redirect-and-index route affects 0.5.0 through 0.5.2.

Now refused at two layers, and every redirect hop is resolved and checked before
it is followed. This is not proof against DNS rebinding, and the code says so
rather than implying more.

A cross-origin redirect bypassed robots.txt, stay_on_host and per-host
politeness
— the same root cause, fixed with it.

Forge described a page as one a person had opened in a browser when nobody
had.
When a live fetch came back behind a bot check, web_fetch served an
indexed copy with the words "already opened in a browser by the user and handed
over"
— inferred from the URL being in the index and the fetch being refused.
A page crawled weeks ago, on a site that has since added a challenge, satisfies
both. That is a claim about a human action reconstructed from a proxy for one,
and it is the strongest thing this project says about that route. The handover
records its own attribution now.

The crawler can prove who it is

Forge already sent an honest user agent with a contact URL and refused to
pretend to be a browser. But a header is a claim, and anyone can write one.

Requests can now carry a Web Bot Auth signature — RFC 9421 HTTP Message
Signatures — so a site can check the claim instead of weighing it. Ed25519
over the authority, with the public key published at
/.well-known/http-message-signatures-directory on a domain you control.
Cloudflare validates these at its edge as part of Verified Bots, which means a
small crawler that behaves well can be recognised without a contract.

Off unless configured. forge-agent --generate-bot-auth-key makes the key —
no openssl needed, which matters because macOS ships LibreSSL and it cannot
generate this kind of key at all.

The crawler obeys what sites actually said

Five places where "we could not tell" was read as "go ahead": a transport error
on robots.txt returned no restrictions with no retry; a 200 was trusted
without being looked at, so a robots.txt redirecting to a login page permitted
everything; there was no size bound; and the cache dropped the scheme.

And two where the site was simply overruled. User-agent: * / Disallow: /
followed by User-agent: forge-search / Disallow: resolved to the blanket ban
rather than the exception — the standard "everyone out except you" idiom, and
exactly the file a site writes after someone asks for access. And
Crawl-delay: 86400 was clamped to 300 seconds: a site asking for one visit a
day got one every five minutes, 288 times what it asked, from code describing
itself as obeying.

429 and 503 were ignored — counted as "missing" while the crawl carried on
at the same rate. Now honoured with Retry-After.

Privacy

A re-crawl returns If-Modified-Since but not If-None-Match.
Last-Modified describes the content and is identical for every visitor; an
ETag is server-chosen and can be minted per visitor, which is a known tracking
technique. agent.send_etag opts in.

The README now states what a site can learn, measured rather than asserted: the
complete header set is Host, Accept and the user agent. No cookies, no
Referer, no Accept-Language, no session identifier.

Three ways the agent could stop being useful while looking fine

A message typed while it was working could be discarded. Stdin was read with
a future that is not cancellation-safe, racing the agent's own output — so an
event arriving mid-line dropped the read and the bytes already consumed. You
waited for an answer to something it never saw.

A message typed mid-turn could end the work instead of steering it. It
arrived as a plain user message, indistinguishable from starting a new
conversation — and answering someone who has just spoken to you means stopping.

A compaction could end the work it interrupted, because the summary read as
a status report and the natural continuation of one is another.

Also: a local-only setup contacted OpenAI twice at startup about a provider it
was not using, and an approved plan could leave the agent with no legal move.

One fewer dependency

scraper is out — 27 packages, including four of the five MPL-2.0 crates in
the tree. Forge has had its own HTML parser since 0.5.0, so web_fetch and
web_search were reading the same page two different ways. They share one now,
and the crawler gained that parser's skip list: navigation and footers are no
longer indexed as though they were what a page is about.

Editor

forge-ide has a changelog entry for the first time in a release where the
editor did not change — the SSRF above was reachable from its agent panel, and
an editor user reading only that file should not have to find out elsewhere.

Forge 0.5.2

Choose a tag to compare

@windingcreek windingcreek released this 28 Sep 02:21

Forge 0.5.2

Windows and Linux arrive, as source installs. macOS remains the only platform
with a downloadable image — see
Platforms for what has been
verified per component.

Windows and Linux

The editor and the terminal client build and run natively on Windows. The
terminal client could not compile there at all — its platform layer was
Unix-only in its entirety — and now has a console-API implementation covering
raw mode, reading keys, terminal size and resize. install.ps1 (or
install.cmd) puts both in place, with optional shortcuts; WINDOWS.md says
what is verified and what is not. Linux gets install.sh for a fresh Ubuntu
system and LINUX.md.

Failure behaviour, exercised rather than assumed

Eleven provider and process fault cases against two streaming formats, on both
platforms, plus SSH disconnection against a loopback server. Streams now require
a completion signal, incomplete tool arguments are never executed, partial
answers survive an error exactly once, transient server failures take part in
bounded retries, and a dropped SSH connection cannot silently move remote work
onto the local machine — the editor offers Reconnect and keeps the remote
workspace and pending conversation.

That matrix runs under cargo test now, including on Windows CI, rather than
being a script someone had to remember to run.

Three ways the agent could stop being useful while looking fine

A message sent while the agent was talking could be discarded. Stdin was
read with a future that is not cancellation-safe, racing the agent's own
outgoing events. An event arriving mid-line dropped the read and the bytes it
had already consumed
; what was left failed to parse and went to stderr. The
client had sent something the agent never saw, and then waited for an answer to
it — more likely the more the agent had to say.

An approved plan could leave nothing legal to do. Approving without clearing
context cleared the plan-mode flag but left PLAN MODE ACTIVE — you CANNOT modify files standing in the transcript, while exit_plan_mode was no longer
offered. The agent refused to implement the plan it had just written and asked
to be taken out of a mode it was no longer in. Both correct, and no move left.

ChatGPT Codex sign-in failed naming a key you never supplied. Forge
exchanged your OAuth token for an sk- API key — on login, on every refresh,
on every token check — and then preferred it for one request that the backend
rejects. Nothing mints it now, and one left by an older build is dropped when
the token file is read.

Knowing where you stand

For ChatGPT Codex, the only provider that reports it, /usage now shows how
much of your weekly allowance is left and when it resets; the editor says so in
the agent panel. An endpoint that reports no allowance says that, rather than
staying silent — silence reads as good news.

Running out used to arrive as API error: Responses API error (429 Too Many Requests) wrapped around a JSON blob, which reads as something broken. It now
reads: "ChatGPT pro plan: the weekly limit is used up. It resets in 5d 14h.
Nothing is wrong with the connection or the login — this allowance is spent."

When a credential is rejected, the error now says which endpoint it went to
and what kind of credential it was — never the secret itself.


Note: WINDOWS.md in this release's source tarball shows an install path under a specific user's home directory. Use a relative path (.\install.ps1) or %USERPROFILE%. Fixed on main; the .dmg is unaffected.

Forge 0.5.1

Choose a tag to compare

@windingcreek windingcreek released this 26 Sep 00:32

Forge 0.5.1

A patch release. Three bugs, each of which stopped work rather than slowed it.

macOS is the supported platform for this release. Windows and Linux are
next; see Platforms for what has
been verified per component.

ChatGPT Codex authentication

Signing in worked and then requests failed with Incorrect API key provided: sk-svcacct… — naming a credential you never supplied and cannot find in any
config file. You didn't supply it: Forge exchanged your OAuth id_token for an
API key on login, on every token refresh and on every token check, wrote it into
chatgpt_auth.json, and then preferred it over the OAuth token for one request.
The backend that accepts the token rejects that key.

It only broke when the mint succeeded, which is why it came and went rather
than failing consistently, and why signing in again did not help — a fresh login
minted a fresh key.

Nothing mints it now, and a key left behind by an older build is dropped when
the token file is read, so it leaves disk on the next save. A credential nobody
asked for should not be sitting in a file.

An approved plan could leave the agent with no legal move

Approving a plan without clearing context cleared the plan-mode flag but left
the PLAN MODE ACTIVE — you CANNOT modify files directive standing in the
transcript. Those are two halves of one state: the flag decides which tools are
offered, the directive is what the model reads. With the flag down,
exit_plan_mode was no longer offered — so the agent was told it must not write
anything, while the one tool that would have let it say so was gone.

It refused to implement the plan it had just written and asked to be switched
out of a mode it was no longer in. Both correct, and no move left that made
progress. Every exit path goes through one place now.

When a credential is rejected, the error says where it came from

A 401 named the key the provider disliked and nothing about which of the several
paths that can supply it had run. It now reports the endpoint, the model,
whether the account header was sent, and the credential's shape — kind, length
and a six-character prefix, which is no more than the provider already echoes
back. Never the secret. FORGE_AUTH_DEBUG=1 traces the same line on every
request, for catching a credential that works now and stops later.

The model-catalog fetch used to swallow every failure and merely look like an
empty catalog. That silence is why the minted key went unseen; it reports auth
failures now.

Editor

forge-ide --version and --help now exist and are answered before any window
opens. An unrecognised argument is taken for a path, so asking the editor its
version used to launch it on a workspace called --version. It reports the
build commit as well as the number, since every build between two releases
reports the same number.

The "newer build is installed" banner wrapped its own dismiss button onto an
empty row once the message filled the first line.


The terminal client has no changes of its own this release and is
version-matched.

Forge 0.5.0

Choose a tag to compare

@windingcreek windingcreek released this 23 Sep 19:45

Forge 0.5.0

A minor release. web_search stops pretending to be a search engine and
becomes what it actually is — a crawler that keeps what it read — and a second
tool searches documents already on your disk. The rest is what using it turned
up.

macOS is the supported platform for this release. Linux and Windows are
next; see Platforms for what has
been verified per component.

Search

The index is a directory of append-only segments. It was one file rewritten
in full on every save — 556 MB to add sixty pages at the page cap. A save now
writes only what it added: 0.98 MB to add sixty pages, whether the index holds
one thousand or fifty thousand.

web_search says plainly that it cannot find a site for you. Naming sites
narrows the answer to them, and a result that finds nothing reports which sites
have been read, when, and which of your words the index has never seen.
Coverage is weighted by rarity, so a document matching six common words and
missing the one that was the question no longer reads as a near miss.

search_documents indexes prose on disk as the sections it is made of, each
named by the heading trail leading to it and carrying the lines it occupies — so
a result is a place to read rather than a file to search. Measured on 176
markdown files: 540 sections in 408 ms, queries in 24–44 ms.

When a site says no

A bot check is reported as a refusal rather than a missing page, and never
indexed as content. In the editor the agent can offer the page to a person, who
clears the check in a real browser and hands it back; the agent is told at
startup whether its client can do that at all, so a terminal is never asked for
something it cannot give.

Nothing on that route claims to be something it is not. No cookie is replayed,
no identity spoofed, and robots.txt is honoured even when the refusal
arrives wrapped in a bot check
— a 4xx carrying the real rules used to be read
as "this site has no robots.txt".

Dependencies

One added: forge-search, which is ours and has none of its own. One outside
requirement removed: web_search no longer needs DuckDuckGo. The in-window
browser cost no new crates — WebKit is linked as an operating system framework.

Also fixed

A long prompt no longer hangs below the window. An input prompt is withdrawn
when the process proves it was not waiting. A split escape sequence no longer
cancels a running turn over a slow link. Three user-facing messages that
rendered with runs of spaces in them.

Full detail in the changelogs:
agent ·
editor ·
terminal


Forge-IDE-0.5.0.dmg is signed with Vulkgryph LLC's Developer ID and notarized
by Apple, so it opens without the unidentified-developer warning. The terminal
client is not distributed this way — install.sh builds it from a checkout.

Forge 0.4.2

Choose a tag to compare

@vulkgryph-dev vulkgryph-dev released this 14 Sep 17:03

Forge 0.4.2

A patch release. Every entry is something that was wrong, found by using the editor — plus one addition that is a menu on something which had none.

macOS download: Forge-IDE-0.4.2.dmg below is signed with Vulkgryph LLC's Developer ID and notarized by Apple, with the ticket stapled to the disk image and to the app inside it, so it opens without the unidentified-developer warning and offline too.

Forge IDE

Added

  • Right-click an editor tab to copy its path. Copy Path, Copy Relative Path (when the file is inside the workspace) and Copy Containing Folder — the tab bar had no context menu at all before. Dragging a tab into another application, which is what VS Code allows, is not possible here: beginning a system drag needs an NSDraggingSession started from a real NSEvent, and the windowing layer this is built on delivers dropped files but offers no way to originate one. So what the drag is for is offered directly, which is also VS Code's own menu entry.

Fixed

  • A .gz file opens, and a binary one says why it cannot. Opening a .gz failed on invalid UTF-8 and reported only to the status bar, so from the outside the tab simply never appeared. They are now decompressed and shown — gzip is the same DEFLATE payload the PNG decoder in this repository already handles, with a different header, so this cost a header parse and no dependency. The trailer's checksum and length are both verified, because a .gz that decompresses to something subtly wrong is worse than one that refuses. Anything else that is not text now says "… is not a text file" rather than reporting a UTF-8 decoding fault, which named the cause without saying what it meant for the file you tried to open.

  • Multi-cursor edits are undoable. That path wrote the buffer's lines directly instead of going through the editor's write path, so no undo step was recorded and Ctrl+Z skipped straight past the edit to whatever came before it. The same mistake as the original undo bug, where the only caller that took a snapshot was the Tab handler — so the test added for it asserts the property (an edit made through the editor path is undoable) rather than naming a call site, since that is the form the bug keeps taking.

  • A saved conversation no longer carries a whole tool result. The panel's transcript is rewritten in full each time an item is added, and one conversation here had reached 11.4 MB on the strength of a single 7.4 MB tool result — so every new message meant re-serialising and writing eleven megabytes on the event-loop thread. Saved tool output is now capped at 64 KB, keeping both ends and saying plainly that the rest is in the agent's own log, which is untruncated and is what rewind reads. The largest conversation here drops from 11.4 MB to 3.4 MB. The file is also written compactly rather than pretty-printed, which was costing about a third of its size for the benefit of nobody who reads it. Across all 35 conversations the total falls 27%, from 32.5 MB to 23.8 MB — the remaining bulk is many ordinary results well under the cap, so this is a fix for the write cost rather than for disk use.

  • The agent composer grows with what you type. The text area was a fixed 56 points inside a fixed 118-point panel, so a message longer than three lines scrolled inside those three lines — and the panel could not grow to show the rest, because its height was a constant. Both now follow the text, bounded at 40% of the panel's height so the transcript above never disappears and a very long message scrolls inside the text area rather than pushing the composer off the bottom. A trailing newline counts as a line, so pressing Enter at the end of a message keeps the caret in view.

  • The send button says the agent is working, and stops it. It sat inert and blue for the whole time a turn ran — the one control a person looks at to find out whether anything is happening. While a turn is in flight it is now red, carries a stop square instead of the send arrow, sweeps a ring so the motion says working, and clicking it interrupts the turn. Escape already did that, but only with the panel focused and only if you knew.

Documentation

  • Why a new terminal starts with old shell history. Terminals are backed by a pty daemon that outlives the editor, so a shell keeps running across restarts — measured on a real machine, two of them had been alive for fifty-one and forty-eight days. Both zsh and bash write their history file when the shell exits, so a shell that never exits never writes one: ↑ recalls everything typed in that terminal, while a new terminal starts from a history file that may be days old. Nothing is lost, it is simply unwritten. The README now says so, with the one-line setopt INC_APPEND_HISTORY (or SHARE_HISTORY) that changes it — and why Forge IDE does not set it for you.

The agent

Fixed

  • A background command's result is actually delivered. shell_exec tells the model "the result will be delivered automatically when it finishes" — and it was not. Ten places in the agent receive from the action channel (the streaming select, each approval wait, the shell_exec loop), and every nested one ended in a catch-all that consumed and discarded whatever it did not recognise. A background command's completion arrives while the model is still working, which is the normal case rather than a race, so it was always swallowed: the turn ended, nothing followed, and a watcher or a long build reported nothing ever. Reproduced against a real agent, where a three-second background command produced one turn and silence.

    Nested loops now park what they cannot handle, and the main loop — the only place with an arm for every action — drains that queue before awaiting the channel. Draining afterwards would leave a finished command waiting on whatever the user next happened to type. Verified the same way it was found: the agent now receives the output and acts on it.

    The regression test is structural rather than behavioural, because the property is: it asserts no action consumer discards silently, which catches the next one added rather than only the case exercised. A first version of it matched a bare _ => {} and so passed against the exact code it was written to catch — the real line carried a trailing comment.

Terminal client (forge)

No changes; released alongside, sharing the version.

Forge 0.4.1

Choose a tag to compare

@vulkgryph-dev vulkgryph-dev released this 13 Sep 17:07

Forge 0.4.1

A patch release. Three changes to Forge IDE, all from one report: a window restarted with Restart Window came back with nothing — no folder, no files, no terminal, no agent conversation.

macOS download: Forge-IDE-0.4.1.dmg below is signed with Vulkgryph LLC's Developer ID and notarized by Apple, with the ticket stapled to the disk image and the app inside it.

Forge IDE

Fixed

  • Undo history stores the edit, not the file. A step was a copy of the whole buffer, capped at a hundred of them — so editing an eighteen-thousand-line file held about 136 MB of undo history for that one buffer, and a window with two such files open accounted for most of its memory. A step is now the one region of lines that differs, with the matching lines around it not stored at all. Measured: two hundred steps on a 1.1 MB file cost 25 KB, against 215 MB as whole copies.

    The cost used to be invisible because undo did not work — snapshot was reachable only from the Tab handler, so the stack stayed empty. Making undo capture ordinary typing is what switched the cost on.

    The depth is now a setting (Undo Steps, default 200) rather than a constant. One step is one edit as a person would count it: adding or removing a line starts a step, a burst of typing becomes a single step, and saving has nothing to do with it. A cap of zero still keeps one step, because an editor with no undo is not a thing to configure by accident.

  • A saved session is a fraction of the size, with nothing dropped from it. Each terminal row was stored with a colour beside every character — {"ch":"E","fg":[204,204,204,255],"bg":null}, about ninety bytes to say one letter — so two hundred rows of scrollback came to 1.8 MB and a config directory held twelve. Terminal output is long stretches of one colour, measured at sixty-two characters per style change, so a row is now stored as its text plus the runs of colour over it. Measured on real saved sessions: 12.3 MB to 0.5 MB, twenty-three times smaller.

    Nothing is lost. Every character keeps its own colours; they are described once per run instead of once per character, and a row is reconstructed cell for cell — there is a test asserting exactly that, including for multi-byte characters, where a run has to count characters rather than bytes. Sessions written by earlier builds are still read, since they are on disk and the alternative is every window coming back empty on the first launch after the change.

  • Restart Window brings the window back as itself, even when its record is missing. A window's saved session — open files, terminals, the agent conversation — is stored under the window's id, and that id travels in the restarted process's arguments. When the shared window record no longer held an entry for it, the fallback reopened the window from its folder and dropped the id, so the new window minted a fresh one and found nothing under it. A megabyte of saved state sat on disk under the old id while the window came back apparently new. The id is carried regardless now; only the geometry and the remote connection depend on the record, and the window says so if it lost them.

    The record can go missing for an ordinary reason: it is shared between every Forge process, and the last one to write it wins, so a window restarted while another process rewrites the file finds no entry for itself.

    Windows are also planned one at a time. They were matched against the record as a set, so one absent entry sent every window down the folder-only path — and if some entries were present while others were not, the windows without them were dropped from the plan and never reopened at all. Restarting several windows for a rebuild is the ordinary case, which made that the ordinary case too.

The agent and the terminal client

No changes; released alongside forge-ide, which shares their version.

Forge 0.4.0

Choose a tag to compare

@windingcreek windingcreek released this 07 Sep 21:47

Forge 0.4.0

A new capability rather than a set of fixes: the agent gets a working area of its own, with new configuration and a changed approval rule. Under semantic versioning that is a minor bump, so this is 0.4.0.

0.3.2 was written up and its manifests committed, but it was never tagged and never released — it existed only as a commit on main, which nobody could install. Everything prepared for it is included here.

macOS download: Forge-IDE-0.4.0.dmg below is signed with Vulkgryph LLC's Developer ID and notarized by Apple, with the ticket stapled to the disk image and the app inside it — so it opens without the unidentified-developer warning, offline included.

The agent

Added

  • The agent has a working area of its own. Everything it wrote previously landed in the user's project, so a throwaway probe script, a scratch copy of a file, or a one-off reproduction either became litter in a real repository or did not get written at all. Each session now gets a directory under the system's temporary directory, created on demand and named after the session, and the model is told where it is. Subagents are handed the same one rather than making their own, so scratch work carries between them.

    Writes into it skip the approval prompt — a scratch area that asks permission for every throwaway file is not a scratch area — and the exemption is deliberately narrow. It covers write_file and edit_file, whose target is a single named path that can be checked. apply_patch names its files inside the diff and can carry several at once, so it is approved as before, and shell_exec is untouched: a command is free to go anywhere once it is running. Copying a file out of the area into a real directory is an ordinary write and is approved like one.

    The containment check is the whole boundary, so it resolves paths rather than comparing strings — lab/../../etc/passwd starts with the area's path as text while naming somewhere else entirely — and compares by path component, so a sibling directory named forge-lab-old is not inside forge-lab.

    On Linux the base is usually /tmp, shared with every account on the machine, which matters more here than for an ordinary temporary file because writes land without asking. The directories are created 0700, and one that is a symlink or belongs to another user is refused outright: left alone, a pre-planted symlink would redirect every auto-approved write the agent makes. Verified on Linux as well as macOS.

    Areas are removed by age at startup rather than on exit, because a session ends in every way a process can and one cleaned up only on a clean exit accumulates. Age is read from the newest file inside, since a directory's own mtime does not move while its contents are edited. Seven days by default; all three switches live under [agent.scratchpad], and older config files still parse.

Fixed

  • A conversation far past the context limit can recover. Two things had to be true for a session to wedge itself, and both were. Nothing capped a single tool result — read_file with no line range returns the whole file — so one call could put the history several times over the window. And the emergency path that sheds oldest messages was told the current size by way of the token count from the last successful request, which by definition describes the conversation before whatever overflowed it. Measured on a history at 1432% of the window: it decided the conversation was comfortably under budget and dropped nothing at all. The turn then failed with an API error, the oversized message stayed in history, and every following turn failed the same way. The further past the limit the conversation was, the less likely it became that anything was shed.

    Size is now taken from the history's own text whenever it may have grown since the last request, so the trigger and the recovery both see what is actually there. Recovery drops oldest turns first and then shortens the largest message still left, because dropping cannot help when the oversized message is also the newest one — and it keeps both ends of what it shortens, since a file says what it is at the top while a command run says whether it passed at the bottom. Compaction's result is fitted to the window too: the summary is small, but the recent messages kept beside it are whatever they were.

    Single results are also bounded as they arrive, at a quarter of the window, so the recovery path is a backstop rather than the thing keeping the session alive. What was cut is stated in the text, so the model knows it is holding a fragment and can read the rest by range.

  • Running a test suite is no longer mistaken for a REPL. The guard that stops an interactive program being started inside a tool call asked whether a command matched a short list of things that counted as work — -c, --version, or an argument ending in .py, .rb or .js. Everything else was called a REPL and refused, which caught a great deal of ordinary work: python3 -m unittest discover -s tests matches none of those patterns, so an agent asked to fix a failing test was told its own test command was interactive and could not verify the fix. python3 manage.py, ruby -e, and the same commands with < /dev/null already on them failed the same way.

    The question is now asked the other way round: an interpreter is a REPL when it is given no work — no module, no program, no script — and -i still asks for a prompt on purpose. Naming the ways an interpreter is given work is a closed set; naming the ways work can look is not.

  • The rendered task list is pinned by a test. Forge IDE and the TUI both parse todo_write's output to draw it as a checklist, and neither can import the function that writes it. Changing the markers or the indentation would have quietly turned both back into flat grey text.

Terminal client (forge)

Fixed

  • The dot on every tool call is a circle again, not an emoji. It was ⏺ (U+23FA, "black circle for record"), which has Emoji_Presentation=Yes and appears in neither Menlo nor Apple Symbols — so it fell through to the colour emoji font, which draws its own glyph in its own colour and discards the one the transcript asked for. That colour is the whole point of the dot: it is what separates a call that is running from one that finished and one that failed. Now ● (U+25CF), which both fonts have and which takes the colour it is given.

    This has been the same mistake three times — ⏵ and ⏸ in the permission-mode line before it — because the media-control glyphs in U+23E9–U+23FA look like plain geometric shapes in an editor and arrive as emoji in a terminal. So there is now a test that renders the transcript and the status line and fails on any of them, rather than a third fix.

  • Approving a plan with auto-accept switches the mode, rather than quietly allowing three tools. Choosing Auto-approve edits added write_file, edit_file and apply_patch to a per-tool allowlist. The edits did land without prompting, but nothing else agreed: the status line still read "Normal mode", so the one indicator of what the session will do without asking said the opposite of what it did. The allowlist was also never consulted against the permission mode, so cycling back to Normal with Shift-Tab left every edit still landing unprompted — the choice could be made but not unmade.

    It now sets auto-accept, the same state Shift-Tab reaches, so approving a plan that way means exactly what choosing it by hand means: it shows in the status line, it is recorded in the transcript, and it can be left again. Leaving plan mode no longer discards it either — the agent's own "left plan mode" event arrives after the approval and used to force the mode back to Normal.

    That also settles an ordering hazard. The plan-mode gate refuses anything that is not a read before the allowlist is consulted, so an edit arriving before that event was denied outright, and the model was told retrying would be refused again. Nothing orders those two messages; now the mode is already set when the edit arrives, so it does not matter which lands first.

  • Escape sequences no longer print as wreckage. Nothing stripped terminal control sequences from tool output, and ESC itself is invisible — so what reached the screen was [32m+++[m and [?1h=, which reads as corruption rather than as colour.

  • Command output is no longer drawn as a diff. Every tool result was passed through the diff renderer, so any line opening with - or + was rendered as a change complete with a red background: a compiler flag, a negative number, a bullet list. A shell_exec result beginning -2 appeared as a deleted line. Tinting is now gated on the content announcing itself as a diff — the edit tools' DIFF: prefix, or a @@ hunk or ---/+++ file header.

  • The transcript no longer redraws itself eight times a second. Every spinner tick erased the whole live block and reprinted it — status line, transcript tail, input box — because one character had changed. Measured on a real session: 11.2 KB/s of escape sequences, sustained, for the length of the run. It now rewrites only the lines that differ, and writes nothing at all when a frame is identical. Redrawing everything for one turning character is what a terminal shows as jitter.

    Patching is refused, and the frame drawn the whole way, whenever anything below a changed line would move — a different number of lines, or a line whose wrapped height changed. The block's own bookkeeping is kept in step, since a following erase climbs from where the caret is, and getting that wrong is what once stranded copies of the status bar down the screen.

  • A dialog no longer leaves a caret blinking in the bottom left. A plan card or an approval prompt asks for no cursor; the new patching path revealed one anyway and parked it at the end of the block.

  • **Two glyphs no...

Read more

Forge 0.3.1

Choose a tag to compare

@windingcreek windingcreek released this 28 Aug 23:15

A patch release over 0.3.0. One of these is a crash, so this supersedes it.

Fixed

  • A wide character in a message could crash the session. A session's title is its first message cut to length, and the cut was made by byte index — which panics whenever that byte lands inside a multi-byte character. Nothing exotic triggers it: an emoji in a prompt, a CJK identifier, an accented word, or the box-drawing characters a model reaches for when it sketches a diagram.

    The same pattern was in seven places across the agent and the editor — session titles, rewind previews, conversation titles in two panels, and three LSP hover truncations — each cutting text a person or a model produced, and each crashable the same way. All seven now cut on a character boundary and count their limits in characters, so a cap of 80 means eighty characters rather than eighty bytes.

  • A window with no folder open named your home directory in the status bar. A folderless window still has a working directory, and the status bar named it — claiming a folder was open while the explorer said "No folder opened" beside it. It now says No folder.

Versions

forge-agent, forge-ide and forge-tui-rs all move to 0.3.1 and ship together. The internal crates are unchanged.

Platform support, install instructions and known limitations are as in
0.3.0.

Forge 0.3.0

Choose a tag to compare

@windingcreek windingcreek released this 28 Aug 19:08

Forge is an AI coding agent with two independent clients: a terminal UI and a
native code editor. All three ship together and share one version number.

macOS on Apple Silicon is the supported platform. Linux x86-64 builds and
passes tests in CI. Linux ARM64 is exercised only as a headless remote, and
Windows and Intel Macs are untested — see the platform
table
for what that means per
component.

This release is large. The three components' own changelogs carry the full
detail: forge-agent,
forge-ide, forge-tui-rs.


forge-ide — the native editor

SSH remote development

  • Connecting to a host uploads a small static daemon over SFTP and runs it there; the remote file tree, editing, and terminal all travel over that one SSH channel.
  • The remote explorer answers a right-click, the folder chooser takes a typed path, and choosing a folder moves the workspace root — with the remote shell following by cd.
  • Open Folder and Quick Open branch on where the session is, instead of listing the local filesystem on a remote host or opening a same-named local file.
  • When this machine lends its model endpoints to a remote agent, the tunnel requires a random per-session token. Your API credential never reaches that machine.

Windows, restart, and session restore

  • Reload Window rebuilds one window in place rather than replacing the whole process, so other windows keep their conversations, language servers, and connections; shells reattach from the pty daemon.
  • A new Window menu adds Restart This Window (one window, new process, current binary), Restart All Windows across every Forge process, and Collect All Windows Into One Process. This is how you roll a new build out one project at a time.
  • Windows come back as they were, and a remote one comes back remote — the whole connection is recorded, not just a host name.

The agent panel

  • Conversations genuinely resume: Forge captures the agent's session id and passes it back on reopen, so the model recovers its prior context instead of only appearing to.
  • Tool calls render as one card each: consecutive reads fold into a checklist, writes show a diffstat and an "Open diff" popup, and long output scrolls rather than truncating.
  • A live strip above the input says what is being done to what, how long it has run, and — for a command that states its own duration — a countdown. Esc interrupts a turn.
  • Permission mode and reasoning effort are per-tab. The status strip no longer wraps onto a second row: context strategy and network state moved into a ⋯ menu, and the row sheds its separators and then its long labels as the panel narrows.
  • Three dead ends became real cards: a plan, which always waits for an explicit decision whatever the permission mode; a question; and a command blocked on a password prompt.

Subagents

  • A subagent's work is visibly a subagent's: a collapsible block, ruled down the left edge, with its activity at the top and its tool calls inside. Nested ones nest inside their parent.
  • A finished subagent no longer says "running" for the rest of the session, and the docked strip is reserved for one awaiting approval.

Terminal

  • The grid is a real fixed-size viewport with separate scrollback, with columns and rows wired through to the pty, so a resize leaves no stale duplicate content behind.
  • VT100 gaps closed: cursor show/hide, faint text, background and truecolor colours so highlight bands appear, and UTF-8 held across read boundaries rather than mangled at chunk edges.
  • Terminals survive a reload with the visible screen and their scrollback. Cmd+C/Cmd+V copy and paste, Ctrl+C still sends SIGINT, and Tab no longer leaks into the editor.
  • The pty daemon's socket is restricted to its owner.

Rendering and performance

  • Long conversations no longer make typing lag: auto-save and a full history re-parse both ran every frame. A 1442-item conversation went from ~85ms per frame to ~6ms, measured.
  • Idle CPU is roughly two orders of magnitude lower, from a demand-driven event loop and cached shaped output. Windows after the first share the graphics device and pipeline.
  • Markdown tables render as tables — the soft-wrap pass had been breaking separator rows before parsing. Wide ones scroll sideways in their own strip instead of forcing the panel open.

Editor and files

  • Redo exists, a trailing newline survives a save, each tab keeps its own scroll position, and JSON, YAML, Shell and Markdown get highlighting. LSP code actions group by kind.

Packaging and first run

  • The packaged app bundles forge-agent, so the agent panel works from the bundle alone, links OpenSSL statically, and finds the Vulkan entry point beside the running executable.
  • A first-run wizard configures a provider.

forge-agent — the agent

Shell commands no longer stall a turn

  • Any top-level shell_exec still running after agent.forced_shell_background_secs (default 5 minutes) moves to the background instead of holding the turn open. It is not killed — it keeps running as bg-N.
  • Subagent and direct shell_exec calls have a hard timeout, and delegate_task a ceiling. A hung pipeline used to return nothing at all, which is what left chats reporting "still running" overnight. Timeouts kill the whole process group, so a pipeline's grandchildren die with it.

Context and checkpoints

  • Fixed a rolling-window bug that could empty a transcript instead of trimming it, dropping context usage from around 90% to 11% in one step and leaving the agent with no memory of what it had been doing. Each dropped message is now charged its own size rather than an average.
  • A project with no git repository gets one initialised at the start of a turn; rewind checkpoints there previously had no real backing to restore to.

Subagents

  • Concurrent subagent approvals go to the right place. Approving one subagent's pending write could wave through a different one's next gated call; every wait site now checks tool_id, and DenyAction carries one — clients must send the id they are answering.
  • A subagent's read-only calls emit ToolRequest/ToolResult like the top level, and results are no longer truncated to 200 characters. Both events carry subagent_id, and SubagentStarted a parent_id, so concurrent and nested subagents can be told apart.

Providers and endpoints

  • Any open_ai-type endpoint with an api_key now actually sends it. The key was being dropped from every request, making authenticated OpenAI-format endpoints unusable, and switching models mid-session lost it too.
  • xAI endpoints discover their model catalogue on launch and prune models no longer offered. A capacity refusal is tagged Provider at capacity: rather than surfacing as a generic API error, and there is an opt-in xai_priority_tier, which xAI bills at 2× its standard rate.

Approvals and remote hosts

  • ToolRequest carries needs_approval from the session's real trust settings, so clients stop showing cards awaiting an approval nothing will answer.
  • The agent asks for explicit permission before installing git or running any package manager on a remote host, in every mode. Where git is absent it says remote revert is unavailable for that path and prefers the file tools, whose writes Forge snapshots independently of git.

Offline mode

  • offline_mode now covers the weekly Codex version self-check, which polled regardless before. web_search identifies itself as forge-agent/<version>.

forge-tui-rs — the terminal client, rewritten in Rust

This is the first release of the Rust terminal client. It replaces the earlier
TypeScript one, which is retired and removed from the repository, and it is
versioned alongside forge-agent and forge-ide rather than restarting at
0.1.0, since the three ship together.

  • It renders itself. Escape-sequence decoding, text measurement, wrapping
    and cursor movement are all its own, with no TUI framework underneath. Text
    cannot overlap, scrollback cannot be damaged, frames cannot tear, and idle
    costs nothing — each was a failure of the previous client, and each is now
    prevented structurally rather than patched.
  • /version tells you which build you are running. It reports the version,
    the commit it was built from and the binary's own build time — a version
    number alone cannot distinguish two builds between releases. It also reports
    the agent's version, since forge and forge-agent install separately and
    running different builds of the two is ordinary.
  • --dangerously-allow-all asks before it starts. The flag bypasses every
    approval gate for the session, so it now prints what that means and waits for
    you to type yes before the agent is spawned. FORGE_SKIP_DANGEROUS_CONFIRM=1
    opts a pipeline out.
  • Editing what you already sent. Up/Down walk back through previous
    messages, and Escape takes back a message the agent has not answered yet —
    only while nothing but thought has come back, since a reply, a tool call or a
    question all mean it has been acted on.
  • Ctrl-Y and /copy put the agent's last message on the system clipboard,
    using the platform helper where there is one and OSC 52 otherwise, which is
    the only route that reaches the right machine over SSH.
  • Tool results are summarised rather than dumped — a read_file shows the
    line its own output leads with, and other results show their first lines and
    say how many they are not showing. Todo lists render as a checklist and are
    never truncated.
  • Fixed: reflow on terminal resize including multi-line input; paste
    detection, so a multi-line paste arrives as one message (a heuristic, since
    Apple's Terminal does not implement bracketed paste); continuous backspace and
    caret movement; and the agent...
Read more

Forge 0.2.1

Choose a tag to compare

@windingcreek windingcreek released this 01 Jul 04:39

Fixed

  • Break the edit_file "old_string not found" death-spiral. When an edit_file target string doesn't match, Forge now returns bounded recovery hints instead of a bare error: it flags whitespace-only differences, shows the single closest-matching region with line numbers (capped — never dumps the file), and on multiple matches lists the occurrence lines to disambiguate. This stops weaker local models from looping after a failed edit.
  • Token usage and auto-compaction restored for OpenAI-compatible streaming. Forge now sends stream_options: { include_usage: true }, so spec-compliant servers (mlx_lm, vLLM, llama.cpp, LM Studio, OpenAI…) report token counts while streaming. Without it those servers sent no usage, leaving /usage and the context footer stuck at 0 and silently disabling auto-compaction — on a long local session context would grow unbounded until the model's real window overflowed.
  • Recover tool calls that misbehaving servers leak as raw text. Some OpenAI-compatible servers (notably mlx_lm at high context) fail to parse a model's <tool_call> block into structured tool_calls, instead leaking the raw markup into the content/reasoning stream and ending the turn with no tool to run — which made the agent appear to stall, loop, or return an empty turn. Forge now recovers a complete leaked <tool_call> block as a real tool call (handling both the JSON/Hermes form and Qwen3-Coder's <function=…><parameter=…> XML dialect), gated so a genuine text answer or a properly-structured call is never affected.

Changed

  • Installer/launcher hardening: the forge wrapper now locates bun robustly (~/.bun/bin/bun or PATH, with a clear error if absent), and install.sh checks for curl and ripgrep up front (the web tools and search_code need them).