Forge 0.6.0
A minor release rather than a patch, because it changes which headers leave your
machine and adds three configuration keys.
macOS is the supported platform for this release. Windows and Linux install
from source — see Platforms.
Security
A crawl could be aimed at the machine Forge was running on. There was no
restriction on which addresses web_search and web_fetch could reach:
localhost, the RFC1918 ranges, and 169.254.169.254 — where a cloud instance
serves its own credentials — were all reachable.
The delivery route needed no cooperation from the model. robots.txt was
checked for the URL that was requested; the fetcher then followed up to five
redirects on its own, and whatever came back was indexed without the
destination ever being checked. One hostile page in an ordinary crawl could
redirect to 127.0.0.1:2375 and have a local service read back to the model as
a page from the web.
Two exposure windows, checked against the tags rather than estimated:
web_fetch has followed redirects since 0.1.0, so a single fetch could be
aimed at a local address in every release there has been; the crawler arrived in
0.5.0, so the redirect-and-index route affects 0.5.0 through 0.5.2.
Now refused at two layers, and every redirect hop is resolved and checked before
it is followed. This is not proof against DNS rebinding, and the code says so
rather than implying more.
A cross-origin redirect bypassed robots.txt, stay_on_host and per-host
politeness — the same root cause, fixed with it.
Forge described a page as one a person had opened in a browser when nobody
had. When a live fetch came back behind a bot check, web_fetch served an
indexed copy with the words "already opened in a browser by the user and handed
over" — inferred from the URL being in the index and the fetch being refused.
A page crawled weeks ago, on a site that has since added a challenge, satisfies
both. That is a claim about a human action reconstructed from a proxy for one,
and it is the strongest thing this project says about that route. The handover
records its own attribution now.
The crawler can prove who it is
Forge already sent an honest user agent with a contact URL and refused to
pretend to be a browser. But a header is a claim, and anyone can write one.
Requests can now carry a Web Bot Auth signature — RFC 9421 HTTP Message
Signatures — so a site can check the claim instead of weighing it. Ed25519
over the authority, with the public key published at
/.well-known/http-message-signatures-directory on a domain you control.
Cloudflare validates these at its edge as part of Verified Bots, which means a
small crawler that behaves well can be recognised without a contract.
Off unless configured. forge-agent --generate-bot-auth-key makes the key —
no openssl needed, which matters because macOS ships LibreSSL and it cannot
generate this kind of key at all.
The crawler obeys what sites actually said
Five places where "we could not tell" was read as "go ahead": a transport error
on robots.txt returned no restrictions with no retry; a 200 was trusted
without being looked at, so a robots.txt redirecting to a login page permitted
everything; there was no size bound; and the cache dropped the scheme.
And two where the site was simply overruled. User-agent: * / Disallow: /
followed by User-agent: forge-search / Disallow: resolved to the blanket ban
rather than the exception — the standard "everyone out except you" idiom, and
exactly the file a site writes after someone asks for access. And
Crawl-delay: 86400 was clamped to 300 seconds: a site asking for one visit a
day got one every five minutes, 288 times what it asked, from code describing
itself as obeying.
429 and 503 were ignored — counted as "missing" while the crawl carried on
at the same rate. Now honoured with Retry-After.
Privacy
A re-crawl returns If-Modified-Since but not If-None-Match.
Last-Modified describes the content and is identical for every visitor; an
ETag is server-chosen and can be minted per visitor, which is a known tracking
technique. agent.send_etag opts in.
The README now states what a site can learn, measured rather than asserted: the
complete header set is Host, Accept and the user agent. No cookies, no
Referer, no Accept-Language, no session identifier.
Three ways the agent could stop being useful while looking fine
A message typed while it was working could be discarded. Stdin was read with
a future that is not cancellation-safe, racing the agent's own output — so an
event arriving mid-line dropped the read and the bytes already consumed. You
waited for an answer to something it never saw.
A message typed mid-turn could end the work instead of steering it. It
arrived as a plain user message, indistinguishable from starting a new
conversation — and answering someone who has just spoken to you means stopping.
A compaction could end the work it interrupted, because the summary read as
a status report and the natural continuation of one is another.
Also: a local-only setup contacted OpenAI twice at startup about a provider it
was not using, and an approved plan could leave the agent with no legal move.
One fewer dependency
scraper is out — 27 packages, including four of the five MPL-2.0 crates in
the tree. Forge has had its own HTML parser since 0.5.0, so web_fetch and
web_search were reading the same page two different ways. They share one now,
and the crawler gained that parser's skip list: navigation and footers are no
longer indexed as though they were what a page is about.
Editor
forge-ide has a changelog entry for the first time in a release where the
editor did not change — the SSRF above was reachable from its agent panel, and
an editor user reading only that file should not have to find out elsewhere.