Releases: Lucasllfs/forjal-releases
Release list
YHITL Forjal 0.19.0
[0.19.0] - 2026-10-01
Added
- Local map (ADR-0059): a deterministic, owner-only index of each granted project and a one-call evidence packet of ranked verbatim spans with
path:line, hash-checked on disk, answered without the runtime or a model. query_local_mapMCP tool, always loaded for Claude Code, plusforjal-codex query|explain|showand a standaloneforjalCLI.- Background enrichment by an explicitly chosen local model (
forjal-codex map enrich|collect), labelled as inferred and used only for ranking. - A
local-map-v1bench contract profile and a reproducible retrieval diagnostic.
Changed
- The resident MCP doctrine now puts the local map before search/read loops; the read worker stays explicit. The resident surface shrank from 8,184 to 8,175 characters.
- Claude and Codex prompt hooks announce the map once per session instead of relying only on curated worker cards.
- The managed Codex execpolicy allows the three read-only map subcommands; building, enriching and dropping a map still ask for approval.
- Onboarding follows the forjal.com landing: gray and slab cards with the landing's radii, regular-weight headings, and labelled pill buttons. It also fits windows down to the 860 x 680 minimum.
- Model surfaces use the landing's pastel skies instead of the red backgrounds, with ink text; the five backgrounds shrank from 12.4 MB to 224 KB.
- The model picker is ordered for the machine it runs on (ADR-0060): the app identifies the Mac by model ("13-inch MacBook Air", "Mac Studio") with its chip, cores and published memory bandwidth, estimates decode speed per quantization, and weighs it against the published Artificial Analysis index and the precision with a Fast-to-Smart control. Estimates and third-party scores are labelled as such; NPU lanes keep their curated order and runtimes are unchanged.
- On a Windows PC with a discrete graphics card (GeForce RTX, Radeon RX, Arc), the app reads the card's name and video memory from its driver and estimates GPU deployments that fit in that memory at the card's published bandwidth, instead of the system memory's. The hardware step names the card.
- Eleven new Apple MLX candidates: Qwen3.5 2B, the first model an 8 GB Mac can run; 8-bit variants of Qwen3.5 4B/9B, Gemma 4 E4B/12B/26B-A4B/31B and Qwen3.6 27B/35B-A3B for Macs with room for them; Gemma 4 E2B; and Qwen3.5 122B-A10B for 96 GB+. Each ships with a draft capability profile. No existing machine's slot picks changed.
Notes
- Software-verified only; not yet installed or evaluated on the test Mac. No session token-savings claim is made.
- Speed figures in the model picker are bandwidth estimates, checked only against one measured run (Qwen3.8 27B 4-bit on an M2). The 8 GB cohort has not been observed.
- Hardware tables and the ranking policy derive from Magnitude (Apache-2.0); see THIRD_PARTY_NOTICES.md.
YHITL Forjal 0.18.6
Forjal 0.18.6 — Qwen3.8 and clearer local runs
This Alpha consolidates the general workspace harness and makes live and historical runs easier to inspect.
- Qwen3.8 27B in Models. The interface shows the resolved deployment and quantization. On Apple Silicon with 24–31 GB of unified memory, explicit selection prepares the normal UD-Q3_K_XL candidate with llama.cpp / Metal and an 8K context. Weights are downloaded separately and verified. Other hardware cohorts retain their existing deployment selection.
- General workspace harness. Eligible native workspace runs use the model-directed tool loop, bounded navigation, command guidance and structured final evidence. Read, write and run capabilities follow the caller's grant. Recovery preserves the admitted policy; unsupported paths retain their existing contract.
- Better run visibility. See the current stage, elapsed time, completed turns, local actions and recorded output tokens. Expand details for prompt, cache and decode counters when available. The timeline distinguishes reads, writes, commands and supervisor tool calls, preserves generation errors, and separates execution completion from collection by the supervisor. A refresh failure keeps the last recorded steps available.
Qwen3.8 and the general harness remain supervised candidates, not certified capabilities. Workspace tools for Qwen3.8 are enabled only on the specified Metal Q3 deployment. The supervisor must review results, evidence, edits and command receipts. This release does not establish a session token-savings percentage or a speed or quality advantage over Qwen3.6.
Alpha distribution: the macOS application is ad-hoc signed, without Developer ID signing or notarization. Windows installers are unsigned. Operating-system launch warnings are expected. Choose the installer for your platform and verify its published checksum.
YHITL Forjal 0.16.0
Forjal 0.16.0 — Ornith 1.5 candidate on Apple Silicon
This release updates the macOS Apple Silicon build. The Windows ARM64 and
Windows x64 downloads remain on version 0.15.0.
Ornith 1.5 9B candidate
Forjal adds ornith-1.5-9b as a text-only candidate for eligible machines with
at least 16 GB of unified memory. After the user or supervising agent explicitly
selects that model_id, the load policy resolves one pinned Apple MLX deployment:
- 16–23 GB: official MLX 4-bit artifact;
- 24–31 GB: official MLX 6-bit artifact;
- 32 GB or more: official MLX 8-bit artifact.
The initial plans use an 8K context, thinking enabled, and MTP disabled. The
resolved deployment is prepared and activated atomically. A failure does not
silently switch quantization, runtime, context, offload, or model.
Prepared-plan fidelity
Prepared deployments now retain the adapter's pinned chat template and sampling
configuration in the immutable execution plan. Existing model profiles and
deployments remain available.
Availability and limits
Ornith 1.5 is a candidate. It is not compatible, certified, or recommended,
and this build has not completed end-to-end validation on physical Apple Silicon
hardware. Model weights are not bundled in the installer; the first load
downloads the selected artifact from its pinned upstream revision.
Alpha build. The macOS application is ad-hoc signed and not notarized, so
macOS asks the user to confirm the first launch.
YHITL Forjal 0.15.0
More local agents
Forjal adds managed adapters for Pi, OpenCode, OpenClaw, and Hermes Agent
alongside Codex and Claude Code. Each adapter follows the host's own
configuration format, timeout, ownership, and reload behavior. The four new
adapters currently have contract, configuration, and fake-host evidence;
real-host end-to-end validation is still pending.
Prepared deployments
Loading an explicitly selected model_id can now resolve one pinned deployment
for the detected hardware and prepare, verify, calibrate, and activate it as one
flow. A failure does not silently switch backend, quantization, context, runtime,
or model.
New candidate path
Qwen3.8 27B and the llama.cpp multi-runtime path enter as candidates on eligible
hardware. They are not certified or recommended yet. Qwen3.8 is excluded from
8 GB systems and Windows ARM64, and availability depends on the deployment
actually packaged for that platform.
Local operational metrics
The desktop adds local operational metrics and an opt-in shareable receipt.
Prompts, outputs, source code, paths, and URLs are excluded.
Alpha build. The installers carry no public signing identity. Windows may
show SmartScreen, and the macOS build is ad-hoc signed and not notarized.
YHITL Forjal 0.14.0
The local model works here now
Forjal used to be a bridge: the local model asked, your cloud agent executed, and a
hook carried the payload back. 120 measured sessions on a MacBook Air M4 said that
shape cannot pay for itself. Every action the local model needs from the parent costs
three of the parent's turns — one to receive the request, one to run it, one to
resume — and a turn re-reads the entire context. Measured against an agent that just
did the work itself: US$ 0.26–0.77 delegating, US$ 0.077–0.11 not.
The one path that won money was the opposite one, where the local model acts on its
own. So that is the product now.
workspace replaces local_read, with three grades ordered by what can change on
your disk:
read— search and read the directories your client declared. Nothing changes.write— also create files. Every one of them is named in the result.run— also run this project's build, tests, linters and read-onlygit. Build and
test output is the largest payload in real work, and it used to be the one thing
only the parent could fetch.
run contains write in fact and not by convention: no test suite runs with a promise
that it leaves your tree untouched.
run_command is not a sandbox and does not claim to be. A project's own test suite
is arbitrary code by construction. What is enforced instead is narrow and checkable:
the model never names a program — only one from an allowlist Forjal wrote, and within
it only the operations Forjal listed; there is no shell, so a pipe or a redirect is a
literal character to the program that receives it; every argument that could be a path
is resolved inside your granted directories and refused against the credential
deny-list; the environment is scrubbed of anything that looks like a secret and of
every interpreter hook before the child process exists; and every command that ran comes
back in the result with its exit status.
Three things the same measurement asked for
min_capability — you name the least capable deployment allowed to answer, and a
machine that cannot serve it refuses instead of trying something smaller. The same
delegation was measured delivering 24 of 24 sections twice and nothing once, purely
because the machine was busy enough for the selection policy to fall to a smaller model.
The policy is right and has not changed; what was missing was your ability to say a weak
answer is worse than none.
shortfall — an answer that is empty, that says it could not do the task, or that
names almost none of the documents it was given comes back marked. The output is still
returned whole; you are just told what it is.
evidence — an answer built on files comes back with the verbatim passage behind
each claim, and Forjal reads each one off the same disk after the answer was written.
A quote that is not really there stays in the list, marked. In 11 of 13 measured
sessions the agent re-read the material after delegating and paid for both paths;
forbidden to re-read it got 60–63% cheaper and was wrong every time, so the re-reading
was the correction it had to apply. This is what removes it.
Fixed
- Two delegations with the same prompt over different files no longer collide. A run's
identity left the files out, so six documents delegated with one prompt collapsed and
five callers received an answer about a document they had not asked about. If you are
on 0.13.0, this is the reason to update. - The Claude Code hook matches again. It compared tool arguments by equality, and
Claude Code attaches adescriptionto everyBashcall it runs — so the two never
matched, the hook stayed silent, and the payload travelled through your context anyway. - Malformed tool calls are repaired before they are refused. Single quotes, a trailing
comma, a code fence, a missing brace. Five unusable calls in a row ended one measured run
with nothing delegated; the intent was right every time. - The read budget now follows the context window of the deployment that will answer,
instead of a flat ceiling that ended searches with the window still nearly empty — and
the model is told what it has left.
Windows catches up
Windows ARM64 and x64 were still on 0.9.0. Both are on 0.14.0 with this release.
What is not measured yet
Everything above is built and tested — 483 tests across the runtime and the bridge — but
the harness itself has not been through a paired measurement on real work. The numbers
quoted here are from the sessions that produced the design, not from the code that came
out of it. Treat workspace: "run" as what it is: a new surface, in an Alpha.
Alpha build. The installers are native to each platform but carry no public signing
identity, so macOS and Windows both ask you to confirm the first launch. Each package
carries a private Python, the local runtime and the MCP bridge. Models are downloaded on
first use and are not bundled.
YHITL Forjal 0.13.0
What changed
This release closes the last gap that produced a wrong answer instead of an error.
Every delegation that let the model search on its own now comes back with
workspace_search: the files the answer was drawn from, how many reads, how many
characters, and which ceiling ended the search when one did. Before, a search that
read 28% of the material and a search that read all of it were the same completed
response with the same confident prose. A search cannot name what it never reached —
nobody enumerated the corpus — but it can name what it did reach, and that is what
lets the agent judge coverage and ask for the rest by name.
A delegation without model_id whose default was decided by free memory now comes
back with model_note, naming the downloaded deployment the policy would otherwise
have chosen. The policy does not change — a default that drags the machine into swap
is worse than a smaller one that runs — but it means the same delegation, repeated
minutes later, can be answered by models of very different capability, and the caller
only found that out by reading the model_id of a result it had already accepted.
forjal-output no longer leaves a zero-byte file when it fails. The recommended
form is now -o <path> instead of shell redirection: the shell creates the
destination before running the command, so every failure — run not finished,
runtime unreachable, wrong run_id — ended with the file the user asked for sitting
at zero bytes.
A prompt refused for size now reports the size sent, the ceiling, and the
alternative. prompt (string_too_long) left no strategy but trial and error, and
trial and error worked: in a real session the agent shrank its request three times and
delivered a task about one of six documents, with nothing recording the scope
degradation.
See ADR-0034.
This build
Alpha channel, macOS Apple Silicon only. Windows stays on 0.9.0.
This installer is not signed and not notarized. macOS will refuse to open it on
first launch: open it from Finder with right-click → Open, or clear the quarantine
attribute manually.
YHITL Forjal 0.12.0
What changed
This release repairs the feature 0.11.0 shipped. If you are on 0.11.0, update.
A delegation that names files reaches the model again. files and the caller's
prompt shared the same 20,000-character ceiling, so a plan the preload budget had
already approved died at validation with string_too_long — after the files had been
read, which is the worst possible moment to fail: the cost is already paid and the
agent gets a schema error instead of an answer. What the engine may be handed is now
the sum of the prompt and the preload, and it stays a backstop: what actually fits is
decided per run, from the loaded model's context window.
A files that matches nothing no longer spends a run. A pattern with no readable
file is refused at admission, with a message naming what was asked for and where the
search happened. Before, the agent got an empty answer and concluded that delegation
does not work, going back to reading the files itself.
Every delegation with named files now reports its read plan beside the answer: what
was read, what was left out because the window ends, and which patterns matched
nothing — each one by name, so they can be passed back in a second call. An answer
built from ten of sixteen files without saying so is the most expensive failure this
product can have: it looks complete and is not.
This build
Alpha channel, macOS Apple Silicon only. Windows stays on 0.10.0.
This installer is not signed and not notarized. macOS will refuse to open it on
first launch: open it from Finder with right-click → Open, or clear the quarantine
attribute manually.
YHITL Forjal 0.11.0
What changed
One call, every file. run_local_model gains files: the agent names workspace
paths or globs, Forjal reads them on this machine and puts them in front of the local
model, without the content ever entering the agent's context. It is the cheapest
delegation there is — one call instead of one read per file — and every reply states
what was read and what did not fit the model's window, with the exact paths for a
second call. files and local_read are independent grants and they compose: with
both, a small deployment answers in a single pass and a capable one keeps working from
the material. Whether the loop is honored is decided by the model's evaluated profile,
not by the bridge — the same call yields more as the user installs better models.
The delegation trigger is now something the agent can evaluate before reading. The
server instructions count files, assessed before the first read, instead of an
accumulated character volume. The old criterion failed on exactly the measured case:
the agent decides file by file, each one looks small, and the total is never assessed.
Five real sessions with Forjal connected and a model loaded delegated zero times.
model_id is optional. Omitted, the run uses that machine's curated default, and
the reply always names which model ran. This is not routing: Forjal still does not
choose by task type and never substitutes a model. A required argument the agent cannot
fill without a discovery call is a toll before the first delegation, and that is where
it gave up.
A delegation over a large file no longer loops. read_file gains offset and
limit, and every reply declares the range delivered and the file's total; a truncated
result says how many lines remain, which offset to continue from, and how much of the
task budget is left. Before, a file larger than one result returned the same beginning
every time and no call could ask for the continuation — the only action in the model's
repertoire was to repeat the read, until the runtime ended the run. A line longer than
an entire result, such as a minified bundle's, is now reported as a character-level
truncation with a grep to use, instead of being announced as a complete file.
This build
Alpha channel, macOS Apple Silicon only. Windows stays on 0.10.0.
This installer is not signed and not notarized. macOS will refuse to open it on
first launch: open it from Finder with right-click → Open, or clear the quarantine
attribute manually.
YHITL Forjal 0.10.0
What changed
A delegation that reads project files no longer costs the agent a turn per read.
With local_read, Forjal reads, lists and searches by itself, inside the directories
the supervising client declared over MCP roots/list — the same boundary the user had
already granted the agent, obtained without spending a context token on either side.
The agent grants it with a single boolean instead of drafting three to five tool
definitions on every delegation. Nothing writes, executes or reaches the network; a
file that looks like a credential is refused even inside a granted directory; and a
model that repeats the same read ends the run with a reason instead of spending its
whole budget on one file. Access is evaluated per model profile, exactly as the context
request already was.
The agent can see delegation as an option again. Clients that load tool schemas on
demand keep only the server's instructions and the tool names resident in context, and
Forjal kept nearly all of its delegation doctrine in the tool descriptions — so the text
that was supposed to cause the decision was only read after it. The instructions now
carry the triggers and the name of every tool, with a criterion the agent can apply
before reading the material: measure the file, do not estimate it.
Delegating no longer requires a call just to discover what to run. The
run_local_model schema names the models installed on that machine, with a short
summary and the curated default for that hardware marked. Forjal still does not route:
the names are reported, the choice is the agent's.
A run the supervising agent never answered no longer hangs. The runtime closes a
suspended run after 15 minutes without a reply, and the clock restarts on every sign
that the agent is working. abandoned is recorded apart from failure — the local model
did its part and nothing on this machine broke — and the activity list gains "End run",
because the agent that started a run may not exist any more.
The product has a public page. The version, size, checksum and per-lane deployment
counts come from the published manifest and the runtime catalog; none of those numbers
are typed into the page.
This build
Alpha channel. This installer is not signed and not notarized. macOS will refuse to
open it on first launch: open it from Finder with right-click → Open, or clear the
quarantine attribute manually.
Windows stays on 0.9.0 in this release.
YHITL Forjal 0.9.0
YHITL Forjal 0.9.0 — Alpha
This release adds a fourth platform lane, replaces the recommendation badge with a declared selection policy, and makes the delegation savings a measured number instead of a claim.
Intel laptops now accelerate
A new openvino lane runs on the AI Boost NPU of Core Ultra chips and on Arc GPUs, with the device declared per deployment and no runtime choice. It carries the same guarantees as the other three lanes: artifacts pinned by commit, SHA-256 manifests, no executable Python inside an artifact, and a refusal to fall back to the CPU when no accelerator is present. Until now an Intel machine only had Foundry Local.
Six curated Intel deployments ship with it. On the NPU: Qwen3 8B and Mistral 7B from 16 GB, and DeepSeek R1 Distill 1.5B from 8 GB — the first time a base-memory machine reaches an accelerator in any lane. On the Arc GPU: Qwen3 4B from 8 GB, Qwen3 14B from 16 GB, and gpt-oss-20b (20.9B total, 3.6B active) from 24 GB.
The artifact format is now a rule the catalog enforces. Intel's NPU only accepts symmetric INT4, so a checkpoint exported with the common asymmetric recipe would download, occupy disk, and only fail at compile time. The schema now rejects that combination when the catalog loads, and the same applies to the context window, which the NPU compiles to static shapes.
Model selection became a policy
Deployment choice is now decided by ADR-0027. The three accelerated lanes used to sort by declared memory and recommend the first, which made the recommendation a function of the computer alone — a 32 GB Mac was pointed at a 31B dense model when a 26B MoE with 4B active fits the same machine and answers far faster. Fit is now arithmetic over the artifacts, order within each slot is curation with written provenance, and the pick is policy: one per slot, defaulting to balanced and never to max-capability.
Recommended is gone from the interface. A recommendation badge asserts more than free memory supports. In its place: the slot, Starts here for the entry point, and a sentence made only of facts the runtime can show — parameters, active parameters when they differ, weight size, and the memory the machine reported. No quality claims, because nothing in this catalog was evaluated on the machine reading it.
Delegation savings are measured, not asserted
Activity used to count tokens spent locally, which is consumption and grows precisely when delegation goes well. What was missing is the counterfactual: the context the supervising agent did not have to carry. Only payloads delivered by the direct hook count — a result relayed inline crossed the agent's conversation and saved nothing. The estimate is declared as an estimate on every surface and is deliberately a floor.
list_forjal_models now adds measured_here to every model with history on this installation — runs, real duration, failures — and get_forjal_status reports the delegation total with the basis of the calculation. A model that has never run reports no field rather than zeros.
A delegation that cost more context than it displaced now says so in its own result. The work is still done and returned; nothing is refused.
Windows no longer replaces itself in place
The update cycle could leave the desktop shortcut pointing at an executable that no longer existed. The cause was the mechanism itself: the app swapped its own installed binary by relaunching the NSIS installer silently — exactly the behavior pattern SmartScreen and Defender exist to flag. Windows now behaves the way ad-hoc-signed macOS already did: it announces the new version and takes you to the download page. Full reasoning in ADR-0029.
Memory accounting you can act on
- A refusal now carries the whole calculation — total required, both parts, free memory, and how much must be freed — instead of naming only the working set while checking working set plus reserve.
- Loading a model with another already resident is no longer refused; the load releases the resident one itself, and the interface says which model will be released before the click. The swap is refused before releasing when even then there would not be enough.
- The runtime returns memory it still holds in the MLX cache before telling anyone to close applications, then re-reads.
- macOS free-memory reads no longer double-count purgeable pages, the error that lets a load through and drags the machine into swap.
- Free memory is now a load prerequisite on all three accelerated lanes, checked before the download rather than after tens of gigabytes.
Hidden reasoning stays out of the agent's context
Reasoning from models like GLM-4 and Qwen3 reached the supervising agent whole. Suppressing thinking through the chat template is a request, not a guarantee. Three shapes are now stripped — the closed block, the orphan close, and the orphan open from a generation that ran out of budget mid-thought. Stripping happens before any interpretation, which also prevents a tool the model merely considered inside the block from being executed as if it had been requested.
Output budget is no longer a flat 2,048 tokens for every model. It is derived from the window of the variant actually loaded, with extra headroom when the model reasons, minus what the prompt occupies, and never below what the flat cap delivered.
Downloads can be cancelled
Choosing another model is itself the cancellation, and the preparation screen has an explicit "Stop and choose another model". Artifacts are fetched file by file instead of through the Hugging Face snapshot helper, which offers no way to stop — partial files stay on disk, so resuming starts where it left off.
This is an Alpha build. The installers are native to each platform but carry no public signing identity: macOS is ad-hoc signed and not notarized, Windows ships without Authenticode. Both ask you to confirm the first launch. Each package carries a private Python, the local runtime, and the MCP bridge. Models are downloaded on first use and are not bundled.