Repository navigation
Releases: akynte/boundedcode
Release list
BoundedCode v0.1.0-alpha.5
BoundedCode v0.1.0-alpha.5
Security fixes, and releases you can verify.
This alpha fixes the defects found in the public-launch
readiness audit. Two of them let an
agent's work reach past BoundedCode's policies on the host. It also moves
to Go 1.27.2, which fixes eight standard-library vulnerabilities.
It is the first release built and attested in CI. You can check that each
file was built from this tag by this repository's workflow, and you can
rebuild it byte for byte.
The feature set is that of v0.1.0-alpha.4. The local
default, Qwen3.6-35B-A3B on Linux with an NVIDIA GPU, is still the only
validated configuration.
Security
- Nested repositories.
- Before: host git ran commands configured inside a git repository
that the agent created in its worktree. Once such a repository was
recorded as a gitlink, its own filter drivers ran on the host during
checkpoint commits. - Now: host git ignores submodules, and a checkpoint refuses a
worktree that holds a nested repository.
- Before: host git ran commands configured inside a git repository
- Quoted file names.
- Before: names that git quotes (non-ASCII characters,
", control
characters) did not match the protected-path, secret-path and deny-path
policies. A new.github/workflows/cié.ymlpassed the diff-scope check. - Now: changed-file lists are read NUL-separated and unquoted.
- Before: names that git quotes (non-ASCII characters,
- Go 1.27.2. It fixes eight standard-library vulnerabilities reachable
from BoundedCode, in net/http, its HTTP/2 transport, crypto/tls and
mime/multipart (GO-2026-6603 to GO-2026-6617). The release binaries are
built with it.
Verification
- Go behavioural evidence.
- Before: a test that failed on the base, but was skipped, failed or
did not run on the change, counted as evidence. - Now: each attributed test must also pass on the change.
- Before: a test that failed on the base, but was skipped, failed or
Installation and first run
bcode setup --checkexits non-zero while any step is missing.- Linux with an NVIDIA GPU but no CUDA toolkit: set-up downloads the
pinned prebuilt CUDA build of llama.cpp, instead of starting a source
build that fails. - Cloud provider: the task contract (and its ambiguity check) now uses
it. It was silently skipped before. - The installers are more careful:
install.shrefuses to replace an unrelatedbcodeorboundedcode
(BC_FORCE=1overrides) and replaces the binary atomically.
install.ps1stages files before replacing them.- Both warn when another
bcodecomes first on PATH, fall back to the
releases feed when the GitHub API rate limit is hit, and accept a
mirror (BC_DOWNLOAD_BASE). install.ps1no longer closes your PowerShell window when it fails.
- Smoke tests for the installers on Linux, macOS and Windows, and for a
first task end to end, now run in CI.
The CHANGELOG lists every change.
Releases you can verify
- Built in CI. A
releaseworkflow builds the release from the tag
withmake dist, after the tests. It creates a draft, which the
maintainer reviews and publishes. - Attested. Every file carries a build-provenance attestation (SLSA,
signed through Sigstore). To check one:gh attestation verify boundedcode-linux-amd64 -R akynte/boundedcode
- Reproducible, SBOMs included.
- The SBOMs now take their creation time from the commit, and each
target has its own document namespace. - Rebuilding from the tag with the Go version in
go.modreproduces
every binary and SBOM. The licenses archive also matches when rebuilt
underumask 022, as on the CI runner: at this tag, its permission
bits still follow the umask. Later releases normalise them. - The steps are in
Verifying a release. - Earlier releases: v0.1.0-alpha.4's binaries and licenses archive
were already reproducible; its SBOMs were not.
- The SBOMs now take their creation time from the commit, and each
Evidence published since alpha.4
- Comparative evaluation:
BoundedCode vs. the same OpenHands agent, model and sandbox without it,
on 8 pre-registered public tasks.- Results: hidden tests passed 7 of 8 vs. 6 of 8. No difference in
success is shown, and BoundedCode took about twice the time and tokens. - What is published: the protocol, the scripts, the raw results, the
deviations and the negative results.
- Results: hidden tests passed 7 of 8 vs. 6 of 8. No difference in
- Evidence demo: a reproducible
demonstration of behavioural verification. It includes a recorded real
local-model run.
Benchmarks for this release. The release checklist asks for a benchmark
re-run. The task benchmark is the comparative evaluation:
- Same code. Its 24 agent runs and 16 gate replays used a build of
commit 86c123d. Since then, no product code has changed except the Go
version. - Speed not re-measured.
bench infrawas not re-run: llama.cpp and
the model are unchanged since alpha.4.
Contributing
- Issue forms for bugs, installation problems, platform reports, wrong
verification results, feature proposals, model reports and evaluation
results. - Contributor onboarding, with first
issues taken from observed defects. - Submitting an evaluation result.
- CI now lists skipped tests, and shows macOS and Windows test failures
instead of hiding them.
Upgrade notes
- Building from source needs Go 1.27.2 or later (
go.mod). - Tasks with nested repositories. Some tasks now fail at their
checkpoint commit with "nested git repository … refusing to run host git
on it", naming the directory:- Which tasks: those where a git repository appears inside the
worktree, for example one the agent created or cloned, or a submodule
it checked out. - Unaffected: task worktrees start without submodules checked out.
- Which tasks: those where a git repository appears inside the
setup --checkin scripts. Scripts that runbcode setup --check
must accept a non-zero exit while set-up is incomplete.
Known limitations
- macOS and Windows remain experimental. Their unit tests do not pass:
5 packages fail on macOS and 14 on Windows. Four of the causes are
filed as good first issues. - Earlier limitations. The limitations of
v0.1.0-alpha.4 and
v0.1.0-alpha.3 still apply.
Please report what you find, using the
issue forms.
Assets
- Binaries:
boundedcode-linux-amd64,boundedcode-linux-arm64,
boundedcode-darwin-amd64,boundedcode-darwin-arm64,
boundedcode-windows-amd64.exeandboundedcode-windows-arm64.exe.
They are static (CGO disabled). SHA256SUMS: checksums, verified by the installers.SBOM-<os>-<arch>.spdx.json: an SPDX 2.3 software bill of materials
for each binary.boundedcode-v0.1.0-alpha.5-licenses.tar.gz:LICENSE,NOTICE,
THIRD_PARTY_NOTICES.md, and the license texts of every Go module
compiled into the binaries (LICENSES/go) and of the upstream
components (LICENSES/upstream).- Attestations: a build-provenance attestation for each of the files
above, stored with GitHub.
Model weights are not distributed. Download them with bcode setup or
bcode model fetch, from their publisher at a pinned revision, and review
their license.
BoundedCode v0.1.0-alpha.4
BoundedCode v0.1.0-alpha.4
Do the repositories still work together?
A multi-repository task can pass every repository's own checks and still
break the contract between them. For example, a provider renames a protobuf
field and updates its server, while a client in another repository still
builds against its vendored copy of the generated code. This alpha adds a
cross-repository compatibility gate (experimental). For every gRPC,
protobuf or OpenAPI link that a task's change affects, it reports
compatible, broken or untested, and it records the evidence: the
commands run, their outcomes, and the repositories and commits involved.
Everything else is as in v0.1.0-alpha.3. The local
default, Qwen3.6-35B-A3B on Linux with an NVIDIA GPU, is still the only
validated configuration. The gate is covered by unit, integration and CLI
tests on fixtures, and was run with the real Docker sandbox. It has not
yet been run on real tasks.
Highlights
-
Only real changes count. The gate compares the change's base and head
commits:.protomessages, fields, enums and RPC signatures are compared
structurally, and breaking changes are recognized: a field removed,
renumbered, retyped or renamed, or an RPC signature changed.- OpenAPI operations are compared together with the schemas they
reference. - Go code is compared by the function that holds the call.
- Comments, other RPCs and unrelated functions do not count.
-
Checked with the repositories' own tests, against each other's new
commits:- gRPC and protobuf (Go): the dependent repository's
go teststage
runs in the sandbox, built against the provider's candidate commit
through a generated Go workspace. A vendored copy or areplace
directive does not apply, and coverage must show the dependent's code
ran. - OpenAPI: the dependent's tests must fail when the operation is
removed from the spec, which shows they actually read it. - Failures: control runs at the base commits decide whether a failure
comes from the change or was already there.
- gRPC and protobuf (Go): the dependent repository's
-
In the task loop:
- A
brokenlink fails the attempt, and the agent is told which link
broke and why. - An
untestedlink withholdsTASK_VERIFIED. When a test could settle
it, the agent is asked once for one. - Results survive a crash or
task resume. A new commit marks them stale.
- A
-
Where to see it:
task status,verify --full(also--json) and the
Verification tab of the interface show the per-link report:cross-repository compatibility: BROKEN (3 broken, 2 untested, 0 compatible) BROKEN grpc_def grpc shop.payments.v1.PaymentService/Charge [changed] checkout@a5f2bdbb99 internal/pay/client.go:23 -> protos@f70b3b206e payments/v1/payments.proto:9 checkout's checks fail with protos's candidate (checkout@a5f2bdbb99 + protos@f70b3b206e) and pass with protos's base commit: internal/pay/client.go:23:77: unknown field AmountCents in struct literal of type paymentsv1.ChargeRequestThis example is the new
contract-breakfixture:
a field rename that all three repositories' own tests pass. See the
design and its limits. -
Vendored Go modules verify offline. Verification no longer forces
-mod=modon a module withvendor/modules.txt. That flag made Go ignore
the vendored copy and try to download it. A vendored module now builds
from its trackedvendor/.
Documentation
The README and docs at this tag were written before two clarifications
that are now on main:
- In a multi-repository task,
TASK_VERIFIEDalso requires every affected
gRPC, protobuf or OpenAPI link to becompatible. See
Verification and
the "How it works" diagram, which now shows the gate. - The sandbox controls
have an entry for the gate and describe vendored Go modules, and
CONTRIBUTING.mdlistsinternal/compatas security-sensitive.
The released binaries are unchanged; only the documentation was updated.
Upgrade notes
- The database migrates itself (migration 6).
- Behaviour change: a multi-repository task whose change affects a
gRPC, protobuf or OpenAPI link now endstask_verifiedonly if every
such link is showncompatible. Otherwise it endstests_green, with the
reason recorded as a decision. To keep the previous behaviour, set
repointel.compat_gate: false. - No new dependencies, and the sandbox image is unchanged.
Known limitations
- The gate checks gRPC and protobuf sides written in Go only: one module
at the repository root, ago teststage, and generated code committed
in a task repository (the gate never runsprotoc). OpenAPI sides are
checked only when a test reads the spec. Every other case is reported
untested, with the reason. compatiblefor a client and server means that both compile against the
same new definition and their own tests execute their side. The client
is never run against the real server.- A wire-incompatible definition change stays
untestedeven when every
task repository agrees. Services already deployed are not tested. - Test code is written by the agent and runs inside the check, the same
trust boundary as for verification. Likewise, the agent can edit a
vendored module'svendor/like any other source; only review catches
that. - The limitations of v0.1.0-alpha.3
still apply.
Please report what you find:
issues. A link reported
broken that was fine, or compatible when it was not, is especially
useful.
Assets
boundedcode-linux-amd64,boundedcode-linux-arm64,
boundedcode-darwin-amd64,boundedcode-darwin-arm64,
boundedcode-windows-amd64.exe,boundedcode-windows-arm64.exe: static
binaries (CGO disabled).SHA256SUMS: checksums, verified by the installers.SBOM-<os>-<arch>.spdx.json: SPDX 2.3 software bill of materials of each
binary.boundedcode-v0.1.0-alpha.4-licenses.tar.gz:LICENSE,NOTICE,
THIRD_PARTY_NOTICES.mdand the license texts of every Go module compiled
into the binaries (LICENSES/go) and of the upstream components
(LICENSES/upstream).
Model weights are not distributed. Download them with bcode setup or
bcode model fetch, from their publisher at a pinned revision, and review
their license.
BoundedCode v0.1.0-alpha.3 (Public Alpha, pre-release)
Your model, your platform, your language.
This alpha removes the "one machine, one model" limits of the first
releases. You can use a cloud model API instead of the local model, choose a
local model that fits your hardware, and run on macOS or (experimentally)
Windows. Verification now works without configuration for Python, Rust,
Java/Kotlin, C/C++, Ruby and PHP as well as Go and JavaScript/TypeScript.
Cross-service analysis now covers gRPC, protobuf, OpenAPI and SQL contracts.
The local default, Qwen3.6-35B-A3B on Linux with an NVIDIA GPU, is still the
only validated configuration. Its validation results from
v0.1.0-alpha.1 still apply. The new features are covered
by unit, integration and container tests but have not yet been validated on
real tasks. That is what this release asks you to help with.
Highlights
-
Cloud model providers (experimental): OpenAI, Anthropic, Google Gemini
or any OpenAI-compatible service, instead of the local model.bcode provider key set anthropic # prompts for the key, no echo bcode provider use anthropic --model claude-opus-5-5 bcode provider test
API keys are kept in the OS credential store (or an owner-only file),
never inconfig.yaml, on command lines or in logs. Only the host-side
gateway uses them: the agent's sandbox still has no network and never sees
a key. With a cloud provider, the agent's conversation, including
repository content, is sent to that provider. -
A model for your hardware (experimental):
bcode model recommend
rates every model profile against this machine (GPU, unified memory on
Apple Silicon, RAM, disk) and suggests one. Five new profiles (Qwen3.5-4B
and 9B, gpt-oss-20b, Devstral Small 2, Qwen3.8-27B), all Apache-2.0 and
pinned by commit and checksum. Downloads are resumable and
checksum-verified (bcode model fetch|use|remove). -
Set-up wizard (experimental):
/setupin the interface walks you
through a local model or a cloud API. For a cloud API it has a masked key
field, the provider's live model list and a connection test. The model and
provider are set up from the interface, with no config file to edit. -
macOS and Windows (experimental): binaries for macOS (Apple Silicon
and Intel) and Windows, with prebuilt checksum-verified tools and llama.cpp
(Metal on Apple Silicon), Docker Desktop resource limits, and Windows path
handling in the Linux sandbox. Install on Windows withscripts/install.ps1. -
Verification for more languages: built-in presets for Python (pytest
or unittest), Rust (Cargo), Java and Kotlin (Maven, Gradle; JDK 11, 17 or
21 chosen per project), C/C++ (CMake + CTest, Meson, Autotools, Make), Ruby
and PHP. Any other language runs its Makefile'stest/checktarget; when
nothing applies, a skippedtestsstage says so instead of passing
silently. Dependencies stay offline: the checkout's.venvorvendor/,
and this machine's Cargo, Maven and Gradle caches, read-only. Behavioural
evidence ("a changed test fails on the original code and passes with the
change") is compared per test for pytest, unittest, Cargo, Maven, Gradle,
minitest, RSpec, PHPUnit, CTest and Meson. One project per language is
verified offline in the real sandbox image by a container test. -
gRPC, protobuf, OpenAPI and SQL contracts: cross-service analysis now
links gRPC clients to servers down to the RPC (Go, Python, Java/Kotlin, C#,
Rust, Ruby, PHP, TS/JS),.protopackages to the code using them, OpenAPI
operations to routes and calls, and SQL tables to queries in other
services. When the agent changes a.proto, a spec or a migration but not
the code depending on it, it gets one round to update it. On Google's
Online Boutique demo it finds all 21 gRPC connections between its services
and nothing else. -
Clearer failures: a missing or stopped Docker, a missing sandbox image,
model weights orllama-server, and broken tools are each reported with
thesetup --only STEPthat fixes them, before a task starts. -
Fewer false ambiguity stops: an ambiguity must be supported by the
request's own text before the task stops to ask you.
Upgrade notes
- Rebuild the sandbox image:
bcode setupdetects the image built by
alpha.2 as outdated and rebuilds it (about 5 GB instead of 2 GB, for the
new language toolchains). - Re-index: run
bcode indexto fill the new cross-service contracts. - The database migrates itself (migrations 4 and 5).
- Behaviour changes:
- A changed test that does not compile or load on the original code, or a
stage that times out there, is no longer behavioural evidence. Such a
task endstests_green, nottask_verified. - Repositories in Python, Rust, Java, C/C++, Ruby or PHP now get test
stages. A project whose dependencies are not installed in the checkout
fails verification with the command to install them, where it used to
pass with no tests run. - If codebase-memory-mcp is missing or at the wrong version,
index
fails with one message namingsetup --only tools; a task run warns
and continues without graph context.
- A changed test that does not compile or load on the original code, or a
- New Go dependencies:
anthropic-sdk-go,go-keyring,godbus/dbusand
their runtime (MIT, BSD, Apache-2.0), listed inTHIRD_PARTY_NOTICES.md.
Known limitations
- Only Qwen3.6-35B-A3B on Linux with an NVIDIA GPU is validated. Cloud
providers, the other model profiles, macOS and Windows have not been run
end to end on real tasks. - Windows: 17 of 31 test packages pass on the CI runner and the full flow
has not been run on a Windows machine. macOS: 26 of 31 pass and the full
flow has not been run on a Mac. Serena is not supported on Windows. - Verification runs offline: a project's dependencies must already be
installed in the checkout or in this machine's package caches. - HTTP routes and calls, topics and environment variables are analyzed in
Go and TS/JS only; the other languages contribute gRPC, protobuf and SQL
contracts. SQL contracts are per table, not per column. - The validation limitations of
v0.1.0-alpha.1 still apply.
Please report what you find:
issues. The
verification state a task ended in (task_verified, tests_green or
failed), and whether it was right, is the most useful thing you can tell us.
Assets
boundedcode-linux-amd64,boundedcode-linux-arm64,
boundedcode-darwin-amd64,boundedcode-darwin-arm64,
boundedcode-windows-amd64.exe,boundedcode-windows-arm64.exe: static
binaries (CGO disabled).SHA256SUMS: checksums, verified by the installers.SBOM-<os>-<arch>.spdx.json: SPDX 2.3 software bill of materials of each
binary.boundedcode-v0.1.0-alpha.3-licenses.tar.gz:LICENSE,NOTICE,
THIRD_PARTY_NOTICES.mdand the license texts of every Go module compiled
into the binaries (LICENSES/go) and of the upstream components
(LICENSES/upstream).
Model weights are not distributed. Download them with bcode setup or
bcode model fetch, from their publisher at a pinned revision, and review
their license.
BoundedCode v0.1.0-alpha.2 (Public Alpha, pre-release)
Install once, then run bcode in any repository.
This alpha adds an interactive way to use BoundedCode, in the style of Claude
Code, Codex or OpenCode. It is a one-command install, a chat for the
repository you are in, and guided set-up of the local prerequisites. The
engine is the same as in
v0.1.0-alpha.1: local inference, sandboxed agent,
deterministic and behavioural verification. Its validation results are
unchanged and still apply. The interface is experimental.
Highlights
-
One-command install:
curl -fsSL https://raw.githubusercontent.com/akynte/boundedcode/main/scripts/install.sh | bashThis installs
boundedcodeand the short namebcodeinto~/.local/bin.
It uses the checksum-verified release binary attached to this release, and
otherwise builds from source with Go. -
Chat (
bcode): run it with no arguments in a git repository. The
repository is registered and indexed on first use. Each message becomes a
task that runs in an isolated worktree inside the sandbox, with its actions
and verification streamed into the conversation. Messages typed during a run
are queued, andescinterrupts. If a task stops on an ambiguous request,
your reply is the answer. After a task completes, the next message starts a
follow-up from its branch. Slash commands:/diff,/apply,/verify,
/review,/new,/open,/status,/model,/setupand more. -
Guided set-up (
bcode setup): checks the configuration, the pinned
tools (codebase-memory-mcp, gitleaks), llama.cpp, the default model weights
and the Docker sandbox image, and installs what is missing. It asks before
every download or build. The installers and the sandbox build context are
embedded in the binary, so no source checkout is needed. -
task apply: brings a completed task's changes into your checkout,
staged for review or committed with--commit. It refuses a checkout with
uncommitted changes, or changes that do not apply cleanly. -
Follow-ups:
task create --from TASKstarts a task from a previous
task's branch. -
Full-screen interface: besides the chat, there are views for tasks
(live activity, diff, verification, escalations), workspaces and indexing,
repository intelligence, the runtime and models, frontier escalations,
stats,doctorand Serena set-up, and a console for any other command.
Actions run the CLI commands in-process, so policy, sandboxing and audit
records are identical to the CLI. Frontier approvals appear as dialogs. -
Audit detail:
agent.eventrecords include a short, redacted summary
of each agent action and its stated reason, which feeds the live activity
view.
Upgrade notes
boundedcodewith no arguments in a terminal now opens the interface. In
scripts (no terminal) it prints help as before.- New Go dependencies: Charm Bubble Tea, Bubbles and Lip Gloss, and their
runtime. All are MIT or BSD-3-Clause and are listed in
THIRD_PARTY_NOTICES.md. - No configuration changes are required.
setupwrites a configuration only
when none exists.
Known limitations
- The interface and chat are experimental. They are tested with model-level
and backend tests and in a pseudo-terminal, not yet on a long real-model
session from the chat. setupbuilds llama.cpp from source, which needs git, cmake and a C++
compiler, plus the CUDA toolkit for GPU inference. To use a server you
already run, setinference.mode: external.- The model download is large (about 22 GB for the default profile).
- Linux x86-64 only.
- The validation limitations of
v0.1.0-alpha.1 still apply.
Assets
boundedcode-linux-amd64: static binary (CGO disabled).SHA256SUMS: checksums, verified by the installer.SBOM.spdx.json: SPDX 2.3 software bill of materials of the binary.boundedcode-v0.1.0-alpha.2-licenses.tar.gz:LICENSE,NOTICE,
THIRD_PARTY_NOTICES.mdand the license texts of every Go module compiled
into the binary (LICENSES/go) and of the upstream components
(LICENSES/upstream).
Model weights are not distributed. Download them with bcode setup, from
their publisher at a pinned revision, and review their license.
BoundedCode v0.1.0-alpha.1 (Public Alpha, pre-release)
Bounded context. Bounded cost. Unbounded codebases.
BoundedCode is a local-first AI software-engineering platform for working
on large repositories. It keeps model context bounded, persists task state,
verifies changes with behavioural evidence and escalates to a frontier model
only when needed (optional).
This is the first public alpha: usable, validated experimentally on a
small sample, tested on one machine with one model. Commands and
configuration may change. It is not production-ready.
Highlights
- Local inference: llama.cpp (
llama-server, supervised), with model
profiles measured on the reference machine. Validated with
Qwen3.6-35B-A3B (UD-Q4_K_M). - Long-running tasks: a persistent task ledger and audit log. A task
resumes after Ctrl-C, a crash or a reboot. - Compact context: small task-specific context packs. Retrieval seeds are
ranked so that quoted error messages lead to their origin. On a 1.8 M-token
repository the model saw 1.6 % of the source. - Repository intelligence:
- codebase-memory-mcp for breadth (graph, impact);
- optional Serena v1.7.0 (MIT, pinned) for semantic depth;
- cross-service contract analysis (HTTP, topics, env, Terraform).
- Behavioural verification: a task is
TASK_VERIFIEDonly when a test it
adds fails on the base commit and passes with the change. Otherwise it is
reported astests_green/ UNVERIFIED. - Isolation: each task works on a git worktree branch
agent/<id>.
Nothing is pushed or merged. - Sandbox: agent tools run in a network-less container, with secret
masking, protected paths, a deterministic command policy and hardened git
handling. - Runaway control:
- per-response caps on thinking and visible output;
- a progress-aware strategy budget;
- stopped strategies are recorded and not repeated.
- Task contract: required behaviour, allowed alternatives and material
ambiguity are surfaced before implementation (task run --clarify). - Optional frontier escalation: a Z1-Z4 policy via the Codex CLI with a
ChatGPT sign-in, or manual packets. API keys are refused, and packets are
sanitized.
Validation
Initial validation: 0/8, then 1/8 after the first defect fixes
(report).
Those failures drove the engineering work. Development reruns are not
counted as validation.
Second independent validation: 6 previously unseen public tasks from
SWE-bench Multilingual and Multi-SWE-bench. They were screened before
execution for consistency between each issue and its acceptance test, frozen,
then run once each:
- 5/6 strict
TASK_VERIFIEDsuccesses with hidden acceptance passing - 6/6 hidden acceptance tests passed
- all five strict successes local-only (0 frontier calls)
- 0 false verification passes; 0 human code intervention
The sixth task (Prometheus) was implemented correctly, but the evidence
checker did not link its data-driven test file to the test that reads it.
It is therefore scored as a strict failure.
Small validation sample; not a statistically comprehensive benchmark.
Full report.
Tested configuration
| Component | Tested configuration |
|---|---|
| Machine | Lenovo LOQ 15IRH8 |
| CPU / GPU | i7-13620H, RTX 4060 Laptop 8 GB |
| RAM / OS | 64 GB, Debian 13 |
| Model | Qwen3.6-35B-A3B UD-Q4_K_M |
| llama.cpp | v0.5.0 |
Not a minimum requirement.
Known limitations
- Small validation sample (8 + 6 tasks, one run each). It covers one
machine and one model. - Ambiguity detection is model-derived and has false positives. Under the
defaultaskpolicy, a false positive costs a clarification question. - Behavioural evidence misses data-driven test files consumed by a test
elsewhere. - Frontier escalation was enabled but not exercised in the second
validation. - The strategy governor has not yet stopped a live runaway; it is
calibrated on development runs.
Install
Source release. Follow the
Quick start. Model
weights are not distributed: download them from their publisher and review
their license.