Skip to content

Releases: akynte/boundedcode

BoundedCode v0.1.0-alpha.5

Pre-release

Choose a tag to compare

@github-actions github-actions released this 10 Oct 14:38

BoundedCode v0.1.0-alpha.5

Security fixes, and releases you can verify.

This alpha fixes the defects found in the public-launch
readiness audit. Two of them let an
agent's work reach past BoundedCode's policies on the host. It also moves
to Go 1.27.2, which fixes eight standard-library vulnerabilities.

It is the first release built and attested in CI. You can check that each
file was built from this tag by this repository's workflow, and you can
rebuild it byte for byte.

The feature set is that of v0.1.0-alpha.4. The local
default, Qwen3.6-35B-A3B on Linux with an NVIDIA GPU, is still the only
validated configuration.

Security

  • Nested repositories.
    • Before: host git ran commands configured inside a git repository
      that the agent created in its worktree. Once such a repository was
      recorded as a gitlink, its own filter drivers ran on the host during
      checkpoint commits.
    • Now: host git ignores submodules, and a checkpoint refuses a
      worktree that holds a nested repository.
  • Quoted file names.
    • Before: names that git quotes (non-ASCII characters, ", control
      characters) did not match the protected-path, secret-path and deny-path
      policies. A new .github/workflows/cié.yml passed the diff-scope check.
    • Now: changed-file lists are read NUL-separated and unquoted.
  • Go 1.27.2. It fixes eight standard-library vulnerabilities reachable
    from BoundedCode, in net/http, its HTTP/2 transport, crypto/tls and
    mime/multipart (GO-2026-6603 to GO-2026-6617). The release binaries are
    built with it.

Verification

  • Go behavioural evidence.
    • Before: a test that failed on the base, but was skipped, failed or
      did not run on the change, counted as evidence.
    • Now: each attributed test must also pass on the change.

Installation and first run

  • bcode setup --check exits non-zero while any step is missing.
  • Linux with an NVIDIA GPU but no CUDA toolkit: set-up downloads the
    pinned prebuilt CUDA build of llama.cpp, instead of starting a source
    build that fails.
  • Cloud provider: the task contract (and its ambiguity check) now uses
    it. It was silently skipped before.
  • The installers are more careful:
    • install.sh refuses to replace an unrelated bcode or boundedcode
      (BC_FORCE=1 overrides) and replaces the binary atomically.
      install.ps1 stages files before replacing them.
    • Both warn when another bcode comes first on PATH, fall back to the
      releases feed when the GitHub API rate limit is hit, and accept a
      mirror (BC_DOWNLOAD_BASE).
    • install.ps1 no longer closes your PowerShell window when it fails.
  • Smoke tests for the installers on Linux, macOS and Windows, and for a
    first task end to end, now run in CI.

The CHANGELOG lists every change.

Releases you can verify

  • Built in CI. A release workflow builds the release from the tag
    with make dist, after the tests. It creates a draft, which the
    maintainer reviews and publishes.
  • Attested. Every file carries a build-provenance attestation (SLSA,
    signed through Sigstore). To check one:
    gh attestation verify boundedcode-linux-amd64 -R akynte/boundedcode
  • Reproducible, SBOMs included.
    • The SBOMs now take their creation time from the commit, and each
      target has its own document namespace.
    • Rebuilding from the tag with the Go version in go.mod reproduces
      every binary and SBOM. The licenses archive also matches when rebuilt
      under umask 022, as on the CI runner: at this tag, its permission
      bits still follow the umask. Later releases normalise them.
    • The steps are in
      Verifying a release.
    • Earlier releases: v0.1.0-alpha.4's binaries and licenses archive
      were already reproducible; its SBOMs were not.

Evidence published since alpha.4

  • Comparative evaluation:
    BoundedCode vs. the same OpenHands agent, model and sandbox without it,
    on 8 pre-registered public tasks.
    • Results: hidden tests passed 7 of 8 vs. 6 of 8. No difference in
      success is shown, and BoundedCode took about twice the time and tokens.
    • What is published: the protocol, the scripts, the raw results, the
      deviations and the negative results.
  • Evidence demo: a reproducible
    demonstration of behavioural verification. It includes a recorded real
    local-model run.

Benchmarks for this release. The release checklist asks for a benchmark
re-run. The task benchmark is the comparative evaluation:

  • Same code. Its 24 agent runs and 16 gate replays used a build of
    commit 86c123d. Since then, no product code has changed except the Go
    version.
  • Speed not re-measured. bench infra was not re-run: llama.cpp and
    the model are unchanged since alpha.4.

Contributing

  • Issue forms for bugs, installation problems, platform reports, wrong
    verification results, feature proposals, model reports and evaluation
    results.
  • Contributor onboarding, with first
    issues taken from observed defects.
  • Submitting an evaluation result.
  • CI now lists skipped tests, and shows macOS and Windows test failures
    instead of hiding them.

Upgrade notes

  • Building from source needs Go 1.27.2 or later (go.mod).
  • Tasks with nested repositories. Some tasks now fail at their
    checkpoint commit with "nested git repository … refusing to run host git
    on it", naming the directory:
    • Which tasks: those where a git repository appears inside the
      worktree, for example one the agent created or cloned, or a submodule
      it checked out.
    • Unaffected: task worktrees start without submodules checked out.
  • setup --check in scripts. Scripts that run bcode setup --check
    must accept a non-zero exit while set-up is incomplete.

Known limitations

  1. macOS and Windows remain experimental. Their unit tests do not pass:
    5 packages fail on macOS and 14 on Windows. Four of the causes are
    filed as good first issues.
  2. Earlier limitations. The limitations of
    v0.1.0-alpha.4 and
    v0.1.0-alpha.3 still apply.

Please report what you find, using the
issue forms.

Assets

  • Binaries: boundedcode-linux-amd64, boundedcode-linux-arm64,
    boundedcode-darwin-amd64, boundedcode-darwin-arm64,
    boundedcode-windows-amd64.exe and boundedcode-windows-arm64.exe.
    They are static (CGO disabled).
  • SHA256SUMS: checksums, verified by the installers.
  • SBOM-<os>-<arch>.spdx.json: an SPDX 2.3 software bill of materials
    for each binary.
  • boundedcode-v0.1.0-alpha.5-licenses.tar.gz: LICENSE, NOTICE,
    THIRD_PARTY_NOTICES.md, and the license texts of every Go module
    compiled into the binaries (LICENSES/go) and of the upstream
    components (LICENSES/upstream).
  • Attestations: a build-provenance attestation for each of the files
    above, stored with GitHub.

Model weights are not distributed. Download them with bcode setup or
bcode model fetch, from their publisher at a pinned revision, and review
their license.

BoundedCode v0.1.0-alpha.4

Pre-release

Choose a tag to compare

@akynte akynte released this 08 Oct 16:14

BoundedCode v0.1.0-alpha.4

Do the repositories still work together?

A multi-repository task can pass every repository's own checks and still
break the contract between them. For example, a provider renames a protobuf
field and updates its server, while a client in another repository still
builds against its vendored copy of the generated code. This alpha adds a
cross-repository compatibility gate (experimental). For every gRPC,
protobuf or OpenAPI link that a task's change affects, it reports
compatible, broken or untested, and it records the evidence: the
commands run, their outcomes, and the repositories and commits involved.

Everything else is as in v0.1.0-alpha.3. The local
default, Qwen3.6-35B-A3B on Linux with an NVIDIA GPU, is still the only
validated configuration. The gate is covered by unit, integration and CLI
tests on fixtures, and was run with the real Docker sandbox. It has not
yet been run on real tasks
.

Highlights

  • Only real changes count. The gate compares the change's base and head
    commits:

    • .proto messages, fields, enums and RPC signatures are compared
      structurally, and breaking changes are recognized: a field removed,
      renumbered, retyped or renamed, or an RPC signature changed.
    • OpenAPI operations are compared together with the schemas they
      reference.
    • Go code is compared by the function that holds the call.
    • Comments, other RPCs and unrelated functions do not count.
  • Checked with the repositories' own tests, against each other's new
    commits:

    • gRPC and protobuf (Go): the dependent repository's go test stage
      runs in the sandbox, built against the provider's candidate commit
      through a generated Go workspace. A vendored copy or a replace
      directive does not apply, and coverage must show the dependent's code
      ran.
    • OpenAPI: the dependent's tests must fail when the operation is
      removed from the spec, which shows they actually read it.
    • Failures: control runs at the base commits decide whether a failure
      comes from the change or was already there.
  • In the task loop:

    • A broken link fails the attempt, and the agent is told which link
      broke and why.
    • An untested link withholds TASK_VERIFIED. When a test could settle
      it, the agent is asked once for one.
    • Results survive a crash or task resume. A new commit marks them stale.
  • Where to see it: task status, verify --full (also --json) and the
    Verification tab of the interface show the per-link report:

    cross-repository compatibility: BROKEN (3 broken, 2 untested, 0 compatible)
      BROKEN     grpc_def     grpc shop.payments.v1.PaymentService/Charge [changed]
                 checkout@a5f2bdbb99 internal/pay/client.go:23 -> protos@f70b3b206e payments/v1/payments.proto:9
                 checkout's checks fail with protos's candidate (checkout@a5f2bdbb99 + protos@f70b3b206e) and pass with protos's base commit: internal/pay/client.go:23:77: unknown field AmountCents in struct literal of type paymentsv1.ChargeRequest
    

    This example is the new
    contract-break fixture:
    a field rename that all three repositories' own tests pass. See the
    design and its limits.

  • Vendored Go modules verify offline. Verification no longer forces
    -mod=mod on a module with vendor/modules.txt. That flag made Go ignore
    the vendored copy and try to download it. A vendored module now builds
    from its tracked vendor/.

Documentation

The README and docs at this tag were written before two clarifications
that are now on main:

  • In a multi-repository task, TASK_VERIFIED also requires every affected
    gRPC, protobuf or OpenAPI link to be compatible. See
    Verification and
    the "How it works" diagram, which now shows the gate.
  • The sandbox controls
    have an entry for the gate and describe vendored Go modules, and
    CONTRIBUTING.md lists internal/compat as security-sensitive.

The released binaries are unchanged; only the documentation was updated.

Upgrade notes

  • The database migrates itself (migration 6).
  • Behaviour change: a multi-repository task whose change affects a
    gRPC, protobuf or OpenAPI link now ends task_verified only if every
    such link is shown compatible. Otherwise it ends tests_green, with the
    reason recorded as a decision. To keep the previous behaviour, set
    repointel.compat_gate: false.
  • No new dependencies, and the sandbox image is unchanged.

Known limitations

  1. The gate checks gRPC and protobuf sides written in Go only: one module
    at the repository root, a go test stage, and generated code committed
    in a task repository (the gate never runs protoc). OpenAPI sides are
    checked only when a test reads the spec. Every other case is reported
    untested, with the reason.
  2. compatible for a client and server means that both compile against the
    same new definition and their own tests execute their side. The client
    is never run against the real server.
  3. A wire-incompatible definition change stays untested even when every
    task repository agrees. Services already deployed are not tested.
  4. Test code is written by the agent and runs inside the check, the same
    trust boundary as for verification. Likewise, the agent can edit a
    vendored module's vendor/ like any other source; only review catches
    that.
  5. The limitations of v0.1.0-alpha.3
    still apply.

Please report what you find:
issues. A link reported
broken that was fine, or compatible when it was not, is especially
useful.

Assets

  • boundedcode-linux-amd64, boundedcode-linux-arm64,
    boundedcode-darwin-amd64, boundedcode-darwin-arm64,
    boundedcode-windows-amd64.exe, boundedcode-windows-arm64.exe: static
    binaries (CGO disabled).
  • SHA256SUMS: checksums, verified by the installers.
  • SBOM-<os>-<arch>.spdx.json: SPDX 2.3 software bill of materials of each
    binary.
  • boundedcode-v0.1.0-alpha.4-licenses.tar.gz: LICENSE, NOTICE,
    THIRD_PARTY_NOTICES.md and the license texts of every Go module compiled
    into the binaries (LICENSES/go) and of the upstream components
    (LICENSES/upstream).

Model weights are not distributed. Download them with bcode setup or
bcode model fetch, from their publisher at a pinned revision, and review
their license.

BoundedCode v0.1.0-alpha.3 (Public Alpha, pre-release)

Choose a tag to compare

@akynte akynte released this 08 Oct 12:09

Your model, your platform, your language.

This alpha removes the "one machine, one model" limits of the first
releases. You can use a cloud model API instead of the local model, choose a
local model that fits your hardware, and run on macOS or (experimentally)
Windows. Verification now works without configuration for Python, Rust,
Java/Kotlin, C/C++, Ruby and PHP as well as Go and JavaScript/TypeScript.
Cross-service analysis now covers gRPC, protobuf, OpenAPI and SQL contracts.

The local default, Qwen3.6-35B-A3B on Linux with an NVIDIA GPU, is still the
only validated configuration. Its validation results from
v0.1.0-alpha.1 still apply. The new features are covered
by unit, integration and container tests but have not yet been validated on
real tasks
. That is what this release asks you to help with.

Highlights

  • Cloud model providers (experimental): OpenAI, Anthropic, Google Gemini
    or any OpenAI-compatible service, instead of the local model.

    bcode provider key set anthropic      # prompts for the key, no echo
    bcode provider use anthropic --model claude-opus-5-5
    bcode provider test

    API keys are kept in the OS credential store (or an owner-only file),
    never in config.yaml, on command lines or in logs. Only the host-side
    gateway uses them: the agent's sandbox still has no network and never sees
    a key. With a cloud provider, the agent's conversation, including
    repository content, is sent to that provider.

  • A model for your hardware (experimental): bcode model recommend
    rates every model profile against this machine (GPU, unified memory on
    Apple Silicon, RAM, disk) and suggests one. Five new profiles (Qwen3.5-4B
    and 9B, gpt-oss-20b, Devstral Small 2, Qwen3.8-27B), all Apache-2.0 and
    pinned by commit and checksum. Downloads are resumable and
    checksum-verified (bcode model fetch|use|remove).

  • Set-up wizard (experimental): /setup in the interface walks you
    through a local model or a cloud API. For a cloud API it has a masked key
    field, the provider's live model list and a connection test. The model and
    provider are set up from the interface, with no config file to edit.

  • macOS and Windows (experimental): binaries for macOS (Apple Silicon
    and Intel) and Windows, with prebuilt checksum-verified tools and llama.cpp
    (Metal on Apple Silicon), Docker Desktop resource limits, and Windows path
    handling in the Linux sandbox. Install on Windows with scripts/install.ps1.

  • Verification for more languages: built-in presets for Python (pytest
    or unittest), Rust (Cargo), Java and Kotlin (Maven, Gradle; JDK 11, 17 or
    21 chosen per project), C/C++ (CMake + CTest, Meson, Autotools, Make), Ruby
    and PHP. Any other language runs its Makefile's test/check target; when
    nothing applies, a skipped tests stage says so instead of passing
    silently. Dependencies stay offline: the checkout's .venv or vendor/,
    and this machine's Cargo, Maven and Gradle caches, read-only. Behavioural
    evidence ("a changed test fails on the original code and passes with the
    change") is compared per test for pytest, unittest, Cargo, Maven, Gradle,
    minitest, RSpec, PHPUnit, CTest and Meson. One project per language is
    verified offline in the real sandbox image by a container test.

  • gRPC, protobuf, OpenAPI and SQL contracts: cross-service analysis now
    links gRPC clients to servers down to the RPC (Go, Python, Java/Kotlin, C#,
    Rust, Ruby, PHP, TS/JS), .proto packages to the code using them, OpenAPI
    operations to routes and calls, and SQL tables to queries in other
    services. When the agent changes a .proto, a spec or a migration but not
    the code depending on it, it gets one round to update it. On Google's
    Online Boutique demo it finds all 21 gRPC connections between its services
    and nothing else.

  • Clearer failures: a missing or stopped Docker, a missing sandbox image,
    model weights or llama-server, and broken tools are each reported with
    the setup --only STEP that fixes them, before a task starts.

  • Fewer false ambiguity stops: an ambiguity must be supported by the
    request's own text before the task stops to ask you.

Upgrade notes

  • Rebuild the sandbox image: bcode setup detects the image built by
    alpha.2 as outdated and rebuilds it (about 5 GB instead of 2 GB, for the
    new language toolchains).
  • Re-index: run bcode index to fill the new cross-service contracts.
  • The database migrates itself (migrations 4 and 5).
  • Behaviour changes:
    • A changed test that does not compile or load on the original code, or a
      stage that times out there, is no longer behavioural evidence. Such a
      task ends tests_green, not task_verified.
    • Repositories in Python, Rust, Java, C/C++, Ruby or PHP now get test
      stages. A project whose dependencies are not installed in the checkout
      fails verification with the command to install them, where it used to
      pass with no tests run.
    • If codebase-memory-mcp is missing or at the wrong version, index
      fails with one message naming setup --only tools; a task run warns
      and continues without graph context.
  • New Go dependencies: anthropic-sdk-go, go-keyring, godbus/dbus and
    their runtime (MIT, BSD, Apache-2.0), listed in THIRD_PARTY_NOTICES.md.

Known limitations

  1. Only Qwen3.6-35B-A3B on Linux with an NVIDIA GPU is validated. Cloud
    providers, the other model profiles, macOS and Windows have not been run
    end to end on real tasks.
  2. Windows: 17 of 31 test packages pass on the CI runner and the full flow
    has not been run on a Windows machine. macOS: 26 of 31 pass and the full
    flow has not been run on a Mac. Serena is not supported on Windows.
  3. Verification runs offline: a project's dependencies must already be
    installed in the checkout or in this machine's package caches.
  4. HTTP routes and calls, topics and environment variables are analyzed in
    Go and TS/JS only; the other languages contribute gRPC, protobuf and SQL
    contracts. SQL contracts are per table, not per column.
  5. The validation limitations of
    v0.1.0-alpha.1 still apply.

Please report what you find:
issues. The
verification state a task ended in (task_verified, tests_green or
failed), and whether it was right, is the most useful thing you can tell us.

Assets

  • boundedcode-linux-amd64, boundedcode-linux-arm64,
    boundedcode-darwin-amd64, boundedcode-darwin-arm64,
    boundedcode-windows-amd64.exe, boundedcode-windows-arm64.exe: static
    binaries (CGO disabled).
  • SHA256SUMS: checksums, verified by the installers.
  • SBOM-<os>-<arch>.spdx.json: SPDX 2.3 software bill of materials of each
    binary.
  • boundedcode-v0.1.0-alpha.3-licenses.tar.gz: LICENSE, NOTICE,
    THIRD_PARTY_NOTICES.md and the license texts of every Go module compiled
    into the binaries (LICENSES/go) and of the upstream components
    (LICENSES/upstream).

Model weights are not distributed. Download them with bcode setup or
bcode model fetch, from their publisher at a pinned revision, and review
their license.

BoundedCode v0.1.0-alpha.2 (Public Alpha, pre-release)

Choose a tag to compare

@akynte akynte released this 07 Oct 01:56

Install once, then run bcode in any repository.

This alpha adds an interactive way to use BoundedCode, in the style of Claude
Code, Codex or OpenCode. It is a one-command install, a chat for the
repository you are in, and guided set-up of the local prerequisites. The
engine is the same as in
v0.1.0-alpha.1: local inference, sandboxed agent,
deterministic and behavioural verification. Its validation results are
unchanged and still apply. The interface is experimental.

Highlights

  • One-command install:

    curl -fsSL https://raw.githubusercontent.com/akynte/boundedcode/main/scripts/install.sh | bash

    This installs boundedcode and the short name bcode into ~/.local/bin.
    It uses the checksum-verified release binary attached to this release, and
    otherwise builds from source with Go.

  • Chat (bcode): run it with no arguments in a git repository. The
    repository is registered and indexed on first use. Each message becomes a
    task that runs in an isolated worktree inside the sandbox, with its actions
    and verification streamed into the conversation. Messages typed during a run
    are queued, and esc interrupts. If a task stops on an ambiguous request,
    your reply is the answer. After a task completes, the next message starts a
    follow-up from its branch. Slash commands: /diff, /apply, /verify,
    /review, /new, /open, /status, /model, /setup and more.

  • Guided set-up (bcode setup): checks the configuration, the pinned
    tools (codebase-memory-mcp, gitleaks), llama.cpp, the default model weights
    and the Docker sandbox image, and installs what is missing. It asks before
    every download or build. The installers and the sandbox build context are
    embedded in the binary, so no source checkout is needed.

  • task apply: brings a completed task's changes into your checkout,
    staged for review or committed with --commit. It refuses a checkout with
    uncommitted changes, or changes that do not apply cleanly.

  • Follow-ups: task create --from TASK starts a task from a previous
    task's branch.

  • Full-screen interface: besides the chat, there are views for tasks
    (live activity, diff, verification, escalations), workspaces and indexing,
    repository intelligence, the runtime and models, frontier escalations,
    stats, doctor and Serena set-up, and a console for any other command.
    Actions run the CLI commands in-process, so policy, sandboxing and audit
    records are identical to the CLI. Frontier approvals appear as dialogs.

  • Audit detail: agent.event records include a short, redacted summary
    of each agent action and its stated reason, which feeds the live activity
    view.

Upgrade notes

  • boundedcode with no arguments in a terminal now opens the interface. In
    scripts (no terminal) it prints help as before.
  • New Go dependencies: Charm Bubble Tea, Bubbles and Lip Gloss, and their
    runtime. All are MIT or BSD-3-Clause and are listed in
    THIRD_PARTY_NOTICES.md.
  • No configuration changes are required. setup writes a configuration only
    when none exists.

Known limitations

  1. The interface and chat are experimental. They are tested with model-level
    and backend tests and in a pseudo-terminal, not yet on a long real-model
    session from the chat.
  2. setup builds llama.cpp from source, which needs git, cmake and a C++
    compiler, plus the CUDA toolkit for GPU inference. To use a server you
    already run, set inference.mode: external.
  3. The model download is large (about 22 GB for the default profile).
  4. Linux x86-64 only.
  5. The validation limitations of
    v0.1.0-alpha.1 still apply.

Assets

  • boundedcode-linux-amd64: static binary (CGO disabled).
  • SHA256SUMS: checksums, verified by the installer.
  • SBOM.spdx.json: SPDX 2.3 software bill of materials of the binary.
  • boundedcode-v0.1.0-alpha.2-licenses.tar.gz: LICENSE, NOTICE,
    THIRD_PARTY_NOTICES.md and the license texts of every Go module compiled
    into the binary (LICENSES/go) and of the upstream components
    (LICENSES/upstream).

Model weights are not distributed. Download them with bcode setup, from
their publisher at a pinned revision, and review their license.

BoundedCode v0.1.0-alpha.1 (Public Alpha, pre-release)

Choose a tag to compare

@akynte akynte released this 06 Oct 00:00

Bounded context. Bounded cost. Unbounded codebases.

BoundedCode is a local-first AI software-engineering platform for working
on large repositories. It keeps model context bounded, persists task state,
verifies changes with behavioural evidence and escalates to a frontier model
only when needed (optional).

This is the first public alpha: usable, validated experimentally on a
small sample, tested on one machine with one model. Commands and
configuration may change. It is not production-ready.

Highlights

  • Local inference: llama.cpp (llama-server, supervised), with model
    profiles measured on the reference machine. Validated with
    Qwen3.6-35B-A3B (UD-Q4_K_M).
  • Long-running tasks: a persistent task ledger and audit log. A task
    resumes after Ctrl-C, a crash or a reboot.
  • Compact context: small task-specific context packs. Retrieval seeds are
    ranked so that quoted error messages lead to their origin. On a 1.8 M-token
    repository the model saw 1.6 % of the source.
  • Repository intelligence:
    • codebase-memory-mcp for breadth (graph, impact);
    • optional Serena v1.7.0 (MIT, pinned) for semantic depth;
    • cross-service contract analysis (HTTP, topics, env, Terraform).
  • Behavioural verification: a task is TASK_VERIFIED only when a test it
    adds fails on the base commit and passes with the change. Otherwise it is
    reported as tests_green / UNVERIFIED.
  • Isolation: each task works on a git worktree branch agent/<id>.
    Nothing is pushed or merged.
  • Sandbox: agent tools run in a network-less container, with secret
    masking, protected paths, a deterministic command policy and hardened git
    handling.
  • Runaway control:
    • per-response caps on thinking and visible output;
    • a progress-aware strategy budget;
    • stopped strategies are recorded and not repeated.
  • Task contract: required behaviour, allowed alternatives and material
    ambiguity are surfaced before implementation (task run --clarify).
  • Optional frontier escalation: a Z1-Z4 policy via the Codex CLI with a
    ChatGPT sign-in, or manual packets. API keys are refused, and packets are
    sanitized.

Validation

Initial validation: 0/8, then 1/8 after the first defect fixes
(report).
Those failures drove the engineering work. Development reruns are not
counted as validation.

Second independent validation: 6 previously unseen public tasks from
SWE-bench Multilingual and Multi-SWE-bench. They were screened before
execution for consistency between each issue and its acceptance test, frozen,
then run once each:

  • 5/6 strict TASK_VERIFIED successes with hidden acceptance passing
  • 6/6 hidden acceptance tests passed
  • all five strict successes local-only (0 frontier calls)
  • 0 false verification passes; 0 human code intervention

The sixth task (Prometheus) was implemented correctly, but the evidence
checker did not link its data-driven test file to the test that reads it.
It is therefore scored as a strict failure.

Small validation sample; not a statistically comprehensive benchmark.
Full report.

Tested configuration

Component Tested configuration
Machine Lenovo LOQ 15IRH8
CPU / GPU i7-13620H, RTX 4060 Laptop 8 GB
RAM / OS 64 GB, Debian 13
Model Qwen3.6-35B-A3B UD-Q4_K_M
llama.cpp v0.5.0

Not a minimum requirement.

Known limitations

  1. Small validation sample (8 + 6 tasks, one run each). It covers one
    machine and one model.
  2. Ambiguity detection is model-derived and has false positives. Under the
    default ask policy, a false positive costs a clarification question.
  3. Behavioural evidence misses data-driven test files consumed by a test
    elsewhere.
  4. Frontier escalation was enabled but not exercised in the second
    validation.
  5. The strategy governor has not yet stopped a live runaway; it is
    calibrated on development runs.

Install

Source release. Follow the
Quick start. Model
weights are not distributed: download them from their publisher and review
their license.