Repository navigation
BoundedCode v0.1.0-alpha.1 (Public Alpha, pre-release)
Pre-releaseBounded context. Bounded cost. Unbounded codebases.
BoundedCode is a local-first AI software-engineering platform for working
on large repositories. It keeps model context bounded, persists task state,
verifies changes with behavioural evidence and escalates to a frontier model
only when needed (optional).
This is the first public alpha: usable, validated experimentally on a
small sample, tested on one machine with one model. Commands and
configuration may change. It is not production-ready.
Highlights
- Local inference: llama.cpp (
llama-server, supervised), with model
profiles measured on the reference machine. Validated with
Qwen3.6-35B-A3B (UD-Q4_K_M). - Long-running tasks: a persistent task ledger and audit log. A task
resumes after Ctrl-C, a crash or a reboot. - Compact context: small task-specific context packs. Retrieval seeds are
ranked so that quoted error messages lead to their origin. On a 1.8 M-token
repository the model saw 1.6 % of the source. - Repository intelligence:
- codebase-memory-mcp for breadth (graph, impact);
- optional Serena v1.7.0 (MIT, pinned) for semantic depth;
- cross-service contract analysis (HTTP, topics, env, Terraform).
- Behavioural verification: a task is
TASK_VERIFIEDonly when a test it
adds fails on the base commit and passes with the change. Otherwise it is
reported astests_green/ UNVERIFIED. - Isolation: each task works on a git worktree branch
agent/<id>.
Nothing is pushed or merged. - Sandbox: agent tools run in a network-less container, with secret
masking, protected paths, a deterministic command policy and hardened git
handling. - Runaway control:
- per-response caps on thinking and visible output;
- a progress-aware strategy budget;
- stopped strategies are recorded and not repeated.
- Task contract: required behaviour, allowed alternatives and material
ambiguity are surfaced before implementation (task run --clarify). - Optional frontier escalation: a Z1-Z4 policy via the Codex CLI with a
ChatGPT sign-in, or manual packets. API keys are refused, and packets are
sanitized.
Validation
Initial validation: 0/8, then 1/8 after the first defect fixes
(report).
Those failures drove the engineering work. Development reruns are not
counted as validation.
Second independent validation: 6 previously unseen public tasks from
SWE-bench Multilingual and Multi-SWE-bench. They were screened before
execution for consistency between each issue and its acceptance test, frozen,
then run once each:
- 5/6 strict
TASK_VERIFIEDsuccesses with hidden acceptance passing - 6/6 hidden acceptance tests passed
- all five strict successes local-only (0 frontier calls)
- 0 false verification passes; 0 human code intervention
The sixth task (Prometheus) was implemented correctly, but the evidence
checker did not link its data-driven test file to the test that reads it.
It is therefore scored as a strict failure.
Small validation sample; not a statistically comprehensive benchmark.
Full report.
Tested configuration
| Component | Tested configuration |
|---|---|
| Machine | Lenovo LOQ 15IRH8 |
| CPU / GPU | i7-13620H, RTX 4060 Laptop 8 GB |
| RAM / OS | 64 GB, Debian 13 |
| Model | Qwen3.6-35B-A3B UD-Q4_K_M |
| llama.cpp | v0.5.0 |
Not a minimum requirement.
Known limitations
- Small validation sample (8 + 6 tasks, one run each). It covers one
machine and one model. - Ambiguity detection is model-derived and has false positives. Under the
defaultaskpolicy, a false positive costs a clarification question. - Behavioural evidence misses data-driven test files consumed by a test
elsewhere. - Frontier escalation was enabled but not exercised in the second
validation. - The strategy governor has not yet stopped a live runaway; it is
calibrated on development runs.
Install
Source release. Follow the
Quick start. Model
weights are not distributed: download them from their publisher and review
their license.