Skip to content

Releases: inferstep/ATLAS

V3.1.5 "Maia"

Choose a tag to compare

@itigges22 itigges22 released this 29 Sep 23:56
v3.1.5

A security release. 3.1.5 fixes a way to skip ATLAS's workspace check with an argument name in another letter case: GHSA-c3p6-m657-h629. Please upgrade.

Security

  • Tool argument names must now be lowercase, as every tool's names are. A name in another letter case (for example "PATH") or with non-ASCII letters is refused. Before, such a name skipped the workspace check, and in 3.1.4 a "Command" argument also skipped the destructive-command deny-list.
  • insert_after's path is now checked at dispatch, like every other write.

Upgrade

  • Existing install: in your ATLAS folder, run git pull first, then atlas upgrade. Or re-run the install command.
  • New install: curl -fsSL https://raw.githubusercontent.com/inferstep/ATLAS/main/scripts/atlas-bootstrap.sh | bash

Verify

  • The tag v3.1.5 is SSH-signed by a key in .github/allowed_signers (in a checkout: git -c gpg.ssh.allowedSignersFile=.github/allowed_signers verify-tag v3.1.5).
  • The images are signed by the inferstep/ATLAS build workflow:
    cosign verify --certificate-identity https://github.com/inferstep/ATLAS/.github/workflows/build-images.yml@refs/tags/v3.1.5 \
      --certificate-oidc-issuer https://token.actions.githubusercontent.com ghcr.io/inferstep/atlas-proxy:3.1.5
    

To hear about releases only, choose Watch → Custom → Releases on the repository page.

Full list: CHANGELOG.md.

V3.1.4 "Maia"

Choose a tag to compare

@itigges22 itigges22 released this 27 Sep 22:32
v3.1.4

The first release from ATLAS's new home, inferstep/ATLAS. It carries the
security fixes below, the move to the new image owner, and the new
contributor setup, on top of the changes on main since 3.1.3 (see the end of these notes).

Security: commands ran without approval, and file search read credential files

  • run_background started any command without the approval prompt in the
    default and accept-edits modes, and was checked against a narrower
    deny-list than run_command (env rm -rf / and (rm -rf /) passed it).
    Outside yolo, every tool that runs a command now asks, and one command
    policy covers both. It looks where a command can start: behind env,
    nohup, nice, time, timeout or exec, in a subshell and in a
    command substitution. grep mkfs notes.txt is no longer refused.
  • search_files returned the contents of credential files that
    read_file refuses (.env, keys, cloud credentials) and followed
    symlinks out of the workspace. It now skips both and reports how many
    credential files it skipped (skipped_credential_files). move_file
    refuses to move a credential file to another name, and insert_after
    gets the write deny-list. The rules follow what a tool does, and a test
    fails when a new tool is outside them. Shell commands are not covered,
    and the docs now say so.
  • Approval prompts cut a command at 100 characters, so the end of a chain
    was never shown. The proxy sends the whole command, and
    stop_background names its job.
  • In the TUI, one "allow for session" answer on a deletion approved every
    later deletion without showing which file, and the proxy honoured
    delete_file in session_allowed_tools from any client. Each deletion
    is now asked about on its own: the TUI never auto-answers or sends a
    session approval for delete_file, the proxy ignores one, and a new
    session starts with no approvals. The TUI prompt shows the whole
    command, wrapped; one too long for the screen keeps its first and last
    lines in view and says how many are not shown.
  • The TUI's chat stream, events stream, raw demo lane and feedback calls
    never sent the service token, so on an install with one they failed
    with 401. Every request to the proxy now sends it, ahead of an api-keys
    token.

Moved to inferstep/ATLAS

  • The repository is now github.com/inferstep/ATLAS. Old links, git
    remotes and the install one-liner redirect.
  • Images are published under ghcr.io/inferstep/atlas-* and signed by
    the inferstep/ATLAS build workflow. ghcr.io/itigges22/atlas-* stays
    published for existing installs but gets no new versions.
  • Existing installs: re-run the install command, or git pull and then
    atlas upgrade. atlas upgrade, atlas config migrate and a bootstrap
    re-run move ATLAS_GHCR_OWNER=itigges22 in .env to inferstep. A
    failed upgrade or a rollback puts it back. An install pinned to a
    release from before the move keeps the old owner until it upgrades, and
    an owner set in the shell is left alone.

Fixed

Contributors

  • New issue forms for bugs, features, tasks, docs, spikes and RFCs, and a
    fuller pull request template. CONTRIBUTING is
    rewritten as the path from an issue to a release. GOVERNANCE
    describes the trust ladder and the RFC flow. New
    TRIAGE and INCIDENT_RESPONSE
    guides.
  • The public Roadmap board
    has a Start Here view. The atlas-bot handles /claim and /unclaim,
    reminds and releases stale claims, adds area labels and welcomes
    newcomers.
  • Pull requests now also need the dependency review and a conventional
    title check. An OpenSSF Scorecard runs weekly. Dependabot targets dev.
  • Releases record a deployment per promotion, and publishing :latest or
    a version tag waits for the release owner's approval.

Docs

  • The V3.0 LiveCodeBench figure (74.6%) is withdrawn. The benchmark runner
    never ran LiveCodeBench's hidden tests (see the notice in
    V3_ABLATION_STUDY). The README says
    ATLAS has no current benchmark result.

Upgrade

  • New install: curl -fsSL https://raw.githubusercontent.com/inferstep/ATLAS/main/scripts/atlas-bootstrap.sh | bash
  • Existing install: in your ATLAS folder, run git pull first, then atlas upgrade. Or re-run the install command. Your .env moves from ghcr.io/itigges22 to ghcr.io/inferstep on its own. Start the TUI with atlas tui, which rebuilds it from the new source.
  • Warning: on 3.1.3, atlas upgrade without git pull stays on the old ghcr.io/itigges22 images, which have none of these fixes, and can still say it succeeded.

Verify

  • The tag v3.1.4 is SSH-signed by a key in .github/allowed_signers (in a checkout: git -c gpg.ssh.allowedSignersFile=.github/allowed_signers verify-tag v3.1.4).
  • The images are signed with the inferstep/ATLAS build workflow's identity:
    cosign verify --certificate-identity https://github.com/inferstep/ATLAS/.github/workflows/build-images.yml@refs/tags/v3.1.4 \
      --certificate-oidc-issuer https://token.actions.githubusercontent.com ghcr.io/inferstep/atlas-proxy:3.1.4
    

This release also contains the changes on main since 3.1.3: see CHANGELOG.md (from "Measured reliability" down).

V3.1.3 "Maia"

Choose a tag to compare

@itigges22 itigges22 released this 06 Jul 18:58
v3.1.3

The production-platform release: the machinery around running, upgrading, and trusting an ATLAS install.

Highlights

  • Staged upgrade & rollback — atlas upgrade records a restore point (tag + image digests + .env backup), verifies cosign signatures on the target images, waits for readiness, runs a smoke check, and automatically restores the previous release if any step fails. atlas rollback returns to the restore point on demand; atlas diagnostics collect produces a shareable, secret-filtered support bundle.
  • SQLite state store (ADR 0007, core implementation from #128 by @HarshalPatel1972) — the lens's pattern cache, co-occurrence graph, router posteriors, task queue, and metrics now live in one WAL-mode SQLite file on the lens-state volume. The Redis service and its configuration are gone; one less dependency, and state backup is a single file.
  • Signed artifact manifests — atlas artifact verify | snapshot | rollback: SSH-signed provenance manifests over lens/ASA bundles with per-file SHA-256 verification, plus one-generation bundle rollback.
  • Observability — structured JSON logs (ATLAS_LOG_FORMAT=json) with X-ATLAS-Request-ID correlation across proxy → llama/v3/lens/sandbox; a stable error-code taxonomy and OpenAPI 3.1 spec for the proxy surface; command-execution trust modes (ATLAS_TRUST_MODE).
  • Interactive permissions & sessions — destructive tool calls prompt for approval in default/accept-edits modes; the TUI saves sessions and atlas --continue / --resume picks them back up.
  • Typed configuration — atlas config validate | migrate: schema-checked .env with forward migration (deprecated keys, including the removed ATLAS_REDIS_*, are cleaned up automatically).
  • Install trust — release-pinned installs: ATLAS_BOOTSTRAP_REF=v3.1.3 pins the checkout to this signed tag and the images to the matching cosign-signed digests. A review-before-running variant is documented beside the one-shot install.
  • Two adversarial review sweeps fixed 33 confirmed bugs across the proxy, CLI, services, and CI — including a trust-mode bypass, several restore-path failures, and CI gates that could never fail.

Upgrading

atlas upgrade --to v3.1.3

Existing installs: the Redis container and redis-data volume are no longer used after this upgrade (docker volume rm <project>_redis-data reclaims the space). Learned lens state starts fresh in SQLite and rebuilds through use.

Full details in CHANGELOG.md.

V3.1.2 — Maia

Choose a tag to compare

@itigges22 itigges22 released this 17 Jun 16:43

ATLAS TUI in action
The ATLAS TUI live, 10× sped up, running the V3 pipeline on a file creation.

V3.1.2 stays in the Maia line. It broadens hardware reach (run ATLAS on AMD, Apple Silicon, Intel, or CPU — not just NVIDIA), closes the bring-your-own-model gap (train Lens + ASA artifacts for any GGUF, including from your own usage), and lands a focused agent-reliability pass that fixes the failure modes that used to silently break a turn on a local 9B. Less new surface, much higher hit-rate per turn.

What shipped

Hardware reach

  • AMD ROCm via llama.cpp — including RDNA4 / RX 9070 (gfx1200/gfx1201) (#26)
  • Apple Silicon — native macOS hybrid Metal path (native llama-server for inference, Docker for the rest) (#32)
  • Vulkan universal fallback — one image covering AMD / Intel / Snapdragon / Apple-via-MoltenVK / CPU (#114)

Bring-your-own-model

  • atlas lens build / atlas lens retrain — local Lens training pipeline; train C(x)+G(x) for a non-default GGUF from bench candidates or from your own rated sessions (#100)
  • atlas asa check / build / publish — per-model ASA control-vector calibration parity (#113)
  • Per-model lens thresholds — off-rails / regression cutoffs now ship with the lens artifact (gx_thresholds.json), auto-calibrated per model, so the lens fires correctly regardless of a model's score scale

In-the-loop lens training

  • Rate a pass in the TUI — /good · /bad (pass-level) and /review + /deny · /accept (per-file) — and the proxy banks weighted, labeled samples. A "retrain available" banner surfaces when enough accumulate, then atlas lens retrain learns your actual workloads.

Agent reliability

  • Tool-result visibility — Gemma's chat template silently dropped role:"tool" messages, so the model never saw tool output and looped; results now render as user turns (model-agnostic). This was the big one.
  • read-dedup false-negative fix; traceback → directed edit (#39 option 3); move_file tool; pip-install and case-mismatch steers; per-turn max_tokens + content-loop bounds
  • Sandbox: shell gate narrowed to catastrophic-only (the container is the real boundary); host-sized cgroup limits (mem + pids); interactive V3 wall-clock cap

Structural reasoning & docs

  • Call-graph reasoning (#39 / #125) — intra-file call edges on outline_file/read_file, a Datalog solver, ast_edit friendly selectors (now incl. <script>/<style>), cyclomatic-complexity tiering
  • ARCHITECTURE.md translated to zh-CN / ja / ko (#25)

Full line-by-line: CHANGELOG.md.

Closed in this release

  • #26 ROCm · #114 Vulkan · #65 GPU vendor lock-in
  • #100 atlas lens build · #113 atlas asa check/build/publish
  • #25 ARCHITECTURE.md translations

Substantially advanced

  • #39 Structural code reasoning — call graph + Datalog solver + ast_edit shipped; the solver-backed reachability tail continues into V3.2

Thank you

Contributor: @yogthos (Dmitri Sotnikov) — #125 structural call-graph reasoning (the #39 foundation), with #124 RPG planning in review for V3.2.

Sponsor on GitHub keeps this sustainable.

Next — V3.2

Architecture-first RPG planning (#120 / #124); solver-backed reachability + wavelet retrieval (#39); reasoning-with-sampling (#9); automated HF submission (#102); ROCm-on-K3s; formal 9B benchmarks (#28).


Maia — first of the Pleiades, Atlas's eldest daughter. The V3.1.x line stays in the Maia series; V3.2 opens the next.

V3.1.0 — Maia

Choose a tag to compare

@itigges22 itigges22 released this 12 May 06:48

ATLAS TUI in action
The ATLAS TUI live, 10× sped up, running the V3 pipeline on a file creation.

Discussion thread — install reports, TUI feedback, hardware requests, bugs welcome.

V3.1.0 ships the TUI (you can see the pipeline now), a one-command install that works on a fresh VM, and a layer of guardrails on top of the V3 stack that catches the failure modes that used to bork an agent turn: tool misselection, lazy stubs, repeated reasoning, sandbox-passing stubs. Less new behavior, more reliability per turn.

What shipped

  • Bubbletea TUI (PC-062) - native Go terminal UI as the canonical chat client; /cancel endpoint; slash commands; mode-aware input
  • One-command bootstrap - curl … | bash tested fresh on Ubuntu 22.04 / 24.04 / Rocky 9; atlas init wizard + atlas doctor + atlas tier
  • Aider removed - ~2000 lines of format-translation glue gone; /v1/chat/completions is now a passthrough; agent loop is on /v1/agent
  • Streaming lens (PC-206 + PC-207) - the Geometric Lens now scores per-token during generation, not post-hoc on the finished candidate. gx_min < 0.05 mid-stream fires a corrective immediately instead of waiting for the full candidate to land. The lens also vetoes a sandbox-passing candidate when its gx_min indicates a stub or placeholder collapse
  • AST-aware surgical edits (#39 v1 + 4 points) - ast_edit tool, structural-verification veto, cyclomatic-complexity tier classification, call-chain repair context, reachability-slice auto-injection
  • BiasBusters - tool descriptions + conditional GBNF + per-step filter + ASA activation steering (always-on llama.cpp control vector built from 1000 contrast-pair prompts)
  • Anti-laziness gates (PC-194–201) - rejects stub writes, completion-claim verification, "stops at the easy fix" detector, run_background tool
  • Plan mode - /v3/plan with adherence scoring; plan-progress reminder per step; TUI renders plan events live
  • Agent loop hardening - 8 detectors (reasoning-repetition, path-aware error breaker, done-without-action gate, truncation recovery, …); absoluteMaxTurns ceiling removed
  • Sandbox + execution stack - run_command is sandbox-only now (PC-188); language-agnostic layout detection (PC-191/192/193)
  • K3s deployment restored - full template set rebuilt after the May 2 repo restructure broke it silently
  • Multilingual docs - README, SETUP, TROUBLESHOOTING in zh-CN / ja / ko
  • CI - ruff, CodeQL, cross-distro install matrix (Ubuntu 22.04 / 24.04 / Rocky 9)

Full line-by-line: CHANGELOG.md.

Closed in this release

  • #108 atlas-proxy public endpoints docs
  • #24 CONTRIBUTING.md
  • #31 CLI reliability L6
  • #22 dead Fox code paths

Substantially advanced

  • #39 Structural code reasoning - v1 + four points shipped; solver-backed layer remains
  • #28 Formal 9B benchmarks - 9B is the canonical model now; formal pass lands V3.1.1

Thank you

Sponsor: @monkeini (Robin Harrison) - ATLAS's first sponsor. Sponsor on GitHub keeps this sustainable.

Inspiration: @yogthos (Dmitri Sotnikov) - chiasmus shaped the #39 work.

Contributors:

V3.1.1

  • macOS and Windows installers
  • AMD ROCm + Apple Metal
  • Formal 9B benchmarks (LiveCodeBench, GPQA Diamond, SciCode)

Maia — first of the Pleiades, Atlas's eldest daughter. Future V3.1.x releases continue the series: Electra, Taygete, Alcyone, Celaeno, Sterope, Merope.

V3.0.1 — Interactive CLI + Documentation Overhaul

Choose a tag to compare

@itigges22 itigges22 released this 06 Apr 04:51

ATLAS Banner

V3.0.1 ships ATLAS as an interactive coding assistant you can download and run today. Type atlas in any project directory and start building — powered by a local 9B model on your own GPU. No API keys, no cloud, no data leaves the machine.

ATLAS CLI


What's New

Interactive CLI with Tool-Call Agent Loop

The biggest change in V3.0.1: ATLAS is no longer just a benchmark runner. It's a full interactive coding assistant.

  • atlas command — starts all services and drops you into an Aider-powered coding session
  • Grammar-constrained agent loop — the model emits structured JSON tool calls (write_file, edit_file, run_command, etc.) with llama-server's response_format:json_object guaranteeing 100% valid output
  • 8 tools: read_file, write_file, edit_file, delete_file, run_command, search_files, list_directory, plan_tasks
  • Per-file V3 routing — config files and short files write instantly (T1), feature files with complex logic automatically route through the full V3 pipeline (T2) for diverse candidate generation, build verification, and energy-based selection
  • Real-time streaming — every tool call, V3 pipeline stage, and build verification visible in the terminal as it happens
  • 95.8% reliability across 8 difficulty levels (24 test iterations)

Docker Compose Deployment

Five commands to a working system:

git clone https://github.com/itigges22/ATLAS.git && cd ATLAS
wget https://huggingface.co/unsloth/Qwen3.5-9B-GGUF/resolve/main/Qwen3.5-9B-Q6_K.gguf -O models/Qwen3.5-9B-Q6_K.gguf
pip install -e .
cp .env.example .env
docker compose up -d
atlas

5 containerized services: llama-server (CUDA), geometric-lens (C(x)/G(x) scoring), v3-service (pipeline), sandbox (8-language code execution), atlas-proxy (agent loop). All orchestrated with health checks and dependency ordering.

Documentation Overhaul

Every documentation file rewritten from scratch and verified against source code:

  • ARCHITECTURE.md — 13 Mermaid diagrams including sequence diagrams showing actual HTTP calls between services
  • API.md — every endpoint across all 5 services with verified request/response formats
  • CLI.md — streaming output format, workflow examples, troubleshooting guide
  • CONFIGURATION.md — every environment variable verified against source code
  • MAP.md — every file in the repo with clickable links and descriptions
  • SETUP.md — Docker Compose, bare metal, and K3s deployment guides
  • TROUBLESHOOTING.md — 20+ issue scenarios with verified fixes

Bug Fixes

  • Geometric Lens Dockerfile port mismatch — container listened on 8001 but docker-compose expected 8099. Fresh Docker Compose deploys had a broken Lens service. Fixed.
  • Python CLI default RAG port — atlas/cli/client.py defaulted to port 31144 (K3s) instead of 8099 (Docker Compose). Fixed.
  • Missing Aider config files — .aider.model.settings.yml and .aider.model.metadata.json were not in the repo. The atlas launcher would fail without them. Restored.
  • hostname -I fails on Arch Linux (#6) — replaced with portable fallback chain: ip addr -> hostname -I -> hostname -i
  • rag-api/models/ does not exist (#10) — resolved by V3.0.1 restructuring (rag-api/ -> geometric-lens/)
  • No Lens weight documentation (#11) — added training docs to SETUP.md + HuggingFace dataset link
  • docker image exists not a real command (#12, PR #13 by @g0dnerd) — fixed to docker image inspect

V3 Pipeline (unchanged from V3.0)

The same pipeline that scored 74.6% LiveCodeBench pass@1-v(k=3) on frozen Qwen3-14B is now integrated into the interactive CLI:

  • Phase 0: Probe with progressive retry (light -> standard -> /nothink)
  • Phase 1: PlanSearch (3 plans) + DivSampling (12 perturbations) + Budget Forcing (5 tiers)
  • Phase 2: Build verification + C(x)/G(x) scoring + S* tiebreaking
  • Phase 3: PR-CoT repair + Refinement Loop + Derivation Chains

Hardware Requirements

Resource Minimum
GPU VRAM 16 GB (NVIDIA, CUDA)
System RAM 14 GB
Disk 20 GB
OS Linux (RHEL, Ubuntu, Arch, Debian)

Tested on RTX 5060 Ti 16GB. See SETUP.md for detailed instructions.


Benchmark Results (V3.0, Qwen3-14B)

Benchmark Score Tasks
LiveCodeBench v5 74.6% pass@1-v(k=3) 599
GPQA Diamond 47.0% 198
SciCode 14.7% (sub-problems) 341

The CLI currently runs Qwen3.5-9B with the same V3 pipeline. Formal benchmarks on the 9B model are V3.1 work.

Full ablation data: v3_ablation_results/ | Traces: HuggingFace


Contributors

Thanks to @g0dnerd for PR #13 fixing the Docker image existence check.

Thanks to @aaronetz (#6), @nguyenhoangthuan99 (#10), and @namp (#11) for reporting issues that made ATLAS better.