Skip to content

Repository files navigation

taracode

DevOps & Cloud SI Assistant
SI, super intelligence, is the new name for AI.
Expert guidance for Kubernetes, Terraform, Docker, and multi-cloud deployments.
Runs locally with Ollama for complete privacy.

Release License: MIT Sponsor tara-vision CI Coverage Go Version

Stars Forks Issues

Quick Start Features Commands Documentation Contributing


Why taracode?

  • Investigate-first - Read-only by default, and never prompts you in that mode
  • Policy-gated operate mode - Mutations pass through protected targets, deny patterns and required dry runs before a remembered permission or a prompt decides
  • Fifteen classified tools - Every call is classified read or mutate from its arguments, not its name
  • Redaction and an audit log - Secrets are stripped from tool output before anything sees it; every mutation is recorded
  • A public scoreboard - Which local models can actually do DevOps work, measured on taracode's own loop, with speed and memory per model
  • Privacy-first - Runs fully local with Ollama, your data never leaves your machine
  • No Account Required - Open source, just install and use

The scoreboard

Which local model can actually do DevOps work? 18 open models, 33 offline tasks each, on one NVIDIA RTX 5090 (32 GB) with Ollama 0.35.0, measured on taracode's own loop. The top five:

# Model Pass rate Tokens/s VRAM
1 gemma4:31b 94% 64 23.1 GiB
2 qwen3.8:27b 91% 117 20.1 GiB
3 glm-4.7-flash 88% 212 19.9 GiB
4 qwen3.6:27b 88% 123 20.0 GiB
5 gemma4:12b 88% 138 9.2 GiB

One run per row, ranked by pass rate, then mean score. The same 18 models were run twice that day, and a model moved by up to four tasks between the two runs (one task is about 3 points of pass rate), so rows within four tasks of each other are a tie: all five above and the next three on the full board. VRAM is measured on the GPU with the model loaded at a 32,768-token context window, on this 32 GB card with nothing else on it; a smaller card was not tested. The full board, with speed, memory and the method, is on code.tara.vision/evals and in docs/evals/scoreboard.md.

Quick Start

1. Install Ollama

# macOS / Linux
brew install ollama

# Or download from https://ollama.ai

2. Pull a Model

ollama pull gemma4:12b     # 16 GB of RAM
ollama pull glm-4.7-flash  # 32 GB of RAM (needs Ollama 0.14.3 or newer)
ollama pull qwen3.6:35b    # 48 GB of RAM and up

Any model that Ollama tags with the tools capability works. On Ollama, taracode refuses a model without that capability at startup or on /model instead of falling back to JSON-in-content tool calls (run taracode doctor to see which installed models qualify). taracode requests its context window itself on every turn (context.window, default auto = 32,768 tokens, or the model's native maximum when that is smaller) instead of relying on the server's default; OLLAMA_CONTEXT_LENGTH only matters for servers taracode does not control, such as vLLM and llama.cpp.

3. Install taracode

Quick install (recommended):

curl -fsSL https://code.tara.vision/install.sh | bash

Homebrew (macOS / Linux):

brew install --cask tara-vision/taracode/taracode

Upgrading from a version installed as a formula? Run brew uninstall taracode once, then the command above.

Go install:

go install github.com/tara-vision/taracode@latest

Manual download:

Download binaries from GitHub Releases. Every release ships checksums, a cosign signature and SLSA provenance; see SECURITY.md to verify a download.

4. Run

cd your-project
taracode

taracode starts in investigate mode right away, no /init needed: read-only tools, nothing saved. Run /init when you want sessions, memory, history and operate mode (it also writes a starter .taracode/policy.yaml). Start asking questions about your infrastructure.

Context window and thinking

Three config.yaml keys control how taracode talks to the model:

context:
  window: auto    # "auto" requests 32768 tokens, or less on a smaller model; set a token count to go higher
think: auto        # auto, off, on, low, medium, or high
keep_alive: ""     # how long Ollama keeps the model loaded; "" = server default, "-1" = keep loaded

context.window: auto (the default) requests 32,768 tokens, or the model's native maximum when that is smaller; it never asks for more than 32,768 tokens on its own, which keeps the KV cache affordable on 16 GB and 32 GB machines. Set a number instead to request an explicit window, clamped to the model's native maximum, on a model that supports going higher. A non-numeric value falls back to auto with a warning. A session warns when the resulting window is below 16,384 tokens, since tool-heavy sessions compact early at that size.

think sets the reasoning mode sent with requests. Change it without restarting taracode with /think (/think alone shows the current mode, /think low changes it).

Run taracode doctor (or /doctor inside a session) to check the server, the installed models and their capabilities, your machine's RAM tier and the registry's recommended model for it, and the external CLIs taracode's tools shell out to.

context.window and keep_alive are controlled on the native Ollama client; vLLM and llama.cpp keep the OpenAI-compatible path, where only think low|medium|high reaches the server (as reasoning_effort).

Features

Modes and policy

taracode starts in investigate mode: only tools with a read form are exposed, and nothing ever prompts you. operate mode exposes every tool; each mutation goes through the policy, in order:

  1. Protected targets - kube contexts, namespaces, cloud accounts, paths and hosts named in the policy are a hard deny, with the reason printed. A kubectl or helm mutation run through shell carries its context and namespace like the dedicated tools (the current ones of the kubeconfig it names when it names none, HELM_NAMESPACE and HELM_KUBECONTEXT for helm), also inside loops, conditionals and subshells. A mutation of every namespace (-A), of several contexts or namespaces, or of one taracode cannot determine before it runs counts as touching the protected ones: a value the shell computes, an option repeated with two values, sudo or xargs in front, a kubeconfig that is not a small regular file, and a context switched earlier on the line (kubectl config use-context, kubectx, kubens, a KUBECONFIG assignment, a sourced file that can set KUBECONFIG, a script). kubectl or helm run from another program's string or file, such as sh -c "kubectl ..." or a script, is not seen; nor is one inside a double-quoted substitution ("$(kubectl ...)" or backticks in double quotes), though the unquoted $(kubectl ...) is. A namespace a kubectl mutation changes as an object (kubectl delete ns kube-system, kubectl label ns/kube-system ...) is the namespace it touches, and on a shell line such a command with a glob or a brace list in any word, even a quoted JSON value (kubectl patch ns shop -p '{"a":1,"b":2}'), counts as touching every namespace; the kubectl tool, which runs no shell, reads the same words literally. Protected paths cover the file write_file or edit_file changes, the directory the terraform tool runs in, and in a shell command the targets of its redirects and the files it hands to a file-writing program (tee, sed -i, cp, mv, rm, touch, chmod, ln, dd, sort -o, curl -o, ...), also after a literal cd. A path a command builds at run time (a variable, a command substitution) is not seen, which is why the built-in deny patterns also refuse any mutation that names .taracode/policy.yaml.
  2. Deny patterns - command globs that are refused outright.
  3. Required dry runs - kubectl apply shows a server-side diff first, terraform apply requires a plan produced in this session and shows its summary, helm upgrade runs --dry-run first (a release with a --post-renderer is refused: the dry run would run the renderer before you approve).
  4. Permission - the remembered allow/ask/deny rule for the tool, or a prompt.
> /mode investigate|operate   # show or switch mode (or --mode at startup)

A call to a tool the session does not offer (one the model repeats from a resumed session, or makes up) is refused. offline hides the two web tools, but it does not reach into shell: curl, wget -O-, dig, nslookup, host and ping still count as reads there.

The policy comes from .taracode/policy.yaml merged over ~/.taracode/policy.yaml (lists unioned, booleans take the stricter value); with neither file, a built-in policy identical to the one below applies. /init writes this starter:

# taracode policy (see: taracode doctor, /policy show).
# The model never sees this file. .taracode/policy.yaml merges over ~/.taracode/policy.yaml:
# lists are unioned and booleans take the stricter value. With no policy file at all, taracode
# uses a built-in policy identical to this one. Patterns are globs (* ? and ** in paths).
version: 1
mode: investigate               # the mode a session starts in: investigate or operate
protected:                      # never mutated in operate mode (hard deny, printed reason)
  kube_contexts: ["*prod*", "*production*"]
  kube_namespaces: ["kube-system"]
  cloud_accounts: []            # AWS account ids, Azure subscription ids, GCP project ids, or *globs*
  paths: ["**/*.tfstate", ".git/**"]
  hosts: []
deny:                           # refused outright; the last pattern keeps the policy files safe
  commands: ["rm -rf /*", "kubectl delete namespace *", "terraform destroy*", "*.taracode/policy.yaml*"]
require_dry_run:                # shown before the permission prompt
  kubectl_apply: true           # kubectl diff first
  terraform_apply: true         # a plan from this session, its summary first
  helm_upgrade: true            # helm --dry-run first (upgrade and install)
redact:
  enabled: true                 # secrets in tool output become [redacted:<kind>]
  extra_patterns: []            # additional Go regular expressions
mcp:
  trust_read_only_hint: true    # a server's readOnlyHint gives its tool a read form (investigate mode)
  read_only: {}                 # per server, the tools (globs) that count as reads when the hint is not trusted

/policy show prints the effective policy and where it came from. /permissions manages the remembered per-tool rules (/permissions allow|deny|ask <tool|all>, /permissions reset). Every mutation, allowed or denied, is appended to .taracode/audit.jsonl before it runs; /audit, /audit all and /audit export json read it.

Redaction runs on every tool result before the model, the session or the history sees it. The live output of a shell command is redacted a line at a time as it reaches the screen, so a secret that spans lines, such as a PEM private key block, is redacted in the tool result but not in the live view (no_stream_commands: true in config.yaml turns the live view off).

Tools

Fifteen tools replace the old 58; every call is classified read or mutate from its arguments, not from the tool's name. Investigate mode exposes the tools that have a read form (thirteen, eleven when offline is set).

Tool Arguments (summary) Read when Mutate when
read_file path, start_line, end_line always never
list_files path, glob, recursive, max always never
search_files pattern, path, glob, max always never
write_file path, content never always
edit_file path, old, new, preview never always
shell command, timeout command matches the read-only allowlist (cat, ls, grep, find, ps, df, du, date, curl GET, dig, nslookup, jq, git read verbs, kubectl read verbs, terraform read verbs, ...) otherwise
git args status, diff, log, show, branch (list), blame add, commit, stash, checkout, reset, push, merge, rebase
kubectl verb, resource, name, namespace, context, args, output get, describe, logs, events, top, explain, api-resources, version, diff, dry-run apply, delete, patch, edit, scale, rollout, exec, cp, drain, cordon
helm args list, status, get, history, show, template, lint, diff install, upgrade, rollback, uninstall
terraform command, dir, args init (with -backend=false), validate, fmt -check, plan (always -json, post-processed), show, state list, output, graph apply, destroy, import, taint, state mv/rm/push, workspace delete
docker args ps, images, logs, inspect, stats, compose ps/config/logs build, run, rm, rmi, exec, compose up/down/restart, push
cloud provider (aws, az, gcloud), args verbs describe, get, list, ls, show otherwise
scan scanner (trivy, gitleaks, tfsec, kubesec, dependency), target, severity always never
web_search query, max always (external, disabled by offline) never
web_fetch url always (external, disabled by offline) never

MCP tools join the same registry. With mcp.trust_read_only_hint: true (the default) a server that annotates a tool readOnlyHint: true gives it a read form; every other MCP tool is a mutation and stays hidden in investigate mode. Set it to false for a server you do not trust and list the tools that may read under mcp.read_only, per server, by name or glob. Protected targets and deny patterns do not apply to MCP tools: the per-tool permission is their only gate.

Project Memory

Remember project-specific knowledge across sessions:

> /remember We use PostgreSQL for production databases
> /remember Always run tests before pushing #workflow
> /memory search database

Commands

Command Description
/init Initialize the project (creates TARACODE.md and .taracode/)
/reload Reload project context from TARACODE.md
/status Show project and session status
/session [new [name]|load <id>|delete <id>|rename <id> <name>] Show or manage the current session
/sessions List all sessions
/clear Clear the conversation (new session)
/model Switch between available models
/think [auto|off|on|low|medium|high] Show or set the reasoning mode
/mode [investigate|operate] Show or switch the operating mode
/permissions [allow|deny|ask <tool|all>|reset] Remembered answers for mutations
/audit [all|export json|clear] Mutations recorded in this project
/policy show Effective policy and where it comes from
/plan Show the active plan
/context Context window budget breakdown
/compact Force conversation compaction
/stats Session statistics
/usage Token usage for this session
/history [n|all] File operation history
/undo [n|--dry-run] Undo file modifications
/diff [export] Show or export session changes
/remember <text> [#tag] Save a memory about this project
/memory [search <q>|delete <id>|export|import <file>|stats|cleanup|clear] Project memories
/mcp [connect|disconnect <name>|tools] MCP servers and their tools
/tools List available tools
/upgrade [check|now|skip|changelog|status] Check for and install updates
/doctor Diagnose the LLM server and tools
/help Show this help

Configuration

Create ~/.taracode/config.yaml:

# The Ollama host (--host and TARACODE_HOST override it)
host: http://localhost:11434

# Generation options for the main chat
generation:
  temperature: 0.7     # Sampling randomness (0.0-2.0)
  top_p: 0.9           # Nucleus sampling threshold (0.0-1.0)
  num_predict: 0       # Max tokens per response (0 = model default)

# Starting mode: investigate or operate (unset starts in investigate, or the mode a policy file names)
mode: investigate

# Security scanning
scan:
  default_severity: ""   # e.g. "HIGH,CRITICAL"

# Search
search:
  primary: duckduckgo
  fallback: searxng
  brave_api_key: ""    # Optional: Brave Search API

# Memory
memory:
  enabled: true
  auto_capture: true

See config.example.yaml for all options.

Supported LLM Backends

Backend Setup Notes
Ollama brew install ollama Recommended, easiest setup
vLLM Self-hosted For production deployments
llama.cpp Self-hosted Lightweight option

Roadmap

Phase 2 of the v3 plan (investigate and operate modes, the classified tools, the policy engine, redaction and the audit log) shipped in 3.0.0-alpha.2. Phase 3 shipped the evals and the first scoreboard in 3.0.0-beta.1, and 3.0.0 made the line stable. 3.1.0 settled the tool set at fifteen and the connection at one Ollama host. 3.2.0 added speed, memory and a board per machine to the evals. Runbooks, the MCP server and the skills pack are next, as 3.3 and 3.4. See ROADMAP.md.

Evals and the scoreboard

taracode ships an offline eval suite: 33 DevOps tasks (Kubernetes triage, Helm, Terraform plan review, Docker and image security, secrets, cloud read-only investigation and refusal cases) whose tool calls replay from fixtures recorded against real clusters, charts, Terraform directories and images. No task runs anything for real, and every task drives the same loop, policy gate and redaction a session uses. Each task scores 0.4 for the tool calls it expected, 0.5 for the answer and 0.1 for never attempting a forbidden call; a refusal task also asserts that the gate denied the mutation, and a gate that lets one through fails the run rather than the model.

The published scoreboard is in docs/evals/scoreboard.md (and on code.tara.vision/evals), regenerated with every release. Every run also records tokens per second, --hardware names the machine so the report can rank models per machine, and --gpu-probe records the GPU memory the machine itself reports (the evals reference shows the command for an NVIDIA card). Run the suite against your own Ollama:

taracode eval run --host http://localhost:11434 --model gemma4:12b --hardware "your GPU or machine"
taracode eval report

Results land in docs/evals/results/, transcripts in evals/runs/. Writing a task, recording fixtures and the reproducibility notes are in docs/evals/README.md.

Development

make deps           # Install dependencies
make build          # Build binary
make test           # Run tests
make coverage-gate   # Check per-package coverage floors
make install         # Install to /usr/local/bin

See CONTRIBUTING.md for development guidelines.

Contributing

Contributions are welcome! Please read our Contributing Guide and Code of Conduct.

Security

For security issues, please see our Security Policy.

Sponsoring

taracode is free, MIT-licensed and built without telemetry or a cloud service. If it saves you time, you can sponsor Tara Vision on GitHub to keep the releases, the model registry and the eval scoreboard coming.

License

MIT License - see LICENSE for details.


Built with ❤️ by Tara Vision · Created by Dejan Stefanoski

About

DevOps & Cloud SI Assistant that runs locally with Ollama for complete privacy.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

51 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages