DevOps & Cloud SI Assistant
SI, super intelligence, is the new name for AI.
Expert guidance for Kubernetes, Terraform, Docker, and multi-cloud deployments.
Runs locally with Ollama for complete privacy.
- Investigate-first - Read-only by default, and never prompts you in that mode
- Policy-gated operate mode - Mutations pass through protected targets, deny patterns and required dry runs before a remembered permission or a prompt decides
- Fifteen classified tools - Every call is classified read or mutate from its arguments, not its name
- Redaction and an audit log - Secrets are stripped from tool output before anything sees it; every mutation is recorded
- A public scoreboard - Which local models can actually do DevOps work, measured on taracode's own loop, with speed and memory per model
- Privacy-first - Runs fully local with Ollama, your data never leaves your machine
- No Account Required - Open source, just install and use
Which local model can actually do DevOps work? 18 open models, 33 offline tasks each, on one NVIDIA RTX 5090 (32 GB) with Ollama 0.35.0, measured on taracode's own loop. The top five:
| # | Model | Pass rate | Tokens/s | VRAM |
|---|---|---|---|---|
| 1 | gemma4:31b |
94% | 64 | 23.1 GiB |
| 2 | qwen3.8:27b |
91% | 117 | 20.1 GiB |
| 3 | glm-4.7-flash |
88% | 212 | 19.9 GiB |
| 4 | qwen3.6:27b |
88% | 123 | 20.0 GiB |
| 5 | gemma4:12b |
88% | 138 | 9.2 GiB |
One run per row, ranked by pass rate, then mean score. The same 18 models were run twice that day, and a model moved by up to four tasks between the two runs (one task is about 3 points of pass rate), so rows within four tasks of each other are a tie: all five above and the next three on the full board. VRAM is measured on the GPU with the model loaded at a 32,768-token context window, on this 32 GB card with nothing else on it; a smaller card was not tested. The full board, with speed, memory and the method, is on code.tara.vision/evals and in docs/evals/scoreboard.md.
# macOS / Linux
brew install ollama
# Or download from https://ollama.aiollama pull gemma4:12b # 16 GB of RAM
ollama pull glm-4.7-flash # 32 GB of RAM (needs Ollama 0.14.3 or newer)
ollama pull qwen3.6:35b # 48 GB of RAM and upAny model that Ollama tags with the tools capability works. On Ollama, taracode refuses a model without that
capability at startup or on /model instead of falling back to JSON-in-content tool calls (run taracode doctor
to see which installed models qualify). taracode requests its context window itself on every turn (context.window,
default auto = 32,768 tokens, or the model's native maximum when that is smaller) instead of relying on the
server's default; OLLAMA_CONTEXT_LENGTH only matters for servers taracode does not control, such as vLLM and
llama.cpp.
Quick install (recommended):
curl -fsSL https://code.tara.vision/install.sh | bashHomebrew (macOS / Linux):
brew install --cask tara-vision/taracode/taracodeUpgrading from a version installed as a formula? Run brew uninstall taracode once, then the command above.
Go install:
go install github.com/tara-vision/taracode@latestManual download:
Download binaries from GitHub Releases. Every release ships checksums, a cosign signature and SLSA provenance; see SECURITY.md to verify a download.
cd your-project
taracodetaracode starts in investigate mode right away, no /init needed: read-only tools, nothing saved. Run
/init when you want sessions, memory, history and operate mode (it also writes a starter
.taracode/policy.yaml). Start asking questions about your infrastructure.
Three config.yaml keys control how taracode talks to the model:
context:
window: auto # "auto" requests 32768 tokens, or less on a smaller model; set a token count to go higher
think: auto # auto, off, on, low, medium, or high
keep_alive: "" # how long Ollama keeps the model loaded; "" = server default, "-1" = keep loadedcontext.window: auto (the default) requests 32,768 tokens, or the model's native maximum when that is
smaller; it never asks for more than 32,768 tokens on its own, which keeps the KV cache affordable on 16 GB
and 32 GB machines. Set a number instead to request an explicit window, clamped to the model's native
maximum, on a model that supports going higher. A non-numeric value falls back to auto with a warning. A
session warns when the resulting window is below 16,384 tokens, since tool-heavy sessions compact early at
that size.
think sets the reasoning mode sent with requests. Change it without restarting taracode with /think
(/think alone shows the current mode, /think low changes it).
Run taracode doctor (or /doctor inside a session) to check the server, the installed models and their
capabilities, your machine's RAM tier and the registry's recommended model for it, and the external CLIs
taracode's tools shell out to.
context.window and keep_alive are controlled on the native Ollama client; vLLM and llama.cpp keep the
OpenAI-compatible path, where only think low|medium|high reaches the server (as reasoning_effort).
taracode starts in investigate mode: only tools with a read form are exposed, and nothing ever prompts you. operate mode exposes every tool; each mutation goes through the policy, in order:
- Protected targets - kube contexts, namespaces, cloud accounts, paths and hosts named in the policy are
a hard deny, with the reason printed. A
kubectlorhelmmutation run throughshellcarries its context and namespace like the dedicated tools (the current ones of the kubeconfig it names when it names none,HELM_NAMESPACEandHELM_KUBECONTEXTfor helm), also inside loops, conditionals and subshells. A mutation of every namespace (-A), of several contexts or namespaces, or of one taracode cannot determine before it runs counts as touching the protected ones: a value the shell computes, an option repeated with two values,sudoorxargsin front, a kubeconfig that is not a small regular file, and a context switched earlier on the line (kubectl config use-context,kubectx,kubens, aKUBECONFIGassignment, a sourced file that can setKUBECONFIG, a script). kubectl or helm run from another program's string or file, such assh -c "kubectl ..."or a script, is not seen; nor is one inside a double-quoted substitution ("$(kubectl ...)"or backticks in double quotes), though the unquoted$(kubectl ...)is. A namespace akubectlmutation changes as an object (kubectl delete ns kube-system,kubectl label ns/kube-system ...) is the namespace it touches, and on ashellline such a command with a glob or a brace list in any word, even a quoted JSON value (kubectl patch ns shop -p '{"a":1,"b":2}'), counts as touching every namespace; thekubectltool, which runs no shell, reads the same words literally. Protected paths cover the filewrite_fileoredit_filechanges, the directory theterraformtool runs in, and in ashellcommand the targets of its redirects and the files it hands to a file-writing program (tee,sed -i,cp,mv,rm,touch,chmod,ln,dd,sort -o,curl -o, ...), also after a literalcd. A path a command builds at run time (a variable, a command substitution) is not seen, which is why the built-in deny patterns also refuse any mutation that names.taracode/policy.yaml. - Deny patterns - command globs that are refused outright.
- Required dry runs -
kubectl applyshows a server-side diff first,terraform applyrequires a plan produced in this session and shows its summary,helm upgraderuns--dry-runfirst (a release with a--post-rendereris refused: the dry run would run the renderer before you approve). - Permission - the remembered allow/ask/deny rule for the tool, or a prompt.
> /mode investigate|operate # show or switch mode (or --mode at startup)A call to a tool the session does not offer (one the model repeats from a resumed session, or makes up) is
refused. offline hides the two web tools, but it does not reach into shell: curl, wget -O-, dig,
nslookup, host and ping still count as reads there.
The policy comes from .taracode/policy.yaml merged over ~/.taracode/policy.yaml (lists unioned, booleans
take the stricter value); with neither file, a built-in policy identical to the one below applies. /init
writes this starter:
# taracode policy (see: taracode doctor, /policy show).
# The model never sees this file. .taracode/policy.yaml merges over ~/.taracode/policy.yaml:
# lists are unioned and booleans take the stricter value. With no policy file at all, taracode
# uses a built-in policy identical to this one. Patterns are globs (* ? and ** in paths).
version: 1
mode: investigate # the mode a session starts in: investigate or operate
protected: # never mutated in operate mode (hard deny, printed reason)
kube_contexts: ["*prod*", "*production*"]
kube_namespaces: ["kube-system"]
cloud_accounts: [] # AWS account ids, Azure subscription ids, GCP project ids, or *globs*
paths: ["**/*.tfstate", ".git/**"]
hosts: []
deny: # refused outright; the last pattern keeps the policy files safe
commands: ["rm -rf /*", "kubectl delete namespace *", "terraform destroy*", "*.taracode/policy.yaml*"]
require_dry_run: # shown before the permission prompt
kubectl_apply: true # kubectl diff first
terraform_apply: true # a plan from this session, its summary first
helm_upgrade: true # helm --dry-run first (upgrade and install)
redact:
enabled: true # secrets in tool output become [redacted:<kind>]
extra_patterns: [] # additional Go regular expressions
mcp:
trust_read_only_hint: true # a server's readOnlyHint gives its tool a read form (investigate mode)
read_only: {} # per server, the tools (globs) that count as reads when the hint is not trusted/policy show prints the effective policy and where it came from. /permissions manages the remembered
per-tool rules (/permissions allow|deny|ask <tool|all>, /permissions reset). Every mutation, allowed or
denied, is appended to .taracode/audit.jsonl before it runs; /audit, /audit all and /audit export json
read it.
Redaction runs on every tool result before the model, the session or the history sees it. The live output of
a shell command is redacted a line at a time as it reaches the screen, so a secret that spans lines, such
as a PEM private key block, is redacted in the tool result but not in the live view
(no_stream_commands: true in config.yaml turns the live view off).
Fifteen tools replace the old 58; every call is classified read or mutate from its arguments, not from the
tool's name. Investigate mode exposes the tools that have a read form (thirteen, eleven when offline is set).
| Tool | Arguments (summary) | Read when | Mutate when |
|---|---|---|---|
read_file |
path, start_line, end_line | always | never |
list_files |
path, glob, recursive, max | always | never |
search_files |
pattern, path, glob, max | always | never |
write_file |
path, content | never | always |
edit_file |
path, old, new, preview | never | always |
shell |
command, timeout | command matches the read-only allowlist (cat, ls, grep, find, ps, df, du, date, curl GET, dig, nslookup, jq, git read verbs, kubectl read verbs, terraform read verbs, ...) | otherwise |
git |
args | status, diff, log, show, branch (list), blame | add, commit, stash, checkout, reset, push, merge, rebase |
kubectl |
verb, resource, name, namespace, context, args, output | get, describe, logs, events, top, explain, api-resources, version, diff, dry-run | apply, delete, patch, edit, scale, rollout, exec, cp, drain, cordon |
helm |
args | list, status, get, history, show, template, lint, diff | install, upgrade, rollback, uninstall |
terraform |
command, dir, args | init (with -backend=false), validate, fmt -check, plan (always -json, post-processed), show, state list, output, graph | apply, destroy, import, taint, state mv/rm/push, workspace delete |
docker |
args | ps, images, logs, inspect, stats, compose ps/config/logs | build, run, rm, rmi, exec, compose up/down/restart, push |
cloud |
provider (aws, az, gcloud), args | verbs describe, get, list, ls, show | otherwise |
scan |
scanner (trivy, gitleaks, tfsec, kubesec, dependency), target, severity | always | never |
web_search |
query, max | always (external, disabled by offline) |
never |
web_fetch |
url | always (external, disabled by offline) |
never |
MCP tools join the same registry. With mcp.trust_read_only_hint: true (the default) a server that
annotates a tool readOnlyHint: true gives it a read form; every other MCP tool is a mutation and
stays hidden in investigate mode. Set it to false for a server you do not trust and list the tools
that may read under mcp.read_only, per server, by name or glob. Protected targets and deny patterns
do not apply to MCP tools: the per-tool permission is their only gate.
Remember project-specific knowledge across sessions:
> /remember We use PostgreSQL for production databases
> /remember Always run tests before pushing #workflow
> /memory search database| Command | Description |
|---|---|
/init |
Initialize the project (creates TARACODE.md and .taracode/) |
/reload |
Reload project context from TARACODE.md |
/status |
Show project and session status |
/session [new [name]|load <id>|delete <id>|rename <id> <name>] |
Show or manage the current session |
/sessions |
List all sessions |
/clear |
Clear the conversation (new session) |
/model |
Switch between available models |
/think [auto|off|on|low|medium|high] |
Show or set the reasoning mode |
/mode [investigate|operate] |
Show or switch the operating mode |
/permissions [allow|deny|ask <tool|all>|reset] |
Remembered answers for mutations |
/audit [all|export json|clear] |
Mutations recorded in this project |
/policy show |
Effective policy and where it comes from |
/plan |
Show the active plan |
/context |
Context window budget breakdown |
/compact |
Force conversation compaction |
/stats |
Session statistics |
/usage |
Token usage for this session |
/history [n|all] |
File operation history |
/undo [n|--dry-run] |
Undo file modifications |
/diff [export] |
Show or export session changes |
/remember <text> [#tag] |
Save a memory about this project |
/memory [search <q>|delete <id>|export|import <file>|stats|cleanup|clear] |
Project memories |
/mcp [connect|disconnect <name>|tools] |
MCP servers and their tools |
/tools |
List available tools |
/upgrade [check|now|skip|changelog|status] |
Check for and install updates |
/doctor |
Diagnose the LLM server and tools |
/help |
Show this help |
Create ~/.taracode/config.yaml:
# The Ollama host (--host and TARACODE_HOST override it)
host: http://localhost:11434
# Generation options for the main chat
generation:
temperature: 0.7 # Sampling randomness (0.0-2.0)
top_p: 0.9 # Nucleus sampling threshold (0.0-1.0)
num_predict: 0 # Max tokens per response (0 = model default)
# Starting mode: investigate or operate (unset starts in investigate, or the mode a policy file names)
mode: investigate
# Security scanning
scan:
default_severity: "" # e.g. "HIGH,CRITICAL"
# Search
search:
primary: duckduckgo
fallback: searxng
brave_api_key: "" # Optional: Brave Search API
# Memory
memory:
enabled: true
auto_capture: trueSee config.example.yaml for all options.
| Backend | Setup | Notes |
|---|---|---|
| Ollama | brew install ollama |
Recommended, easiest setup |
| vLLM | Self-hosted | For production deployments |
| llama.cpp | Self-hosted | Lightweight option |
Phase 2 of the v3 plan (investigate and operate modes, the classified tools, the policy engine, redaction and the audit log) shipped in 3.0.0-alpha.2. Phase 3 shipped the evals and the first scoreboard in 3.0.0-beta.1, and 3.0.0 made the line stable. 3.1.0 settled the tool set at fifteen and the connection at one Ollama host. 3.2.0 added speed, memory and a board per machine to the evals. Runbooks, the MCP server and the skills pack are next, as 3.3 and 3.4. See ROADMAP.md.
taracode ships an offline eval suite: 33 DevOps tasks (Kubernetes triage, Helm, Terraform plan review, Docker and image security, secrets, cloud read-only investigation and refusal cases) whose tool calls replay from fixtures recorded against real clusters, charts, Terraform directories and images. No task runs anything for real, and every task drives the same loop, policy gate and redaction a session uses. Each task scores 0.4 for the tool calls it expected, 0.5 for the answer and 0.1 for never attempting a forbidden call; a refusal task also asserts that the gate denied the mutation, and a gate that lets one through fails the run rather than the model.
The published scoreboard is in docs/evals/scoreboard.md (and on
code.tara.vision/evals), regenerated with every release. Every run also
records tokens per second, --hardware names the machine so the report can rank models per machine, and
--gpu-probe records the GPU memory the machine itself reports (the evals reference shows the command for
an NVIDIA card). Run the suite against your own Ollama:
taracode eval run --host http://localhost:11434 --model gemma4:12b --hardware "your GPU or machine"
taracode eval reportResults land in docs/evals/results/, transcripts in evals/runs/. Writing a task, recording fixtures and
the reproducibility notes are in docs/evals/README.md.
make deps # Install dependencies
make build # Build binary
make test # Run tests
make coverage-gate # Check per-package coverage floors
make install # Install to /usr/local/binSee CONTRIBUTING.md for development guidelines.
Contributions are welcome! Please read our Contributing Guide and Code of Conduct.
For security issues, please see our Security Policy.
taracode is free, MIT-licensed and built without telemetry or a cloud service. If it saves you time, you can sponsor Tara Vision on GitHub to keep the releases, the model registry and the eval scoreboard coming.
MIT License - see LICENSE for details.
Built with ❤️ by Tara Vision · Created by Dejan Stefanoski