Build LLM apps, agents, RAG, evals, and MCP with production discipline.
Real session, real output. Recorded with VHS.
Quick Start · Skills · Agents · Commands · Plugins · Memory · Model · Docs · Roadmap
LLMFoundry is a complete kit to develop production-quality LLM applications on DeepSeek (cheap, high-volume) with engineering discipline. It turns opencode into a team: an orchestrator that understands your intent and delegates to specialist agents, a living memory that learns and feeds itself, and gates that block bad output.
Built for $2.01. The entire kit, 30 skills, 14 agents, memory, gates, 60 tests, and full docs, developed on DeepSeek V4. This is what cost discipline buys.
Looking for: AI agents, RAG pipeline, prompt engineering, LLM evals, MCP servers, DeepSeek coding, AI engineering best practices, agent skills for opencode? You're in the right place.
You send a rough idea. The orchestrator turns it into a plan, verifies instead of guessing, and produces work you can trust. This is how it actually feels:
You: "I want an agent that pulls fresh docs for libraries and answers with sources"
Orch: Before I design this, I need one decision. What interface?
A) CLI B) web chatbot C) API D) MCP server inside opencode
You: "D"
Orch: Good, an MCP server. Let me specify:
SPEC: tech-docs-agent, scope, done criteria, out of scope.
Delegating to ai-architect with full context...
[ai-architect] returns a design with trade-offs
[llm-security-reviewer] reviews prompt injection and exfiltration
[deep-researcher] validates the cited sources
Orch: 3 agents ran in parallel. Contradictions resolved.
Here is the final plan, approved to implement.
And every claim in that flow was verified, not guessed. That is the LLMFoundry difference: a team that reasons, delegates, checks its own work, and writes like a human.
Four disciplines that make it different:
- Cost discipline, DeepSeek only (V4 Pro/Flash). 35-322x more requests per dollar than kimi-k3 on the same plan. Explicit PRO vs FLASH routing.
- Mode discipline, PLAN before BUILD when there's ambiguity or stakes. Never build what you don't understand.
- Anti-delirium, every claim has concrete proof (
file:line, command output, URL) or an honest[UNVERIFIED]marker. No hallucination, ever. - Human voice, output never reads like AI-generated text. No dashes, no AI vocabulary, no template structure.
git clone git@github.com:Pl3ntz/llmfoundry.git
cd llmfoundry
./scripts/install.shThat is it. Three commands and the kit is live in opencode.
The installer, in one pass:
- symlinks
agents/andcommands/into~/.config/opencode/ - registers
skills/viaskills.paths, so your existing skills stay untouched - installs Python deps for the memory engine (fastembed for semantic search)
- registers the gates, memory, voice, and verify plugins in your opencode config
Restart opencode. The orchestrator becomes your default agent, and you are working with a team instead of a single model.
What you had: one generic agent, prompt by prompt
What you get: an orchestrator + 13 specialists + a memory that learns + gates that protect
Requires: opencode, python3, DeepSeek V4 (Go plan or API key). Full guide in docs/INSTALL.md.
You ──→ AI Orchestrator (the Captain) ──→ specialist subagents
│
├─ deep-researcher (deep research with correlation + anti-injection)
├─ ai-architect (LLM system design with trade-offs)
├─ ai-evals-runner (prove it works)
├─ llm-security-reviewer(security before shipping)
├─ reverse-engineer (binary, firmware, malware analysis)
├─ red-team-agent (authorized offensive security)
├─ bug-bounty-hunter (scope to validated report)
├─ security-defensive (defensive audit, hardening)
├─ database-engineer (full PostgreSQL stack)
├─ data-model-engineer (data modeling, tenancy)
├─ backend-architect (API design, middleware, queues, caching)
├─ platform-engineer (infra, Docker, CI/CD, cloud, monitoring)
└─ api-contract-engineer(deep API contract work)
│
└─ Living Memory (SQLite + embeddings, self-feeding)
The flow: you send a raw idea → the orchestrator captures intent, asks one question at a time until it understands → defines the SPEC with you → rewrites into a master prompt → delegates with full context → synthesizes results and presents options.
30 skills in 5 categories. All follow the transversal standards.
| Skill | Use when |
|---|---|
ai-engineering-standards |
Start of every task, tone, evidence, anti-fabrication |
ai-dev-process |
Writing code: SPEC, worktree, TDD, atomic commit |
interview-me |
Request is underspecified or high-stakes |
ai-orchestration |
Routing, delegation protocol, fan-in synthesis |
human-voice |
Write in a natural human voice, never looks AI-generated |
anti-delirium |
Prove it or don't say it, evidence or confidence marker on every claim |
git-workflow |
Committing, branching, merging, resolving conflicts |
pull-request |
Creating and updating effective PRs |
code-review |
Extremely effective review, five-axis method, before merge |
| Skill | Use when |
|---|---|
ai-prompt-engineering |
Writing/iterating prompts |
ai-agent-patterns |
Designing agentic systems |
ai-context-engineering |
Managing context windows |
ai-rag-pipeline |
Building retrieval systems |
ai-evals |
Proving behavior, guarding regressions |
ai-model-integration |
Wiring providers, streaming, fallback |
ai-mcp-development |
Building MCP servers |
ai-agent-safety |
Sandboxing, permissions, fail-closed |
ai-llm-app-security |
Defensive LLM security (OWASP LLM Top 10) |
ai-llm-observability |
Tracing, cost/token tracking |
| Skill | Use when |
|---|---|
ai-research |
Deep research with correlation |
doubt-driven-development |
High-stakes review: CLAIM/EXTRACT/DOUBT/RECONCILE |
source-driven-development |
Grounding decisions in official docs |
debugging-and-error-recovery |
5-step debugging |
| Skill | Use when |
|---|---|
re-binary-analysis |
Identify format, arch, packing |
re-decompilation |
Recover logic from disassembly (radare2/Ghidra) |
re-algorithm-recovery |
Reconstruct crypto/checksums/serials with proof |
re-dynamic-analysis |
Confirm behavior under controlled execution |
re-malware-analysis |
Malware triage, IOC extraction, safe detonation |
re-firmware-analysis |
Extract and analyze device firmware |
pdf-processing |
Fast local PDF to text/Markdown (pdf-inspector, MIT, free) |
Full catalog: SKILLS.md
| Agent | Mode | Role |
|---|---|---|
| ai-orchestrator | primary (default) | The Captain, interprets, discusses, delegates, synthesizes |
| deep-researcher | subagent | Deep research with correlation + anti-injection |
| ai-architect | subagent | LLM system architecture with trade-offs |
| ai-evals-runner | subagent | Build and run evals |
| llm-security-reviewer | subagent | Security review of LLM apps |
| reverse-engineer | subagent | Binary, firmware, malware analysis |
| red-team-agent | subagent | Enterprise red team, pentest, exploitation |
| bug-bounty-hunter | subagent | Bug bounty, web and API hunting |
| security-defensive | subagent | Defensive audit, hardening, remediation |
| database-engineer | subagent | Full PostgreSQL: schema, indexes, EXPLAIN, RLS, migrations |
| data-model-engineer | subagent | Data modeling, normalization, partitioning, tenancy |
| backend-architect | subagent | Backend design: APIs, middleware, jobs, caching, queues |
| api-contract-engineer | subagent | Deep API contracts: OpenAPI discriminators, hypermedia, rate limit RFCs |
| platform-engineer | subagent | Infrastructure: Terraform, Docker, K8s, CI/CD, cloud, monitoring |
/ai-spec · /ai-build · /ai-evals · /ai-review · /ai-research · /ai-memory · /ai-re
| Plugin | What it does |
|---|---|
gates.ts |
Blocks commit without tests, secret files staged, secrets in outbound fetch/search |
memory.ts |
Captures errors→gotchas, commits→memory; injects recall into every session |
voice-guard.ts |
Flags output that reads like AI-generated text (dashes, AI vocabulary) |
verify-guard.ts |
Flags conjecture-as-grounding (probably, should be, i assume) per anti-delirium |
publish-guard.ts |
Injects mandatory human-voice + anti-delirium + standards gate into every system prompt |
delegation-guard.ts |
Validates subagent spawns: 4 mandatory parts + routing table check |
research-guard.ts |
Warns only when the ORCHESTRATOR fetches research directly; subagents (deep-researcher) research freely |
SQLite + FTS5 + local semantic embeddings (fastembed/ONNX). The living feedback loop: encode → consolidate → retrieve → reconsolidate. Auto-captures errors and agent findings; recall is injected into every session. 100% local, never versioned. See docs/MEMORY-SPEC.md.
| Role | Model | Requests/mo | Cost/MTok out |
|---|---|---|---|
| Default | opencode-go/deepseek-v4-pro |
17,150 | $0.87 |
| small_model | opencode-go/deepseek-v4-flash |
158,150 | $0.28 |
Routing rule: reasoning to PRO, mechanical to FLASH. Mode rule: ambiguity or stakes → PLAN first, clear + approved → BUILD. See docs/MODEL-POLICY.md.
llmfoundry/
├── agents/ # 14 agents (orchestrator + 13 specialists)
├── commands/ # 7 slash commands
├── skills/ # 30 skills (5 categories)
├── plugins/ # 7 plugins: gates, memory, voice-guard, verify-guard, publish-guard, delegation-guard, research-guard
├── evals/ # golden-sets, rubric, baseline
├── docs/ # architecture, model policy, memory spec, RE spec
├── references/ # shared checklists
├── templates/ # sanitized MEMORY templates (placeholders only)
├── scripts/ # install.sh, memory engine, eval runner, routing scorer
├── .github/ # CI workflow
└── assets/ # logo
The kit tests itself. scripts/eval-runner.py runs 44 deterministic checks, no model
calls: engine unit tests, routing golden-set validation, scorer cases, plugin compile
checks, and K=5 stability checks. A GitHub Actions CI runs them on every push/PR, so a
regression is caught before it ships. Baseline: evals/baseline.json.
python3 scripts/eval-runner.py # full suite
python3 scripts/eval-runner.py --baseline # show the number to beat| Doc | Covers |
|---|---|
| docs/ARCHITECTURE.md | System architecture and data flow |
| docs/MODEL-POLICY.md | Why DeepSeek, PRO/FLASH + BUILD/PLAN routing |
| docs/MEMORY-SPEC.md | Memory architecture, living loop, privacy |
| docs/INSTALL.md | Full install/uninstall guide |
| docs/CI-LOCAL.md | CI na VPS via Docker + cron, sem GitHub Actions |
| docs/REVERSE-ENGINEERING-SPEC.md | Reverse engineering specialist design |
| SKILLS.md | Skill catalog |
| CONTRIBUTING.md | How to add skills/agents/commands/evals |
- Core kit: skills, agents, commands, plugins, memory, evals
- Orchestrator (the Captain) as default agent
- Semantic memory (local embeddings)
- Gates as real commit blockers
- Routing eval (golden-set + deterministic scorer, validated manually)
- Regression CI (44 checks on every push)
- Human-voice + anti-delirium disciplines
- PRO/FLASH + BUILD/PLAN routing policies
- Stability checks (K=5, deterministic engine)
- Reverse engineering specialist (6 skills + agent + command)
- Recall includes memories + facts (not just findings/gotchas) so imported knowledge enters agent context automatically
- Fleet redesign: database-engineer, backend-architect, platform-engineer, security-defensive (14 agents, balanced coverage)
- Publish guard: mandatory human-voice + anti-delirium + standards gate on every system prompt
- Delegation guard: validate subagent spawns (4 mandatory parts + routing table)
- Research guard: enforce research delegation policy (no direct webfetch from orchestrator)
Routing is validated manually, one question at a time, to avoid batch sessions touching the user's Chrome. Batch automation of model-routing tests was removed.
See CONTRIBUTING.md. Follow the kit's own discipline: SPEC → TDD → worktree → atomic commit.
Honest numbers, sourced from a live deep-researcher pass over GitHub and npm (Aug 2026).
| Project | Stars | Focus | Gates | Anti-delirium | Evals | DeepSeek-first |
|---|---|---|---|---|---|---|
| LLMFoundry | new | complete kit for opencode | ✅ runtime plugins | ✅ | ✅ 44 checks | ✅ |
| agent-skills (addyosmani) | 81.2k | skills pack (multi-tool) | ✅ | ✅ evals/ | ❌ | |
| hiai-opencode | 12 | multi-agent + gates | ✅ runtime | ❌ | ✅ 986 tests | ❌ |
| GoopSpec | 37 | spec-driven workflow | ✅ contract gates | ❌ | ❌ | ❌ |
| CrewBee | 16 | agent teams | ❌ | ❌ | ❌ | |
| maestria | 2 | cross-IDE management | ❌ | ❌ | ❌ |
What no competitor combines: DeepSeek-first cost, an orchestrator + 13 specialists, living semantic memory, runtime quality gates, anti-delirium, and human-voice in one install for opencode. agent-skills is the closest in quality, but it is a skills pack, not a team with memory and gates, and it is not cost-optimized for DeepSeek.
Do I need to pay for expensive models? No, the kit is built for DeepSeek (V4 Pro/Flash), which is 35-322x more requests per dollar than kimi-k3.
Is my code/memory sent anywhere? No. Memory is 100% local (SQLite + local embeddings), never versioned. See docs/MEMORY-SPEC.md.
Does it work alongside my existing opencode setup? Yes, the installer uses
skills.paths and per-file symlinks, coexisting with existing skills/agents.
Is it only for opencode? Built for opencode, but skills follow the Agent Skills open standard (agentskills.io), portable to Claude Code, Codex, Cursor, etc.
How does it prevent hallucination? Every claim must have proof or a confidence marker
(anti-delirium skill + verify-guard plugin). The orchestrator verifies before asserting.
MIT. See LICENSE.
