Harness your code.
Quick Start · Architecture · Usage · 中文文档
ReIN is an open-source agentic coding runtime that implements a complete harness architecture — the control plane that orchestrates LLM calls, tool execution, hook lifecycle, permission control, and plugin systems.
It supports both Anthropic Claude (cloud) and local LLMs (Ollama, LM Studio, llama.cpp, vLLM) for fully offline operation.
Rein (n.) — a strap fastened to a bit, used to guide a horse. In software, the harness that guides an AI agent: intercepting, evaluating, executing, and extending every action it takes.
Most agentic coding tools are closed-source black boxes. ReIN opens up the full runtime:
- See exactly how LLM tool calls are orchestrated
- Hook into every lifecycle event (PreToolUse, PostToolUse, Stop, etc.)
- Control permissions at 5 layers (admin → user → project → command → hook)
- Extend with plugins (commands, agents, skills, hooks, MCP)
- Run offline with any local model — no API key needed
- Python 3.11+
- An LLM backend (choose one):
pip install rein-harnessOr from source:
git clone https://github.com/BDeMo/ReIN.git
cd ReIN
pip install -e .# Cloud mode (Anthropic Claude)
export ANTHROPIC_API_KEY=sk-ant-xxx
rein direct
# Fully offline (Ollama)
ollama pull qwen2.5-coder:7b
rein direct --local
# Custom local server (LM Studio / llama.cpp / vLLM)
rein direct --local --local-url http://localhost:1234/v1 --local-model my-modelrein/
├── core/
│ ├── harness.py Core orchestrator — the heart of ReIN
│ ├── config.py Multi-layer settings hierarchy
│ └── conversation.py Session and message management
├── llm/
│ ├── provider.py Abstract LLM interface
│ ├── anthropic_llm.py Anthropic Claude (streaming + tool use)
│ └── local_llm.py Local LLM (Ollama / LM Studio / llama.cpp / vLLM)
├── tools/
│ ├── registry.py Tool registry and base class
│ ├── file_tools.py Read / Write / Edit
│ ├── bash_tool.py Bash with command filtering and security
│ └── search_tools.py Grep / Glob
├── hooks/
│ ├── engine.py Hook execution engine (command + prompt based)
│ └── types.py 9 lifecycle event types
├── permissions/
│ └── manager.py 5-layer permission model (allow / deny / ask)
├── plugins/
│ └── loader.py Plugin discovery and loading
├── server/
│ └── app.py FastAPI server with WebSocket streaming
├── client/
│ └── cli.py Terminal client with rich Markdown rendering
└── main.py CLI entry point
Every tool call passes through the full harness pipeline:
User Input
→ [UserPromptSubmit Hook] Validate / preprocess
→ LLM streaming response Generate text + tool calls
→ Tool call detected
→ [PreToolUse Hook] Validate / modify / block
→ [Permission Check] 5-layer allow / deny / ask
→ Tool Execution Run the tool
→ [PostToolUse Hook] React / log / feedback
→ LLM continues Feed result back
→ [Stop Hook] Validate task completion
| Event | When | Purpose |
|---|---|---|
PreToolUse |
Before tool execution | Validate, modify, or block |
PostToolUse |
After tool execution | React, log, feedback |
Stop |
Before agent stops | Verify task completion |
UserPromptSubmit |
User sends message | Input preprocessing |
SessionStart / SessionEnd |
Session lifecycle | Init / cleanup |
PreCompact |
Before context compaction | Preserve critical info |
Notification |
Any notification | Logging, monitoring |
SubagentStop |
Subagent completes | Validate subagent output |
Layer 1 managed-settings.json Enterprise admin (MDM deployable)
Layer 2 ~/.claude/settings.json User global preferences
Layer 3 .claude/settings.json Project-level settings
Layer 4 YAML frontmatter Command / Agent tool whitelist
Layer 5 PreToolUse Hook Runtime dynamic decisions
# Anthropic Claude
rein direct
# Local LLM (Ollama)
rein direct --local --local-model qwen2.5-coder:7b
# Custom system prompt
rein direct --system-prompt "You are a Python expert."# Terminal 1: start server
rein server --port 8765
# Terminal 2: connect client
rein client --url ws://localhost:8765/ws/chat
# Local LLM server
rein server --local --local-model llama3.1:8b| Endpoint | Method | Description |
|---|---|---|
/health |
GET | Health check |
/api/tools |
GET | List tools with schemas |
/api/settings |
GET | Current settings |
/api/chat |
POST | Non-streaming chat |
/ws/chat |
WebSocket | Streaming chat (full harness) |
ReIN supports two tool-use modes:
| Mode | How | Models |
|---|---|---|
| Native | OpenAI tool_call format |
qwen2.5, llama3.1, mistral, functionary |
| Prompt-based | Schemas in prompt, parses ```tool_call blocks |
Any model |
Auto-detected from model name. Force native with --native-tools.
| Server | Default URL | Install |
|---|---|---|
| Ollama | http://localhost:11434/v1 |
ollama serve |
| LM Studio | http://localhost:1234/v1 |
GUI |
| llama.cpp | http://localhost:8080/v1 |
./llama-server -m model.gguf |
| vLLM | http://localhost:8000/v1 |
vllm serve model |
| LocalAI | http://localhost:8080/v1 |
Docker |
| Model | Size | Tool Use | Notes |
|---|---|---|---|
qwen2.5-coder:7b |
4.7 GB | Native | Best coding model at this size |
qwen2.5-coder:1.5b |
1.0 GB | Native | Fast, lightweight |
llama3.1:8b |
4.7 GB | Native | Strong general purpose |
deepseek-coder-v2:16b |
9.0 GB | Prompt | Excellent at code |
codellama:7b |
3.8 GB | Prompt | Meta's code model |
| Variable | Description | Default |
|---|---|---|
ANTHROPIC_API_KEY |
Anthropic API key | — |
CLAUDE_MODEL |
Override model name | claude-sonnet-4-20250514 |
ANTHROPIC_BASE_URL |
Override API URL | — |
LOCAL_LLM_URL |
Local server URL | http://localhost:11434/v1 |
LOCAL_LLM_MODEL |
Local model name | qwen2.5-coder:7b |
| Package | Purpose |
|---|---|
| anthropic | Anthropic Claude API |
| httpx | Async HTTP for local LLMs |
| fastapi | API server |
| uvicorn | ASGI server |
| websockets | WebSocket client |
| pyyaml | YAML parsing |
| rich | Rich terminal output and Markdown rendering |
| prompt-toolkit | Input history and editing |
ReIN is inspired by and built upon ideas from:
- Anthropic — the Claude Code open-source plugin ecosystem and harness architecture
- Ollama — making local LLMs accessible to everyone
- FastAPI — elegant async Python web framework
- llama.cpp — efficient local model inference
- OpenAI — the tool-calling API convention adopted by local LLM servers
CC BY-NC-SA 4.0 — Non-commercial use. Share alike.