Skip to content

Repository files navigation

AI Code Reviewer

Multi-agent code review system that orchestrates multiple LLMs to produce comprehensive, consensus-based code reviews.

License: MIT


Overview

AI Code Reviewer takes a different approach to automated code review: instead of relying on a single model, it orchestrates multiple specialized agents that review code from different perspectives — security, performance, and code quality — then combines their findings into a unified, confidence-scored review.

Key Features

  • Multi-Agent Architecture: Run 2–5+ LLM agents in parallel, each with a specialized focus area
  • Consensus-Based Scoring: Findings are weighted by how many agents agree, reducing false positives
  • Anthropic Messages API: All models (Claude Sonnet 5, Haiku 4.5) accessed directly via the official anthropic SDK, streamed via messages.stream, with prompt caching, JSON-schema structured output, and tool use
  • GitHub Integration: Automatic PR reviews via webhooks, with inline comments and thread resolution
  • Incremental Reviews: Delta tracking detects new, fixed, and open findings across pushes — with convergence logic that stops reviewing when findings stabilize
  • Documentation Review: Rule-based check that flags missing doc updates on architecture-impacting PRs — works out-of-the-box on any repo by probing for CLAUDE.md, AGENTS.md, and architecture folders (zero LLM cost)

For the full technical deep-dive — pipeline flowcharts, scoring formulas, convergence state machine, and prompt engineering — see the Architecture Documentation.


Quick Start

# Install
pip install ai-code-reviewer

# Export credentials
export ANTHROPIC_API_KEY=sk-ant-...
export GITHUB_TOKEN=ghp_...

# Review a GitHub PR
ai-reviewer review-pr calimero-network/core 123

# Review a local diff
git diff main | ai-reviewer review --output markdown

How It Works

All LLM agents call Anthropic's Messages API directly via the official anthropic SDK. Security, performance, patterns, and logic agents run on claude-sonnet-5 (security and logic with adaptive extended thinking on); the style agent uses claude-haiku-4-5. Repo exploration happens through Claude tool use (read_file / glob / grep) backed by the GitHub Contents API — no cloning, no extra infrastructure.

flowchart LR
    PR["PR Diff"] --> Anthropic["Anthropic Messages API\n(claude-sonnet-5 / haiku-4-5)"]

    subgraph Agents["Parallel Agent Execution"]
        A1["Sonnet\n(Security)"]
        A2["Sonnet\n(Performance)"]
        A3["Sonnet\n(Patterns)"]
        A4["Sonnet\n(Logic)"]
        A5["Haiku\n(Style)"]
    end

    Anthropic --> Agents

    Agents --> Agg["Review Aggregator\n• Cluster similar findings\n• Compute consensus scores\n• Rank by severity × agreement"]
    Agg --> Delta["Delta Tracking\n• New / fixed / open findings\n• Convergence detection"]
    Delta --> Out["Consolidated Review\n(GitHub / JSON / MD)"]
Loading

For a detailed breakdown of the pipeline, scoring formulas, and convergence logic, see the Architecture Documentation.


Configuration

Create config.yaml:

anthropic:
  api_key: ${ANTHROPIC_API_KEY}
  default_model: claude-sonnet-5
  enable_prompt_caching: true
  max_combined_context_tokens: 80000

github:
  token: ${GITHUB_TOKEN}  # or Classic PAT for thread resolution (see below)

agents:
  - name: security-reviewer
    model: claude-sonnet-5
    focus_areas: [security, architecture]

  - name: performance-reviewer
    model: claude-sonnet-5
    focus_areas: [performance, logic]

  - name: style-reviewer
    model: claude-haiku-4-5
    focus_areas: [style, readability]
    allow_tool_use: false

orchestrator:
  timeout_seconds: 300
  min_agents_required: 2

# Documentation review (rule-based, no LLM cost)
doc_review:
  enabled: true
  architecture_paths: ["architecture/", "docs/", "doc/"]
  convention_files: ["AGENTS.md", "CLAUDE.md", "CONTRIBUTING.md"]

Local Review (before a PR exists)

Run the same multi-agent pipeline against uncommitted work, from inside a Claude Code session - no API key and nothing posted to GitHub. Reviewers are subagents in your session; clustering, consensus scoring, confidence floors and fix validation stay in Python, so local and PR reviews apply the same rules.

uv tool install git+https://github.com/calimero-network/ai-code-reviewer
/plugin marketplace add calimero-network/ai-code-reviewer
/plugin install ai-review@calimero
/ai-review                 # uncommitted changes, including untracked files
/ai-review --staged        # the index only
/ai-review --base main     # main...HEAD

Setup, scopes, configuration and the non-session CLI workflow: Local Review.


CLI Commands

# Review a GitHub PR (includes doc review by default)
ai-reviewer review-pr <owner/repo> <pr-number>

# Skip documentation review
ai-reviewer review-pr <owner/repo> <pr-number> --no-doc-check

# Force documentation review even if disabled in config
ai-reviewer review-pr <owner/repo> <pr-number> --doc-check

# Start webhook server
ai-reviewer serve --port 8080

# Configuration
ai-reviewer config validate
ai-reviewer config show

Output Example

Reviewed by 3 agents  |  Quality score: 87%

CRITICAL (1)
  SQL Injection in auth/login.py:45  [3/3 agents]
  User input interpolated directly into SQL query without parameterization.

WARNING (2)
  Missing rate limiting on /api/login  [2/3 agents]
  Inefficient O(n²) loop in process_batch()  [2/3 agents]

SUGGESTION (3)
  Add type hints to process_user()
  Extract magic number 86400 to a named constant
  Add docstring to AuthHandler

Repository Configuration

Add .ai-reviewer.yaml to the root of any reviewed repository to customize behavior:

# Exclude generated or vendored files from review
ignore:
  - "**/*.generated.rs"
  - "**/vendor/**"

# Append custom instructions to a specific agent's prompt
agents:
  - name: security-reviewer
    custom_prompt_append: |
      This is a Rust codebase using eyre for errors.
      Flag all unwrap() calls.

# Review policy
policy:
  require_human_review_for: [security]
  block_on_critical: true

GitHub Actions Setup

Basic Setup (GITHUB_TOKEN)

The default GITHUB_TOKEN provided by GitHub Actions is sufficient for most features:

  • Posting reviews and inline comments
  • Adding reactions
  • Posting "Resolved" replies

It cannot resolve review threads (collapsing them in the UI), which requires a Classic PAT.

Full Features (Classic Personal Access Token)

Note: Fine-grained PATs do not support the resolveReviewThread GraphQL mutation. Use a Classic PAT with repo scope.

  1. Create a Classic Personal Access Token with the repo scope.

  2. Add it as a repository secret named GH_PAT:

    Settings → Secrets and variables → Actions → New repository secret
    Name: GH_PAT
    Value: ghp_xxxxxxxxxxxxxxxxxxxx
    
  3. The workflow uses GH_PAT automatically when present, falling back to GITHUB_TOKEN.

For production deployments, prefer a dedicated service account and rotate tokens regularly. GitHub Apps with fine-grained permissions are the recommended long-term approach.


Development

git clone https://github.com/calimero-network/ai-code-reviewer
cd ai-code-reviewer

pip install -e ".[dev]"

pytest
ruff check .
mypy src/

AI Rules & Documentation

The repository ships structured AI context to help coding assistants work with the codebase effectively.

.ai/
├── context.md           # Codebase overview — read first
├── doc-bot.md           # Documentation bot instructions
├── prompts/             # Reusable AI prompts
└── rules/               # Per-module design rules
    ├── architecture.md  # High-level design & invariants
    ├── agents.md        # Agent module patterns
    ├── orchestrator.md  # Orchestration rules
    ├── github.md        # GitHub integration patterns
    ├── models.md        # Data model conventions
    └── conventions.md   # Coding style guide

The review-pr command includes a built-in documentation review that runs alongside the AI code review. It detects architecture-impacting changes (new modules, manifest edits, CI changes, infrastructure files) and checks whether convention files like CLAUDE.md or AGENTS.md were updated. For repos with .ai-reviewer.yaml, it also checks explicit source_to_docs_mapping rules. The check is rule-based (no LLM calls) and posts a separate PR comment with suggestions. Disable with --no-doc-check or doc_review.enabled: false in config.

Automatic Doc Update PRs

When stale documentation is detected, the reviewer can automatically generate updated file content and open a PR with the real changes — not just a comment telling you to update them.

Flow — a four-stage pipeline (Understand → Route → Apply → Verify):

  1. Feature PR merges to main; CI triggers ai-reviewer update-docs on the merge
  2. Understand — reads the full merged PR (title, body, commits, diff) once into a structured change summary (no blind diff truncation)
  3. Route — maps each change to a doc action: update an existing section, add a new section, or create a new page (honouring source_to_docs_mapping)
  4. Apply — drafts the edits: surgical FIND/REPLACE for HTML, additive sections, or whole new pages wired into nav.js/index.html
  5. Verify — a confidence gate; edits that don't reflect the change are flagged for a human, not shipped
  6. Commits confident updates to a new branch docs/auto-<sha> and opens a PR against main, assigned to the original PR author. The PR body previews each edited page as a GitHub-style diff (added doc text in green, removed in red) so a reviewer sees what changed without opening the raw HTML; flagged docs are listed in their own section. Nothing auto-merges.

Setup: two steps

  1. Enable generation in the repo's .ai-reviewer.yaml:
documentation:
  enabled: true
  source_to_docs_mapping:
    "src/**": [docs/api.md, README.md]

doc_generation:
  enabled: true
  understanding_model: claude-sonnet-5     # stage 1: full-PR comprehension
  apply_model: claude-haiku-4-5            # stage 3: drafting edits/pages
  allow_new_pages: true                    # may create + wire new pages
  verify_confidence_threshold: medium      # below this, flag for a human
  max_files: 5
  1. Add a one-line workflow to each repo (calls the shared reusable workflow):
# .github/workflows/doc-update.yaml
on:
  push:
    branches: [main, master]
jobs:
  doc-update:
    uses: calimero-network/ai-code-reviewer/.github/workflows/doc-update.yaml@main
    secrets:
      ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
      GH_PAT: ${{ secrets.GH_PAT }}

No copy-pasted workflow logic — the shared workflow lives in this repo and all repos call it via uses:.

Test locally before wiring up CI:

# Preview what would be generated — no PR opened, no files committed
ai-reviewer update-docs calimero-network/auth-frontend 42 --dry-run

# Actually generate and open the PR
ai-reviewer update-docs calimero-network/auth-frontend 42

Cost: ~$0.05–0.15 per merge with stale docs (Claude Sonnet, 1–5 files). Zero LLM cost when no stale docs are detected.


Related Projects


License

MIT License - see LICENSE for details.


Built with ❤️ by Calimero Network

About

Multi-agent code review system that orchestrates multiple LLMs

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages