v3.0.0 — a code review bot with a comment budget
v3 rebuilds this action around a single idea: the scarce resource in code review
is the reviewer's attention, and a bot that posts forty comments spends more of
it than it returns.
Two changed defaults. mode now defaults to review, so a workflow that
never set mode gets inline comments instead of one summary comment — pin
mode: files to keep the old output. And provider now defaults to gemini,
though a workflow passing the deprecated openai_api_key without a provider
stays on OpenAI.
Every v2 input still works, and @v2 itself is untouched. See
the migration guide.
Multiple providers
provider selects Google Gemini (the default), OpenAI, Anthropic, or any
OpenAI-compatible server — Ollama, vLLM, Azure OpenAI, OpenRouter — via base_url. New inputs:
provider, model, api_key, base_url, temperature, max_tokens. The
openai_* inputs remain as aliases.
Transient failures are retried with backoff; permanent ones are not, because
retrying a rejected key only delays the message you need. A provider failure is
now reported in the pull request and fails the check, rather than surfacing only
in a log — a review that silently did not happen looks exactly like a review
that found nothing.
Inline comments on the right lines
With mode: review, the model returns findings matching a JSON schema — file,
line range, severity, category, confidence, title, rationale, optional
suggestion — validated in code, with one repair attempt when it does not match.
Findings are posted as inline comments grouped into a single review.
Review pressure controls
max_comments(default 5), ranked by severity then confidence. Suppressed
findings are counted and listed, never dropped.min_severityandmin_confidence.- One sticky summary comment, edited in place instead of a new one per push.
- Incremental review: on a push, only the new commits are reviewed, and findings
already raised are not repeated. ignore_paths, defaulting to lockfiles, generated and vendored code,
snapshots and minified assets.- Large diffs are chunked, and anything that did not fit is reported rather than
silently truncated. panel(experimental): two providers merged, with agreement as a confidence
signal. Roughly double the cost.
Measured
| Model | Precision | Recall | Noise rate | Cost / PR |
|---|---|---|---|---|
gemini-3.8-flash |
100% | 93% | 0% | $0.0009 |
gpt-oss-120b |
70% | 93% | 83% | n/a |
21 fixtures, 15 with a seeded defect and 6 clean. Recall did not separate the
two models; noise did. Full numbers and limits in docs/results/.
Measurement
Every run writes a JSON log — provider, model, tokens, estimated cost, latency,
findings by severity, posted versus suppressed, coverage — containing no
credentials, no diff and no source code.
evals/ holds 21 fixtures with seeded, labelled defects plus clean diffs, and
make eval scores precision, recall, noise rate, cost per pull request and
latency.
Fixed
Everything in v2.1: the container entrypoint path (#47), the unpinned httpx
that broke client construction, and the deprecated default model. Also the patch
parser, which split on the bare substring diff and broke on any diff that
mentioned the word (#28).
Security
The diff is treated as untrusted input: fenced with a per-request random
sentinel, with instructions placed after the data, and the model told that
injection attempts are themselves reportable. Use pull_request, not
pull_request_target, for pull requests from forks.
Quick start
on:
pull_request:
types: [opened, synchronize]
permissions:
contents: read
pull-requests: write
jobs:
review:
runs-on: ubuntu-latest
steps:
- uses: cirolini/genai-code-review@v3
with:
api_key: ${{ secrets.GEMINI_API_KEY }}
github_token: ${{ secrets.GITHUB_TOKEN }}
github_pr_id: ${{ github.event.number }}No actions/checkout step is needed.
Known limits
Two adapters have not been exercised against a live API. openai shares its
class with openai-compatible, which is verified against Groq, so the code path
is covered — but gpt-5.6-luna itself has never been called. anthropic
authenticates and reaches Anthropic's billing layer, but the account used for
testing had no credit. Neither is the default.
The comment budget is unmeasured: no eval fixture produced more than five
findings, so max_comments never bound during testing.
Thanks
@arjunsuresh, @Asthethi, @fredlemieux, @tapegram, @keenan-repo and @dlidstrom,
who between them reported, diagnosed and fixed problems in v2 while this
repository was unmaintained.