Skip to content

v3.0.0 — a code review bot with a comment budget

Choose a tag to compare

@cirolini cirolini released this 19 Sep 18:54
· 20 commits to main since this release

v3 rebuilds this action around a single idea: the scarce resource in code review
is the reviewer's attention, and a bot that posts forty comments spends more of
it than it returns.

Two changed defaults. mode now defaults to review, so a workflow that
never set mode gets inline comments instead of one summary comment — pin
mode: files to keep the old output. And provider now defaults to gemini,
though a workflow passing the deprecated openai_api_key without a provider
stays on OpenAI.

Every v2 input still works, and @v2 itself is untouched. See
the migration guide.

Multiple providers

provider selects Google Gemini (the default), OpenAI, Anthropic, or any
OpenAI-compatible server — Ollama, vLLM, Azure OpenAI, OpenRouter — via base_url. New inputs:
provider, model, api_key, base_url, temperature, max_tokens. The
openai_* inputs remain as aliases.

Transient failures are retried with backoff; permanent ones are not, because
retrying a rejected key only delays the message you need. A provider failure is
now reported in the pull request and fails the check, rather than surfacing only
in a log — a review that silently did not happen looks exactly like a review
that found nothing.

Inline comments on the right lines

With mode: review, the model returns findings matching a JSON schema — file,
line range, severity, category, confidence, title, rationale, optional
suggestion — validated in code, with one repair attempt when it does not match.
Findings are posted as inline comments grouped into a single review.

Review pressure controls

  • max_comments (default 5), ranked by severity then confidence. Suppressed
    findings are counted and listed, never dropped.
  • min_severity and min_confidence.
  • One sticky summary comment, edited in place instead of a new one per push.
  • Incremental review: on a push, only the new commits are reviewed, and findings
    already raised are not repeated.
  • ignore_paths, defaulting to lockfiles, generated and vendored code,
    snapshots and minified assets.
  • Large diffs are chunked, and anything that did not fit is reported rather than
    silently truncated.
  • panel (experimental): two providers merged, with agreement as a confidence
    signal. Roughly double the cost.

Measured

Model Precision Recall Noise rate Cost / PR
gemini-3.8-flash 100% 93% 0% $0.0009
gpt-oss-120b 70% 93% 83% n/a

21 fixtures, 15 with a seeded defect and 6 clean. Recall did not separate the
two models; noise did. Full numbers and limits in docs/results/.

Measurement

Every run writes a JSON log — provider, model, tokens, estimated cost, latency,
findings by severity, posted versus suppressed, coverage — containing no
credentials, no diff and no source code.

evals/ holds 21 fixtures with seeded, labelled defects plus clean diffs, and
make eval scores precision, recall, noise rate, cost per pull request and
latency.

Fixed

Everything in v2.1: the container entrypoint path (#47), the unpinned httpx
that broke client construction, and the deprecated default model. Also the patch
parser, which split on the bare substring diff and broke on any diff that
mentioned the word (#28).

Security

The diff is treated as untrusted input: fenced with a per-request random
sentinel, with instructions placed after the data, and the model told that
injection attempts are themselves reportable. Use pull_request, not
pull_request_target, for pull requests from forks.


Quick start

on:
  pull_request:
    types: [opened, synchronize]

permissions:
  contents: read
  pull-requests: write

jobs:
  review:
    runs-on: ubuntu-latest
    steps:
      - uses: cirolini/genai-code-review@v3
        with:
          api_key: ${{ secrets.GEMINI_API_KEY }}
          github_token: ${{ secrets.GITHUB_TOKEN }}
          github_pr_id: ${{ github.event.number }}

No actions/checkout step is needed.

Known limits

Two adapters have not been exercised against a live API. openai shares its
class with openai-compatible, which is verified against Groq, so the code path
is covered — but gpt-5.6-luna itself has never been called. anthropic
authenticates and reaches Anthropic's billing layer, but the account used for
testing had no credit. Neither is the default.

The comment budget is unmeasured: no eval fixture produced more than five
findings, so max_comments never bound during testing.

Thanks

@arjunsuresh, @Asthethi, @fredlemieux, @tapegram, @keenan-repo and @dlidstrom,
who between them reported, diagnosed and fixed problems in v2 while this
repository was unmaintained.