Six opt-in review features drove this release: a deterministic scan of the added lines for leaked secrets and risky IaC settings, the issues a pull request links and their acceptance criteria, the CI checks that had already failed on the head commit, the maintainer decisions earlier reviews were taught (review learnings), a per-call choice of the review dimensions a call's files can use, and an outgoing notification when a review completes or fails. The summary changed shape alongside them: every review now ends with the dedicated summary call, and the summary comment is edited in place on every round instead of being posted once. Most of the Fixed list is older follow-up-round behaviour that work brought to light: findings re-posted beside their own open threads, a follow-up that called the pull request clean or did not block on an open critical finding, and the prompt's own vocabulary (dimension numbers, context-block names, notes about how the input reached the model) showing up in posted text.
Every new setting is off by default, so an upgrade with an unchanged environment reviews without any of the new material: REVIEW_SECRET_SCAN_ENABLED and REVIEW_IAC_SCAN_ENABLED (with REVIEW_SECRET_SCAN_ENTROPY_THRESHOLD and REVIEW_SECURITY_SCAN_SKIPPED_FILES), REVIEW_TICKET_CONTEXT_ENABLED (with REVIEW_TICKET_CONTEXT_PROVIDER, REVIEW_TICKET_CONTEXT_MAX_ISSUES, REVIEW_TICKET_CONTEXT_MAX_CHARS and REVIEW_TICKET_CONTEXT_FROM_BRANCH), REVIEW_CI_CONTEXT_ENABLED (with REVIEW_CI_CONTEXT_INCLUDE_LOGS and REVIEW_CI_CONTEXT_MAX_CHARS), REVIEW_LEARNINGS_ENABLED (with REVIEW_LEARNINGS_MAX_PER_REPO, REVIEW_LEARNINGS_PROMPT_MAX_ITEMS and REVIEW_LEARNINGS_PROMPT_MAX_CHARS), REVIEW_DIMENSION_ROUTING_ENABLED, and the NOTIFICATIONS_WEBHOOK_* group, which stays off until NOTIFICATIONS_WEBHOOK_URL is set. Each feature's bounds are validated at boot while it is on, and turning learnings on while REVIEW_DECLINE_RECHECK_ENABLED is off fails startup. No new GitHub App permission is needed. One existing setting changes meaning: REVIEW_MAX_AI_CALLS now counts the summary call every review makes, so at 1 a review that makes its review call gets a counts-only summary.
There are no migration scripts; the schema is still managed by Hibernate. The first start creates one table, review_learning, whether learnings are on or not, and it stays empty until they are. Visible after the upgrade: a pull request small enough for one review call now also makes the summary call, so it costs one more model call, and the first round on a pull request reviewed before the upgrade edits that pull request's existing summary comment, found by its heading, instead of leaving it as it was. The image moves to the 2026-09-27 UBI9 micro base.
Added
- Opt-in per-call review dimension routing (#665): every review call carried all ten review dimensions of the system prompt, whatever its files were, so a documentation or test batch paid for the IaC rules, the pagination rules and the config-key documentation rules, and on a large pull request that cost was paid once per batch. The prompt is now a core plus one block per dimension, and with
REVIEW_DIMENSION_ROUTING_ENABLED=trueeach call carries only the blocks its files can use. The router reads each file's kind (documentation, configuration or infrastructure, script, source, test) and probes its patch for API or list code and for mocks. Correctness, security and regressions are always included, an unrecognized file type brings every block in, and a batch of mixed files gets the union, so doubtful cases resolve to more blocks, not fewer. A documentation-only call's system prompt shrinks by about 27%, a source-only call's by 19–22%. The budget planner sizes the shared overhead from the pull request's routed prompt, which no batch's exceeds, so a pull request without configuration, test or API code also fits more diff per call. A routed prompt puts the whole core first, so every call shares a ~7,000-token prefix for a provider's prompt cache. The verifier's dimension carve-outs follow the files its candidates are anchored in. Each call's routing is logged at INFO with the file that brought each block in. Off by default: off, both prompts are byte for byte what they were. The eval corpus gains a must-find case for each routed dimension that had none (comment contradicts code, quality and complexity, pagination, config/IaC, config-key documentation), and a unit test now fails when routing would drop a dimension a corpus case needs - Opt-in outgoing notification when a review completes or fails (#73): set
NOTIFICATIONS_WEBHOOK_URLand each final review outcome is posted to it once, as a structured JSON payload or as a Slack or Discord incoming-webhook message (NOTIFICATIONS_WEBHOOK_FORMAT). By default only metadata leaves the process: repository, PR number, head commit, verdict, finding counts by severity, failure category, dashboard link, timestamp and bot version;NOTIFICATIONS_WEBHOOK_INCLUDE_CONTENTadds the PR title and finding titles, never finding prose or code. Requests can be signed with HMAC-SHA256 (NOTIFICATIONS_WEBHOOK_SECRET, sent asX-Thrillhousebot-Signature-256), run off the review thread with a bounded timeout and retry, never follow redirects, and never log the URL beyond its scheme, host and port. A run superseded by a newer push sends nothing, and a verdict held on CI is sent once, not again when CI turns green. The configuration is validated at boot - Opt-in CI-failure context in the review (#59): with
REVIEW_CI_CONTEXT_ENABLED=true, the checks that had already failed on the head commit when the review started reach the review prompt, so a finding can name the failing test or build step it explains. Each failing check contributes its name, conclusion, output title and summary, and a page of annotations;REVIEW_CI_CONTEXT_INCLUDE_LOGS=trueadds the tail of up to two failing GitHub Actions job logs under theActions: Readpermission the app already has. The checks come from the CI gate's own read, the section is capped byREVIEW_CI_CONTEXT_MAX_CHARS(validated at boot) and counted in the token budget, and the output is stripped of control characters and fenced as untrusted data. Checks still running add only a count next to a failure, and a CI failure after the review does not trigger a new one - Opt-in review learnings (#38): with
REVIEW_LEARNINGS_ENABLED=true, a maintainer's decline of a finding is remembered for the repository, with its reason, once it has survived the decline re-check, and a later review of a related file is told not to raise the same finding while that reason still holds. A decline whose reason rests on a premise the code could refute (a "this cannot run concurrently" answer to a race finding, PR #160's case) is never stored, nor is one that stood only because the maintainer answered the re-check twice; a 👍/👎 alone is not a learning./remember <text>stores an explicit convention,/learningslists what is remembered with source links, and/forget <id>retracts one; all three need write access, andGET /api/dashboard/learningsshows the history. Learnings are scoped to one repository and installation, refused when they contain anything credential-shaped, capped per repository (REVIEW_LEARNINGS_MAX_PER_REPO) and per prompt (REVIEW_LEARNINGS_PROMPT_MAX_ITEMS,REVIEW_LEARNINGS_PROMPT_MAX_CHARS), and fenced as untrusted data in the review call only. Startup fails if the store is on whileREVIEW_DECLINE_RECHECK_ENABLEDis off. The newreview_learningtable is created on first start - Opt-in linked-issue context and ticket compliance (#58): with
REVIEW_TICKET_CONTEXT_ENABLED=true, the issue(s) a pull request links (closing keywords in the PR body, then GitHub's closing references, and optionally the head branch name) are read from GitHub Issues and their title, acceptance criteria and body reach the review. The review call reads them as intent; the summary call lists acceptance criteria the change does not address under Description vs. Implementation, never as findings. Same-repository issues only, read-only, capped byREVIEW_TICKET_CONTEXT_MAX_ISSUESandREVIEW_TICKET_CONTEXT_MAX_CHARS(validated at boot) and counted in the token budget; the text is stripped of control and bidi characters and fenced as untrusted data. The tracker sits behind a provider interface (REVIEW_TICKET_CONTEXT_PROVIDER) so an external tracker can be added later - Opt-in deterministic scan for leaked secrets and risky IaC in the added lines (#60):
REVIEW_SECRET_SCAN_ENABLEDmatches AWS access key IDs, GitHub, Slack, Google API and Stripe live tokens, private keys, JWTs, and credential-named assignments whose literal passes an entropy threshold (REVIEW_SECRET_SCAN_ENTROPY_THRESHOLD, default 3.5);REVIEW_IAC_SCAN_ENABLEDmatches admin ports open to0.0.0.0/0or::/0, public S3 buckets, IAMAction: "*"onResource: "*", privileged pods and host namespaces, a Dockerfile whose final stage runs as root, and Terraformencrypted = false. No model call is made: each match is a finding at a fixed grade (a known-format credential is critical) that joins the review after the verifier, so no model demotes it and it is never marked unverified. A matched secret is shown only as its first characters and length in every comment, log line, stored session and dashboard view, and is scrubbed from model findings that quoted it; a model finding on the same defect is dropped, and a later round does not post the same finding again. The review's ignore globs apply, fixtures and snapshots are skipped (REVIEW_SECURITY_SCAN_SKIPPED_FILES), placeholders such asexample,changeme,<...>and${...}are ignored, and athrillhousebot:allow-secret/thrillhousebot:allow-iaccomment exempts a line. Both halves are off by default and the threshold is validated at boot
Changed
- The summary comment is updated in place on every review round (#868): the summary was posted on the first round only, so after a few pushes it described code the pull request no longer had while the current state was spread across review bodies. Each later round now finds the bot's summary comment by a
<!-- thrillhousebot:summary -->marker (or, for summaries posted before this release, by its heading) and replaces its body with the current render, disclosures included; a round that renders the same body spends no write. When no summary is left (a maintainer deleted it) the round posts a new one, and a failed edit falls back to posting one, logged at WARN. The comment keeps no history: inline findings, review bodies and the opt-in delta comment stay per round, and GitHub keeps the edit history. An edit notifies nobody, so a follow-up round still posts its review and delta comment as before - Every review ends with the dedicated summary call, and the review call returns findings only (#664): the single-call lane asked one response to carry the findings and the whole summary object (counts, verdict, purpose, description gaps, file summaries, labels, diagram), so the hottest call in the system spent input on that spec and output on the object, and a response cut at its length cap (as one was in production, inside
file_summaries) lost the summary with it. Both lanes now run review, refine, then the summary call the multi-call lane already made: the review prompt asks forfindingsandprevious_findings_statusand nothing else, the label list and diagram request ride the summary call only, and the summary is written from the findings after verification, so its counts describe what is posted. A cut review response keeps its summary, and a summary call that fails, is cut or meets the spend ceiling leaves the same disclosed counts-only summary on either lane instead of failing the review. The rendered summary comment is unchanged.REVIEW_MAX_AI_CALLSnow counts the summary call on both lanes; at1a review that makes its review call skips the summary call and says so, while a review whose every file exceeded the budget, which makes no review call, still gets its summary.REVIEW_OUTPUT_BUFFER_TOKENSkeeps its 8192 default: it reserves room for the response cap, not the expected response, and on a shared window it must also cover the concise summary call's own 8192 cap - The summary omits the Description vs. Implementation section when the check found no mismatch (#867): the clean state's heading and "No mismatch found" line appeared on most reviews and pushed the findings down. A summary that came back means the check ran, and a counts-only summary (no model prose) is visibly different, so an absent section on a summary with model prose means the description matched. Listed gaps and the collapsed "reported as a finding below" state render as before
Fixed
- Notes on how the review's input reached the model, and its account of the confidence rule it applied, no longer reach review text (#975): posted findings read "I am rating this claim at the required arithmetic-claim confidence", quoted a withheld path as "rust/Cargo.lock (excluded from review scope by the ignore list)" as if the pull request said it, and cited "the repository's config key definition for POLL_INTERVAL (supplied from …/Config.kt)" and an issue "as provided with this review". The review prompt now asks the model not to explain its confidence or describe how material reached it, and to cite a key's defining file by path; the guard behind it (#918) drops the confidence-rule clause, the withheld-path note and "as provided with this review", and rewrites "(supplied from )" to "(in )" and "config key definition for" to "definition of".
- The prompt's context-block names no longer reach review text (#950): the review prompt names its context blocks ("config key definitions", "the provided material"), so posted findings read "line 16 of the config-key definition section". The prompt now lists these names among the vocabulary not to cite, and the guard behind it (#918) rewrites what still leaks, in findings, review bodies, the summary and conversational replies alike: the block's name becomes "the configuration code" and "the provided material" becomes "the reviewed code".
- A "Things to double-check" bullet is linked "Same issue as" only to the inline finding it restates (#951): matching a bullet's title against any finding's description had linked configuration-doc items to a Dockerfile finding and an events-table item to an unrelated poll finding. A bullet is now linked only when its title restates the finding's title, or when the two sit within three lines of each other in one file and the bullet restates the description.
- A follow-up round requests changes while an earlier critical finding is still open (#948): the verdict was computed from the round's new findings and the previous round's
unresolvedstatuses, while the findings only the approval backstop kept open (every earlier round's, once a later round raised anything) could turn an approval into a comment but never request changes. On ThrillhouseBot-test#151 the summary listed an open CRITICAL SQL injection while the round, whose new findings were all low, ended as COMMENT. The verdict is now computed over the same still-open set the summary lists, so a held finding blocks at its own severity under the configured blocking strictness. - A follow-up round no longer re-posts findings that are still open on their threads (#939): a follow-up is shown one earlier round's findings by number, the newest round that raised any, so once a round raised even one new finding, the findings the rounds before it left open were no longer listed to the model, and the next
/reviewposted most of them again on new threads beside the open ones. The retry after an invalid response on ThrillhouseBot-test#146 was incidental; the retry already resends the first attempt's prompt unchanged, and a test now pins that. Findings earlier rounds left open on their own threads are now listed in the prompt as already posted, and before a round is summarized and published it drops any finding that restates a prior finding, from any round, still open on its own thread with its code still in the diff. - The approval backstop holds each open finding on its own, at its most severe rating (#934): it merged two findings one round raised into a single hold, so a reply on either thread cleared both, and when it held one defect raised again at a new severity it kept the first rating rather than the more severe one. Each finding is now held separately, and a defect raised again is held at the more severe rating.
- GitHub's "submitted too quickly" 422 is retried as a throttle instead of spending the finding's fallback routes (#919): GitHub's content-creation limit on review comments also answers
422 Validation Failedwith "was submitted too quickly", which the write retry read as a refusal of the payload. So the comment was not repeated, and the finding went down its suggestion-less and file-level routes, each refused the same way within milliseconds; with twelve pull requests reviewed at once, 118 writes were refused and 42 findings ended with no thread. A 422 carrying that phrase or GitHub's secondary-limit wording is now a throttle, waited out and repeated on the same comment and line, while any other 422 still fails fast. The suggestion-less and file-level routes run only after GitHub refuses the payload or the anchor ("could not be resolved"), never after a throttle. A throttle's wait now also holds the process-wide write pacer, so writes from other reviews queue behind it instead of drawing refusals of their own. A finding still throttled when its retries run out, or once the review's write-retry budget is spent and it is not retried, is listed in the review body with a note that GitHub's rate limit refused it and a re-run may post it, and the review logs once how many findings ended without a thread - Findings no longer cite the review prompt's own guidance blocks (#918): the prompt numbered its review dimensions and cited them by number ("dimension 4") throughout, and the model cited them back, so posted findings read "Dimension 7 artifact-name mismatch", "Comment-contradiction (dimension 4)" or "Parsing-rule probe (heuristic section)", references to instructions the maintainer has never seen. The leak predates dimension routing (#665): pull requests reviewed in August without routing carry it at a similar rate. The dimension headings are now unnumbered, the rules name each dimension by what it checks, and the prompt tells the model not to name, number or quote its instructions in a finding. A deterministic guard also removes the labels from finding titles and descriptions, previous-finding notes, the summary's prose, the summary comment, the delta comment and the review body before anything is stored or posted. It only matches shapes the prompt produced, so
dimensionas an ordinary word, a numbered tensor axis, code spans and code blocks are left as written. The "previous findings remain unresolved" sentence and the directive acknowledgements showedpath/to/File.java:42as their example on C and Rust pull requests; they now show<path>:<line> - A follow-up review no longer says the pull request has no issues while earlier findings are open (#933): a
/reviewthat opened nothing new over open threads posted "ThrillhouseBot found no issues in this PR" under CHANGES_REQUESTED whenever CI also held the verdict. The review body now leads with "No new issues in this revision, but N previous finding(s) remain unresolved" and lists the CI hold after it without the no-issues claim, and the held check-run summary adds the same open count. Both read the count the delta comment's "still open" line uses. With nothing open, the wording is unchanged - A bracket in the deliberation ahead of the answer no longer costs the review its findings (#894): the parser and the truncation salvage both started reading the response at its first
{or[, wherever it was. A 256,302-character production response opened with deliberation whose first[was a[LOW]severity tag at index 242, 252,138 characters before the answer, so the salvage of that cut response recovered nothing and the review failed with a complete findings array in hand. Both now start at the answer, recognized as the tail of the body: the earliest object from which the rest reads as a run of JSON documents to the end (a cut counting as the end), whose first document opens on or holdsfindings,previous_findings_statusorsummary(the verifier's salvage usesverdicts). That finds the answer whether it is fenced, unfenced or cut before its closing fence, past any fenced Swift, diff or JSON excerpts, and past a previous round's findings quoted back as JSON in the deliberation, which must not be published again. A root that opens on another key is still read whole, cut or not; a response split over several documents is still merged with the same warning; and a response with no such object is read exactly as before - The reasoning step-down no longer bills the cap a second time on a model that writes its deliberation into the response (#893): a call that stops at its length cap with no content is repeated once with reasoning disabled (#839), which frees the output allowance only where the provider keeps reasoning in its own channel. In production a model on the Ollama endpoint answered that repeat by writing its deliberation into
contentinstead: 252,380 characters of prose before the answer opened, then the same cap again, and the review was billed 65,536 completion tokens twice with nothing posted. The repeat is now watched: once its content reaches 8,000 characters with no JSON object opened in it, it is stopped and its connection closed. The call then fails with the first call's truncation, the one billed at the cap, and the log line, the failed check run and the pull-request notice state its cap and usage, what the repeat showed, and that loweringAI_REASONING_EFFORT(orAI_REASONING_EFFORT_CONCISEon the summary lane) is the other lever. A batch or summary call stopped this way is disclosed as before, and the summary's review-scope note says the repeat was stopped, apart from the existing note for a repeat that ran. A provider whose repeat answers from its first characters, including after a short lead-in or a code fence, is not affected, and deliberation that quotes a JSON object before the bound is left to run as it did before - A cut response states the cap the request licensed and what the provider billed, and advises raising the cap only when the stop reached it (#895): on a length stop the log line, the failed check run and the pull-request notice told the operator to raise
max-output-tokenswithout saying what it was set to or what the provider had honoured. In production every request carriedmax_tokens=96000and the provider stopped both passes at exactly 65,536 completion tokens, so the advice pointed at a setting that could not have changed the outcome, and the only figure given was a character count nobody can compare with a token setting. All three surfaces now state the licensedmax_tokensand the setting that supplied it (the active model's per-modelmax-output-tokenskey, orREVIEW_CONCISE_MAX_OUTPUT_TOKENSon the concise lane), the billed completion tokens and the prompt tokens. The setting is named as the remedy only when the billed completion reached the cap; a stop short of it is reported as a provider-side bound that a higher setting will not move. A request that carried no cap says the provider default applied, and a provider that reports no usage gets today's advice with an explicit "usage not reported". The blocking lanes (/improveand siblings, the verifier, replies) now carry the provider's usage on the failure as the streaming lane already did, and the character count survives only as the size of the partial body that reached the salvage step. The concise-lane advice no longer says to leaveREVIEW_CONCISE_MAX_OUTPUT_TOKENSunset to reach the provider default: unset, it falls back to 8192, and only an empty value drops the cap - A finding the verifier never ruled on no longer reads as if it had been screened (#885): verification fails open, so an empty response body, a cut response, an error or a call skipped at the review's spend ceiling keeps the findings the reviewer raised, and #623 made the round's banner say how many went unverified. The finding itself still posted at the severity and confidence the review pass gave it, with no second opinion behind it, as one did on #871. Such a finding now posts with its confidence capped at medium and a line in its own text saying the second-pass audit returned no verdict on it. Nothing is dropped and nothing moves: medium keeps a medium-risk finding on the diff, where low would have sent it to the collapsed "Things to double-check" block and made it eligible for the litigated-anchor withhold, a loss caused by the verifier's silence alone. Under the default
balancedstrictness a capped finding can no longer request changes on its own;strictstill blocks on critical or high risk at any confidence, as it does for verifier demotions. The same applies to a candidate a partial or cut response left without a verdict, so the marked findings are exactly the ones the banner counts, and the banner now says they post capped. Production logs show the path is rare (none in a week on 0.6.9, the ceiling disabled there), which is why holding the findings for a retry on the next round was declined - The release's Trivy scans no longer leave a stale warning in the Security tab (#869): code scanning measures a category's freshness against the default branch, and the release scans run on a tag, so
trivy-releaseandtrivy-release-distrolessnever received an analysis onmainafter June and were reported as months out of date. The release scans still gate on CRITICAL and HIGH in both images and now upload no SARIF; the reports are kept as the run'strivy-release-sarifartifact, and the job no longer holdssecurity-events: write.mainstays covered by the image scans inci.ymlandsecurity-scan.yml. - The weekly image scan reads
.trivyignore(#1005): the scheduled scan of:latestand:latest-distrolessran without the ignore file the pull-request and release scans use, so every documented no-fix CVE reopened its code-scanning alert each Monday. It now applies the same file - A malformed count in the summary response no longer drops the whole PR summary (#944): the model sent an array for the summary's
highcount, the summary node failed its schema mapping, and the salvage dropped it, so the summary comment had no "What this PR does" paragraph and every Changed Files row read "Not summarized". A count field that is not a whole number is now coerced when it is a numeric string and dropped otherwise, since the risk table recomputes the counts from the findings, and the rest of the summary maps normally.
Security
- The dashboard shows each login only the repositories it can access (#991): only
/feedbackand/learningschecked the repository;/sessionslisted every installed repository's reviews (and any repository named in?repository=),/sessions/{id}returned any session by id with its findings, quoted code, PR title and error,/costs,/tokensand/summaryaggregated across all repositories, and the WebSocket sent every live review to every connection, so a collaborator on one installed repository could read another private repository's reviews. Present since v0.1.0. The owner still sees every installed repository; anyone else sees the installed repositories they collaborate on, resolved once per request from the cached installation snapshot and applied in the queries, so paging totals and aggregates count only those. A?repository=outside the set answers 403 and a session id from one answers 404, the same as a missing id. The installation-wide skip counters in the summary go to the owner only - The native image moves to the 2026-09-27 UBI9 micro base, and the image CVEs with no upstream fix are documented instead of failing the release (#915): the new base ships the fixed
libacl(CVE-2026-54369), so that ignore and itslibattrcompanion are gone.pcre210.40 (CVE-2026-89161, CVE-2026-86145) has no fixed el9 package, and Red Hat marks CVE-2026-86145 "will not fix"; it is in the image only because the base'slibselinuxrequires it, and the bot's native binary linkslibzandlibcalone, with its regular expressions compiled fromjava.util.regex.libssl3in the distroless image (CVE-2026-84782) is a DTLS-only defect; the bot uses TLS over TCP through the JVM's own implementation and never loadslibssl3. Each ignore in.trivyignorestates the condition for removing it - Bumped
devaluefrom 5.9.2 to 5.9.4 (#999) anddompurifyfrom 3.4.13 to 3.4.16 (#998) inwebsite/for seven advisories: CVE-2026-92708, GHSA-mcm9-63f2-9j32 and GHSA-x5rw-q4pp-hg5g (high), GHSA-hx4r-w6wj-j8fg and GHSA-4q55-j62x-fr9h (medium), GHSA-wf3x-273g-mvxv and GHSA-p98j-92pf-mc4p (low). Both are transitive dependencies of the documentation site (through Astro and Mermaid); neither is in the bot image. - CVE-2026-103111, a third
pcre2out-of-bounds write in the UBI9 micro base (fixed upstream in 10.49, no fixed el9 package), joins the twopcre2CVEs already documented in.trivyignore(#1001). As with those, the native binary does not link PCRE2; the library is only in the base image forlibselinux, and the distroless image does not ship it. - CVE-2026-14456 (OpenSSL QUIC server) is resolved for the published images (#1002): Debian's tracker now marks
libssl33.0.x in Debian 12 as not affected, since that branch has no QUIC server code, and Trivy no longer reports it, so its.trivyignoreentry is removed. - Bumped
undicifrom 8.10.0 to 8.11.2 inwebsite/(#892) andfrontend/(#905) for eleven advisories, three of them high: CVE-2026-19534 (denial of service through an unrequested WebSocket subprotocol), CVE-2026-84961 (TLS certificate validation bypassed inBalancedPool) and CVE-2026-85152 (cross-origin cache poisoning in the interceptors). The others are CVE-2026-18149, CVE-2026-84890, CVE-2026-84933, CVE-2026-85014 and CVE-2026-85024 (medium) and CVE-2026-18540, CVE-2026-84947 and CVE-2026-85008 (low). It reacheswebsite/through the docs build's font loader andfrontend/throughjsdomin the test toolchain, and is not in the bot image
Documentation
AGENTS.mdrecords the invariants the build and tests enforce, where each review pipeline stage lives, and the alternatives that were tried and rejected, andCONTRIBUTING.mdpoints to it before a change to the review pipeline, prompts or model calls.docs/ARCHITECTURE.mdnow names the patch-coverage class correctly,PatchCoverageResolver(#673, #910)- The README, architecture, comparison and contributor docs are updated for 0.7.0 and simplified: shorter sentences, each fact stated once, and the release's new review features described where they apply
Dependencies
- Bumped Quarkus from 3.39.3 to 3.39.4 (#887)
- Bumped the actions group:
codecov/codecov-action7.1.0 to 7.1.1 andgithub/codeql-action4.38.0 to 4.38.1 (#891) - Bumped the frontend npm group:
next16.3.5 to 16.3.6,@types/node26.5.1 to 26.6.3,jsdom30.0.1 to 30.1.1 andvitest5.0.0 to 5.0.2 (#913) - Bumped the website docs group:
astro7.3.2 to 7.3.3 and@astrojs/starlight0.42.0 to 0.42.2 (#890)
What's Changed
Features
- refactor(notification): make the webhook request a record by @devops-thiago in #997
- release: 0.7.0 by @devops-thiago in #980
Fixes
- fix(review): read typed and array declarations in the secret scan by @devops-thiago in #928
- fix(review): keep every still-open finding in the summary edited in place by @devops-thiago in #929
- fix(review): stop findings citing the review prompt's own guidance blocks by @devops-thiago in #931
- fix(github): treat GitHub's "submitted too quickly" 422 as a throttle by @devops-thiago in #930
- fix(review): stop follow-up rounds calling the PR clean while findings are open by @devops-thiago in #936
- fix(review): keep a redacted secret finding open on an unchanged head by @devops-thiago in #937
- fix(review): carry description gaps across rounds and keep linked-issue numbers real by @devops-thiago in #938
- fix(review): give extensionless files a type so learnings reach them by @devops-thiago in #941
- fix(review): carry the summary's earlier findings by identity, not similarity by @devops-thiago in #943
- fix(review): stop follow-up rounds re-posting findings still open on their threads by @devops-thiago in #942
- fix(review): keep the PR summary when a count field is malformed by @devops-thiago in #949
- fix(review): keep an earlier round's declined finding open and block on it by @devops-thiago in #953
- fix(review): count an open double-check item once when a later round raises it again by @devops-thiago in #955
- fix(review): keep learning ids and context-block names out of review text by @devops-thiago in #954
- fix(review): read slice and array type annotations in the secret scan by @devops-thiago in #966
- fix(review): replace a double-check item only with the same defect by @devops-thiago in #968
- fix(review): read a bare findings array as the answer's findings by @devops-thiago in #967
- fix(review): strip carried-gap labels from description gaps by @devops-thiago in #973
- fix(review): rewrite carried-gap labels used as a subject or a reference by @devops-thiago in #976
- fix(review): keep confidence-rule narration and input notes out of review text by @devops-thiago in #977
- fix(review): keep a declined security-scan finding settled across rounds by @devops-thiago in #984
- fix(dashboard): scope sessions, costs and summaries to the repositories the user can access by @devops-thiago in #995
- fix(review): keep open earlier findings through the text scrub and read long redaction markers by @devops-thiago in #996
Other changes
- fix(docker): move to the 2026-09-27 UBI9 micro base and document the unfixable image CVEs by @devops-thiago in #915
- chore(release): prepare 0.7.0 by @devops-thiago in #978
- docs: update the comparison and contributor process docs for 0.7.0 by @devops-thiago in #988
- docs: update the architecture overview and AGENTS.md for 0.7.0 by @devops-thiago in #990
- refactor(review): clear static-analysis findings before 0.7.0 by @devops-thiago in #992
- docs: update and tighten the README for 0.7.0 by @devops-thiago in #989
- deps(docs): bump devalue and dompurify in /website by @devops-thiago in #1000
- chore(security): document CVE-2026-103111 and record the website advisory bumps by @devops-thiago in #1001
- chore(security): drop the stale CVE-2026-14456 Trivy ignore by @devops-thiago in #1002
- refactor: replace collection loops with streams (Sonar S9391) by @devops-thiago in #1003
- docs(changelog): list in 0.7.0 only the fixes for behaviour 0.6.9 shipped by @devops-thiago in #1004
- ci(security): apply .trivyignore to the weekly image scan by @devops-thiago in #1005
Full Changelog: v0.6.9...v0.7.0