You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
feat(code-review): show how each reviewer has actually performed
The settings page could say what a preset promises and nothing about what it
delivered. A lens that raises forty findings and keeps two is not thorough, it
is expensive, and from inside the product there was no way to tell those apart.
Each lens now reports raised versus kept across the workspace's recent
findings, as a bar and a count. Shown only once there is history — "0 of 0"
teaches nobody anything — and fetched once when the page opens rather than
polled, because "is this reviewer worth its cost" is a question you ask
occasionally. A failed fetch is silent: this is context, not a control, and a
settings page that cannot show a statistic must still let somebody change a
setting.
A finding counts once PER LENS that raised it, so two lenses agreeing counts
for both. That is the point rather than a rounding choice: agreement is the
signal the dedupe key exists to preserve, and splitting the credit would punish
a lens for being corroborated.
Aggregated in SQL with a lateral unnest of the jsonb array. The alternative
ships every finding's row to count a string, over a whole workspace's history.
Two things the tests caught that a typecheck could not. `execute` returns a
different SHAPE per driver — postgres-js hands back rows, pglite wraps them in
`{ rows }` — so the existing cast-to-array pattern in this repo works in
production and cannot be covered by the suite that would catch it breaking;
both shapes are accepted now. And an unvalidated finding counts as raised but
not kept, which is what stops every lens looking flattered while a review is
still running.