Skip to content

Releases: Gilbert09/talyn

0.3.5

Choose a tag to compare

@github-actions github-actions released this 28 Sep 18:25
62eb0ff
v0.3.5

feat(marketing): put code review on the site, led by the thing it doe…

0.3.4

Choose a tag to compare

@github-actions github-actions released this 28 Sep 12:30
3462934
v0.3.4

docs: session notes for the review outcome events and the fix commit …

0.3.3

Choose a tag to compare

@github-actions github-actions released this 28 Sep 09:43
2ef8738
v0.3.3

feat(code-review): re-show the introduction, now that it introduces s…

0.3.2

Choose a tag to compare

@github-actions github-actions released this 28 Sep 00:48
49fb65d
feat(code-review): only dispatch the reviewers a change actually needs

Every lens read every change. A pull request of three documentation files still
booted a security sandbox, a reliability sandbox and an operability sandbox,
each spending six to ten minutes to report nothing — and a reviewer with
nothing to say either returns empty or invents something the judging pass then
pays to reject.

Files are classified coarsely (code, test, config, migration, docs, asset) and
each lens declares which classes it needs at least one of. Coarse on purpose:
the finer the classification, the more confidently it is wrong.

THE RULE IS DELIBERATELY ASYMMETRIC, and that is the whole safety argument. A
lens is skipped only on POSITIVE EVIDENCE that it has nothing to look at, never
because the classifier is unsure. Running a reviewer that finds nothing costs
money, which is the status quo. Skipping one that would have found something
costs a bug, and looks exactly like a clean review. Those are not comparable
mistakes, so every uncertain case takes the expensive branch — including a lens
nobody has classified yet, which always runs.

correctness is never skipped: any change to anything can be wrong, and that is
the one reviewer whose absence is a hole rather than a saving. At least one
lens always runs, so "we reviewed nothing" is not an outcome this can produce
quietly.

Decided in prepareCycle, which is both where the file list first exists AND
where runsTotal is written — so the progress denominator counts the passes that
will actually happen. Selecting anywhere later would mean a bar that promised
five steps and delivered four.

The skips are RECORDED with their reasons, and the tab says which reviewers
read the change and how many sat out. A reviewer that never ran finds nothing,
and nothing reads exactly like a clean bill of health, so the absence is stated
rather than left to pass for a result.

A rule rather than a model call, deliberately. A model would judge better and
would also be another unit, another failure mode and another bill before the
review starts — and its decisions could not be shown to the user as a reason
they can check.

The tests are mostly about what must NOT be skipped, which is where the risk
is: correctness under every input, the at-least-one floor, an unclassified
lens, a migration keeping every runtime reviewer, and one source file among
documentation keeping all of them.

0.3.1

Choose a tag to compare

@github-actions github-actions released this 27 Sep 21:17
a1ce898
v0.3.1

feat(code-review): open the detail sheet from the panel, and give it …

0.3.0

Choose a tag to compare

@github-actions github-actions released this 27 Sep 20:12
539e795
feat(code-review): show how each reviewer has actually performed

The settings page could say what a preset promises and nothing about what it
delivered. A lens that raises forty findings and keeps two is not thorough, it
is expensive, and from inside the product there was no way to tell those apart.

Each lens now reports raised versus kept across the workspace's recent
findings, as a bar and a count. Shown only once there is history — "0 of 0"
teaches nobody anything — and fetched once when the page opens rather than
polled, because "is this reviewer worth its cost" is a question you ask
occasionally. A failed fetch is silent: this is context, not a control, and a
settings page that cannot show a statistic must still let somebody change a
setting.

A finding counts once PER LENS that raised it, so two lenses agreeing counts
for both. That is the point rather than a rounding choice: agreement is the
signal the dedupe key exists to preserve, and splitting the credit would punish
a lens for being corroborated.

Aggregated in SQL with a lateral unnest of the jsonb array. The alternative
ships every finding's row to count a string, over a whole workspace's history.

Two things the tests caught that a typecheck could not. `execute` returns a
different SHAPE per driver — postgres-js hands back rows, pglite wraps them in
`{ rows }` — so the existing cast-to-array pattern in this repo works in
production and cannot be covered by the suite that would catch it breaking;
both shapes are accepted now. And an unvalidated finding counts as raised but
not kept, which is what stops every lens looking flattered while a review is
still running.

0.2.111

Choose a tag to compare

@github-actions github-actions released this 27 Sep 18:23
077a63c
v0.2.111

feat(code-review): make the panel usable, and settings a flow rather …

0.2.110

Choose a tag to compare

@github-actions github-actions released this 27 Sep 13:23
b69e4ab
docs: record what the first real code review taught us

A session note and four additions to the code-review concept in claude.md,
because four of the things below were declared in the design and never actually
ran — and nothing in a typecheck or a test was going to say so.

The session note covers the shadowing bug in full: fleetd's terminal event is
its own summary of the RUN, not the agent's words, and preferring it meant the
agent's real final message was never read. True of the fleet and not of the
Agent SDK, which is why the reading looked right.

claude.md gains what a future session would otherwise re-derive or repeat:

- Depth means model tier, the judging passes escalate on every preset, and
  escalation must happen after the model ladder — with the cross-vendor trap
  spelled out, because fleetProviderForModel's anthropic default is correct for
  choosing a credential and dangerous for choosing a model.
- The judge's verdict carries a severity and a reason, and both used to be
  discarded.
- Auto-fix exists, is off, and what bounds it — including that each bound came
  from a finding that was actually wrong.
- Lens names are user-facing now, and why that reverses the original rule
  without contradicting the reasoning behind it.

0.2.109

Choose a tag to compare

@github-actions github-actions released this 27 Sep 12:26
e2e8209
v0.2.109

feat(code-review): fix findings automatically, bounded by what the ev…

0.2.108

Choose a tag to compare

@github-actions github-actions released this 27 Sep 09:58
7c5af54
v0.2.108

fix(code-review): gate the Findings tab on the flag, like every other…