Releases: Gilbert09/talyn
Releases · Gilbert09/talyn
Release list
0.3.5
0.3.4
v0.3.4 docs: session notes for the review outcome events and the fix commit …
0.3.3
v0.3.3 feat(code-review): re-show the introduction, now that it introduces s…
0.3.2
feat(code-review): only dispatch the reviewers a change actually needs Every lens read every change. A pull request of three documentation files still booted a security sandbox, a reliability sandbox and an operability sandbox, each spending six to ten minutes to report nothing — and a reviewer with nothing to say either returns empty or invents something the judging pass then pays to reject. Files are classified coarsely (code, test, config, migration, docs, asset) and each lens declares which classes it needs at least one of. Coarse on purpose: the finer the classification, the more confidently it is wrong. THE RULE IS DELIBERATELY ASYMMETRIC, and that is the whole safety argument. A lens is skipped only on POSITIVE EVIDENCE that it has nothing to look at, never because the classifier is unsure. Running a reviewer that finds nothing costs money, which is the status quo. Skipping one that would have found something costs a bug, and looks exactly like a clean review. Those are not comparable mistakes, so every uncertain case takes the expensive branch — including a lens nobody has classified yet, which always runs. correctness is never skipped: any change to anything can be wrong, and that is the one reviewer whose absence is a hole rather than a saving. At least one lens always runs, so "we reviewed nothing" is not an outcome this can produce quietly. Decided in prepareCycle, which is both where the file list first exists AND where runsTotal is written — so the progress denominator counts the passes that will actually happen. Selecting anywhere later would mean a bar that promised five steps and delivered four. The skips are RECORDED with their reasons, and the tab says which reviewers read the change and how many sat out. A reviewer that never ran finds nothing, and nothing reads exactly like a clean bill of health, so the absence is stated rather than left to pass for a result. A rule rather than a model call, deliberately. A model would judge better and would also be another unit, another failure mode and another bill before the review starts — and its decisions could not be shown to the user as a reason they can check. The tests are mostly about what must NOT be skipped, which is where the risk is: correctness under every input, the at-least-one floor, an unclassified lens, a migration keeping every runtime reviewer, and one source file among documentation keeping all of them.
0.3.1
v0.3.1 feat(code-review): open the detail sheet from the panel, and give it …
0.3.0
feat(code-review): show how each reviewer has actually performed
The settings page could say what a preset promises and nothing about what it
delivered. A lens that raises forty findings and keeps two is not thorough, it
is expensive, and from inside the product there was no way to tell those apart.
Each lens now reports raised versus kept across the workspace's recent
findings, as a bar and a count. Shown only once there is history — "0 of 0"
teaches nobody anything — and fetched once when the page opens rather than
polled, because "is this reviewer worth its cost" is a question you ask
occasionally. A failed fetch is silent: this is context, not a control, and a
settings page that cannot show a statistic must still let somebody change a
setting.
A finding counts once PER LENS that raised it, so two lenses agreeing counts
for both. That is the point rather than a rounding choice: agreement is the
signal the dedupe key exists to preserve, and splitting the credit would punish
a lens for being corroborated.
Aggregated in SQL with a lateral unnest of the jsonb array. The alternative
ships every finding's row to count a string, over a whole workspace's history.
Two things the tests caught that a typecheck could not. `execute` returns a
different SHAPE per driver — postgres-js hands back rows, pglite wraps them in
`{ rows }` — so the existing cast-to-array pattern in this repo works in
production and cannot be covered by the suite that would catch it breaking;
both shapes are accepted now. And an unvalidated finding counts as raised but
not kept, which is what stops every lens looking flattered while a review is
still running.
0.2.111
v0.2.111 feat(code-review): make the panel usable, and settings a flow rather …
0.2.110
docs: record what the first real code review taught us A session note and four additions to the code-review concept in claude.md, because four of the things below were declared in the design and never actually ran — and nothing in a typecheck or a test was going to say so. The session note covers the shadowing bug in full: fleetd's terminal event is its own summary of the RUN, not the agent's words, and preferring it meant the agent's real final message was never read. True of the fleet and not of the Agent SDK, which is why the reading looked right. claude.md gains what a future session would otherwise re-derive or repeat: - Depth means model tier, the judging passes escalate on every preset, and escalation must happen after the model ladder — with the cross-vendor trap spelled out, because fleetProviderForModel's anthropic default is correct for choosing a credential and dangerous for choosing a model. - The judge's verdict carries a severity and a reason, and both used to be discarded. - Auto-fix exists, is off, and what bounds it — including that each bound came from a finding that was actually wrong. - Lens names are user-facing now, and why that reverses the original rule without contradicting the reasoning behind it.
0.2.109
v0.2.109 feat(code-review): fix findings automatically, bounded by what the ev…
0.2.108
v0.2.108 fix(code-review): gate the Findings tab on the flag, like every other…