Releases: nearform/lastlight
Release list
v0.38.2 — @libsql/client 0.18
Highlights
- @libsql/client 0.17 → 0.18 (#433). 0.18 replaces the local client's connection relay with a real connection pool, and two state-layer pieces are adapted to it:
- SQLite busy timeout now covers every connection. The write lock's
PRAGMA busy_timeoutre-arm reached one pooled connection of many, so a write on another failedSQLITE_BUSYat once instead of waiting out another process (lastlight state, a backup). The timeout is now set as the client'stimeoutoption, which libsql applies to every connection the pool opens. - The legacy
messaging_sessionsrebuild runs as oneclient.migrate(), with its foreign-key check enforced inside the transaction. A violation now rolls the rebuild back instead of being detected after COMMIT.
- SQLite busy timeout now covers every connection. The write lock's
No schema or config changes.
Packages
lastlight0.38.2 ·lastlight-core0.38.2lastlight-evals0.18.2lastlight-sharedunchanged (0.11.1) ·lastlight-code-factsunchanged (0.8.1) ·lastlight-workflow-engineunchanged (0.11.0) ·agentic-piunchanged (0.7.0)
Full changelog: v0.38.1...v0.38.2
v0.38.1 — dependency refresh
Highlights
- Dependency refresh (#432). Folds the open Dependabot PRs (#411–#419) into one change with a single lockfile regeneration.
- vitest 5 in every package (the CLI was already on it).
- Runtime: zod 4.6, hono 4.13.9, OpenTelemetry 0.222 / 2.11, react 19.3, astro 7.3, @clack/prompts 1.8, drizzle-orm 0.45.3, @slack/bolt 5.1, @ff-labs/pi-fff 0.11, plus patch bumps.
- Tooling: vite 8.3, wrangler 4.139, biome 2.5.14, drizzle-kit 0.31.11, pglite 0.5.8, tsx, @types/node.
- Held back:
@libsql/client0.18. It changes connection-release behaviour under the #421 SQLite write lock, so it will ship in its own PR after a closer look.
No behaviour changes intended.
Packages
lastlight0.38.1 ·lastlight-core0.38.1lastlight-code-facts0.8.1 ·lastlight-shared0.11.1lastlight-evals0.18.1lastlight-workflow-engineunchanged (0.11.0; dev-only bump) ·agentic-pinpm unchanged (0.7.0; the sandbox image vendors the workspace copy, so it picks up the new deps anyway)
Full changelog: v0.38.0...v0.38.1
v0.38.0 — re-reviews converge
Highlights
- Re-reviews converge (#429, #431). With
review.analysison, pr-review now remembers what it reviewed and re-reviews only what changed.- A review ledger per PR. Each review records the units it covered, per-file line hashes, and every finding (posted or withheld). The next review is dispatched with it. Finding status (
open/withheld/addressed/resolved) comes only from structured signals — the quoted code gone at the new head, our thread resolved — never from reply text. No table or migration: the ledger rides the run's scratch. - Delta-scoped sites. Units get a stable identity and a delta against the last review (
new/changed/affected/unchanged). Rows of unchanged units are carried, not investigated, so a small push gets a few sites and a push that changed nothing the last review covered plans none. - A per-line convergence gate. A re-review finding whose lines were all already there at the last review is a late discovery: it is withheld as
convergedunless it is must-fix on non-low-risk code, which posts labelled "Missed in an earlier review". A re-found open finding is withheld asalready-raised, and the summary opens with what was addressed and what is still open.
- A review ledger per PR. Each review records the units it covered, per-file line hashes, and every finding (posted or withheld). The next review is dispatched with it. Finding status (
- Risk tiers (
review.risk.rules). First-match globs — a repo's.lastlight/rules, then the operator's, then built-ins (docs/tests/generated low; migrations/schema/auth/CI high) — raised one tier by a security/state obligation or high fan-in. They weigh site ranking, gate strictness and coverage; they never decide whether a changed unit is surveyed. - Review coverage. Each run writes per-unit surveyed vs investigated, risk-weighted, and the in-scope units nobody investigated. The run detail page gets a Review tab showing coverage and the ledger.
- Evals: chained re-review cases. A pr-review case can list
rounds: [{ head_commit }]to replay a PR's review history through the real workflow, carrying the ledger forward. Each round is diffed from its own merge base, and seeded human discussion can carryfrom_roundso an early round never reads the review a later one is graded on. The scorecard and dashboard show late discoveries, converged / already-raised withholds, cumulative gold recall and per-round coverage and cost.
Measured on real skillspro review chains: a push that changed nothing reviewed costs about a third of a first review (~35 s, no investigators); a 1-of-34-unit push investigates one file; across both repeats 1 of 8 later-round comments landed on already-reviewed lines.
Packages
lastlight0.38.0 ·lastlight-core0.38.0lastlight-code-facts0.8.0 ·lastlight-shared0.11.0lastlight-evals0.18.0lastlight-workflow-engineunchanged (0.11.0) ·agentic-piunchanged (0.7.0)
Full changelog: v0.37.0...v0.38.0
v0.37.0 — dynamic fan-out
Highlights
-
Fan-outs can take their branch list from run-time data (#423). A
type: fanoutphase can declarebranches_from: { file, max }and abranch:template instead of a fixedbranches:list.- The template is rendered once per item of a JSON manifest an earlier phase wrote into the workspace, with each item available as
{{item.*}}. maxcaps the spend.- An empty manifest is a no-op.
- A malformed manifest fails that phase, loudly, rather than the run.
- The template is rendered once per item of a JSON manifest an earlier phase wrote into the workspace, with each item available as
-
pr-review runs one site investigator per real site (#423).
site-reviewno longer declares sixteen fixed slots and pre-closes the unused ones.site-planwrites the list of sites it actually formed, so a PR with two sites runs two investigators. Withmodels.review-site-pairset, each site's second investigator issite-00N-b. -
Compact fan-out view in the dashboard (#423). A fan-out with five or more branches is drawn as one block:
- status counts and one chip per branch, including chips for branches that haven't started yet;
- clicking a chip opens that branch;
- an expand toggle brings back the full cards.
The workflow diagram labels dynamic fan-outs
fan-out · dynamic ≤ N. -
Evals record the right model for every fan-out branch (#423). Branch results now carry the model template they ran under, so pair investigators are scored on
review-site-pair. Stored runs in the old sixteen-slot format still replay.
Packages
lastlight0.37.0 ·lastlight-core0.37.0lastlight-workflow-engine0.11.0 ·lastlight-code-facts0.7.0 ·lastlight-shared0.10.3lastlight-evals0.17.0agentic-piunchanged (0.7.0)
Full changelog: v0.36.0...v0.37.0
v0.36.0 — superseded reviews stop cleanly, kinder re-reviews, range-anchored comments
Highlights
- A superseded review stops cleanly (#426). When a new push supersedes an in-flight pr-review, the cancel is now final. Before, the killed phase's failure could flip the run from
cancelledtofailed, and the dead run carried on through its remaining phases. A fan-out also stops launching branches, gates and gate re-runs into a killed sandbox, and the phase showssuperseded: …instead of a misleadingSandbox agent failed (exit 137). - No model call for an empty selection (#426). When the site investigators find nothing,
selectis skipped: its only correct answer is fixed. This removes the "empty completion" failure it used to hit. Workflowskip_ifgains an anchoredstartsWith(...)expression. - Brief but kind review summaries (#427). A clean first review reads "Looks good — no issues to raise." A clean re-review thanks the author and says it's good to merge, but only when nothing from an earlier review is still open: an open ledger point, an already-raised finding, or an unresolved inline thread of ours. An
@bot reviewof an unchanged head isn't treated as a re-review. The default soul gains a "brief, not curt" line. - Review comments can highlight a line range (#428). Site investigators can give an optional
startLine, so a comment highlights the defect's own stretch of code rather than a single line with unrelated context above it. A range that doesn't fit the diff falls back to exactly the placement it would have had without one, so asking for a range never costs a finding its inline comment. - micro-select eval (#424). A new
lastlight-evalsscript replays pr-review'sselectphase over preserved runs, to compare selection models on identical input, with gold judging and a dashboard view.
Packages
lastlight0.36.0 ·lastlight-core0.36.0lastlight-workflow-engine0.10.0 ·lastlight-code-facts0.6.0 ·lastlight-shared0.10.2lastlight-evals0.16.0agentic-piunchanged (0.7.0)
Full changelog: v0.35.2...v0.36.0
v0.35.2 — resumed runs keep their context
Highlights
- A resumed run keeps its dispatch's template context (#425).
resumeSimpleRunrebuilt only the base template fields and dropped everything the dispatch had rendered onto the run: the PR snapshot,analysisEnabled,triageEnabled, and a comment'scommentBody. Every re-entry goes through resume: boot recovery after a restart, admission of a run queued at the concurrency cap, and Retry. So a pr-review on a busy instance, or one in flight during a deploy, skipped all its analysis phases and silently posted the light single-pass review, or failed atpost-reviewwith no findings. Resume now restores the persisted context underneath the fields it owns.
Packages
lastlight0.35.2,lastlight-core0.35.2,lastlight-evals0.15.2
Full changelog: v0.35.1...v0.35.2
v0.35.1 — SQLite write lock, review supersede, sweep grace window
Highlights
- SQLite write lock (#421). Every SQLite write now holds one in-process lock. Before, a plain write racing an open transaction could fail with
SQLITE_BUSY: database is locked, and the next transaction then failed to commit withcannot commit transaction - SQL statements in progress. The worst case was the boot resume sweep re-dispatching several orphaned runs at once, which failed every one of them. - A newer commit replaces an in-flight review (#422). When a review of an older commit is still running and the PR's current commit is due a review, the old run is cancelled: its containers are killed, and the new review waits for it to stop before using the shared workspace. A review of the same commit is never replaced, so an
@bot reviewsent mid-flight gets the "already working" reply. - The review sweep no longer reviews mid-CI (#422). New
review.sweepPendingGraceMinutes(default 60): the 30-minute sweep reviews a PR whose checks are still pending only once they have been pending that long. It is still the safety net for checks that never finish, but it no longer fires a minute after a push. review.placeholderCheck(#422). Set it tofalse(withpostsCheck: true) and thelast-light/reviewcheck appears only once a review actually starts, with noqueuedplaceholder while CI runs.- Workspace guard (#422). A run now refuses to reset a per-PR workspace that another still-running run owns (
WorkspaceBusyError). If two runs collide, one fails instead of both.
Packages
lastlight0.35.1,lastlight-core0.35.1,lastlight-shared0.10.1,lastlight-evals0.15.1
Full changelog: v0.35.0...v0.35.1
v0.35.0 — pr-review: units + sites, paired investigators
Highlights
- One pr-review analysis path: units + sites. With
review.analysis.enabled, the review now runsfacts→seed→ a unit survey (one bounded model call per changed function or module region, in parallel, cached) → sites →reconcile→post-review. The survey's rows are treated as a signal of where to look: rows cluster into sites, and each top site gets one investigator that reads the code, runs probes, and writes 1–3 grounded findings, or anoneit has to back with an executed probe. Oneselectpass then merges duplicates, sets importance and writes the comments. The five-branch agent survey, the adjudicator,dossierandjev-classifyare gone. - Paired investigators. Set
models.review-site-pairand every site gets a second investigator on a different model. Two models miss different things, so the pair finds more than either one, andselectmerges what both report. - Coverage and context.
review.analysis.siteTop(1–8, default 5) sets how many sites get an investigator. Test-file sites now fill slots the code sites leave free instead of being skipped. Investigators read the PR's title, description and closed issues as a claim to check.selectreads the PR's earlier reviews, threads and comments, so a point someone already raised is recorded rather than posted again.
On the Martian held-out set (18 PRs across 5 repos, 54 known defects, two repeats), judged on the posted review: sites with gpt-6-luna investigators matched 24/19 defects at $0.37/PR, deepseek-v4-flash 26/28 at $0.61, and the luna + deepseek pair 28 in both repeats at $0.70. The earlier Haiku-investigated arm matched 13 of 44. The full record of what was tried is in docs/plans/pr-review-units-sites.md.
Fixes
- A single fan-out branch that failed (for example a provider's transient 404) no longer marks a posted review as failed. A failed run left the head unassessed, so the review sweep would re-dispatch it and post again. Branches that hit a provider error are also retried once.
- A light re-review now runs
reviewand posts, instead of posting nothing. - On the in-process backends, a timed-out bash phase now kills its whole process group, and bash phases no longer block the event loop.
Deploy notes
Bump deploy.version to v0.35.0 in the overlay. This changes what pr-review runs when review.analysis.enabled is on:
- Pin the investigator model.
models.review-siteis unset by default and falls back tomodels.review-survey. Measured, cheap models beat Haiku here:openai/gpt-6-luna(withvariants.site-review: medium) oropencode/deepseek-v4-flash.models.review-selectfalls back tomodels.review, so pin it too ifreviewpoints at a model you don't want writing the comments. - Removed keys are ignored with a warning:
review.analysis.surveyEngine,reviewEngine,independentReview,adjudicate,jevModel,admit,jevTimeoutSeconds,surveyPasses, andmodels.review-adjudicate. Delete them from overlays. surveyConcurrencyis nowsiteConcurrency(the old name is still read). With the pair, a PR runs up to 10 investigators, and thedockerbackend runs at most 6 at once.probesis inert for now. Falsify is not yet attached to the sites engine; the investigators run their own probes.- Kubernetes:
review.analysis.enabledis refused at startup on thekubernetesbackend.
Packages
lastlight 0.35.0 · lastlight-core 0.35.0 · lastlight-evals 0.15.0 · lastlight-code-facts 0.5.0 · lastlight-shared 0.10.0 · lastlight-workflow-engine 0.9.0 · agentic-pi 0.7.0 (unchanged)
Full changelog: v0.34.1...v0.35.0
v0.34.1 — sandbox CPU and memory per phase
Highlights
- Sandbox CPU and memory, per phase and per workflow. Docker and Kubernetes sandboxes now record the CPU time and memory high-water mark of every phase, read from the sandbox's own cgroup. The home page gains a CPU card and a Sandbox CPU time chart, with the day's largest memory peak as a line. Recent Workflows shows each run's CPU (and its peak memory in the tooltip), and phase detail shows CPU Time and Peak Memory against the limit. A fan-out's shared container is recorded once rather than split across its branches. The in-process backends (gondolin / none) report nothing (#409).
- Token stats count cache writes. The Tokens total and token chart now include cache-write tokens, shown as a cache write bar. Anthropic reports the uncached prompt as a cache write while OpenAI-compatible providers (OpenCode Zen) report it as input, so without them a switch to open models looked like a jump in input tokens. Historical totals rise by the Anthropic cache-write volume. That's a correction, not new usage (#409).
Fixes
- PR review no longer puts prose in "Apply suggestion" blocks (prompt-level). The adjudicator had lost the definition of
suggestionand was filling it with a description of the fix, which pressing Apply would have committed into the source file. The field is now defined everywhere it's written as exact replacement code, or omitted (#410).
Deploy notes
Bump deploy.version to v0.34.1 in the overlay. The release adds three nullable executions columns (additive migration on SQLite and Postgres, applied at boot). No config changes.
On Kubernetes, the usage reading travels on the pod's log stream. A pod that dies early (OOM-kill, deadline) records none, and it's a best-effort metric rather than a security boundary. See spec/09-sandbox.md → Resource usage.
Packages
lastlight 0.34.1 · lastlight-core 0.34.1 · lastlight-evals 0.14.1 · lastlight-shared 0.9.1 · lastlight-workflow-engine 0.8.1 · lastlight-code-facts 0.4.0 · agentic-pi 0.7.0
Full changelog: v0.34.0...v0.34.1
v0.34.0 — PR review on open models, command policy, fewer and better review comments
Highlights
- PR review on open models. pi 0.87 and OpenCode Zen (
opencode/…models,OPENCODE_API_KEY): the whole pr-review pipeline can now run without an Anthropic key (#402). - Command policy per phase. Workflow phases can
blockorloginstalls, test runs, and (new) any bash call that reaches outside the agent's workspace (host: whole-diskfind,~/.nvm, globalnode_modules,PATHpointing outside).pr-reviewblocks installs and test runs in the phases that must not run code, and blockshoston survey, review and adjudicate (#403, #404). - Fewer, better review comments (#405):
- Impact rule: style, locale, "no tests", dead-code and convention findings are recorded but no longer posted. A finding adjudicate calls a
defectis never demoted this way. - Verdict hygiene: falsify's
reproducednow means the scenario was executed; a grep or file read iscorroborated. - Derived severity: severity is computed from the evidence, not the adjudicator's label, and the posting caps rank on it, with a tie-break computed in code within each severity.
- Summary after the caps: the review's opening summary is written after the caps decide what posts, so it never mentions a finding that was held back.
- Impact rule: style, locale, "no tests", dead-code and convention findings are recorded but no longer posted. A finding adjudicate calls a
- Calibration harness for pr-review findings, and the two fixes it found (#398).
On the Martian cal.com set (3 cases × 3 repeats, all-open stack): posted comments per case are down about 20%, precision .220 → .246, cost at or below the previous release.
Fixes
{{artifactUrl}}links no longer 400 on docker/kubernetes withbuildAssets.location: server(#400).- The survey gate rewrites hypothesis ids to their canonical form, which removes a mismatch that sent adjudicate into a second full pass in about half of cases.
Deploy notes
Bump deploy.version to v0.34.0 in the overlay. Config changes that can matter:
review.analysis.maxInlineCommentsdefault 10 → 5. Combined withmaxBodyComments: 0, a review now posts at most 5 comments. PinmaxInlineCommentsin the overlay to keep the old cap.review.analysis.thresholdsandinternalFloorare removed. They're accepted and ignored with a warning; delete them from overlays.review.analysis.probestakes three values (off|static|full).truestill reads asstatic, andfalseor unset asoff.- New model key
models.review-summary(the post-cap summary; one tool-free call). Unset, it usesmodels.default. Pin it ifdefaultpoints at a provider this deployment has no key for. - New
review.analysiskeys:adjudicate(legacy|dossier|jev, defaultlegacy),jevModel,falsifyTimeoutSeconds(600),jevTimeoutSeconds(120). All have safe defaults. - To run pr-review on open models, set
OPENCODE_API_KEYin the host.env, then recreate the agent container (lastlight server start agent); a restart doesn't pick up a changed secret.
Packages
lastlight 0.34.0 · lastlight-core 0.34.0 · lastlight-evals 0.14.0 · lastlight-code-facts 0.4.0 · lastlight-shared 0.9.0 · lastlight-workflow-engine 0.8.0 · agentic-pi 0.7.0
Full changelog: v0.33.2...v0.34.0