Skip to content

Releases: nearform/lastlight

v0.38.2 — @libsql/client 0.18

Choose a tag to compare

@cliftonc cliftonc released this 02 Oct 07:20
v0.38.2
a1e923f

Highlights

  • @libsql/client 0.17 → 0.18 (#433). 0.18 replaces the local client's connection relay with a real connection pool, and two state-layer pieces are adapted to it:
    • SQLite busy timeout now covers every connection. The write lock's PRAGMA busy_timeout re-arm reached one pooled connection of many, so a write on another failed SQLITE_BUSY at once instead of waiting out another process (lastlight state, a backup). The timeout is now set as the client's timeout option, which libsql applies to every connection the pool opens.
    • The legacy messaging_sessions rebuild runs as one client.migrate(), with its foreign-key check enforced inside the transaction. A violation now rolls the rebuild back instead of being detected after COMMIT.

No schema or config changes.

Packages

  • lastlight 0.38.2 · lastlight-core 0.38.2
  • lastlight-evals 0.18.2
  • lastlight-shared unchanged (0.11.1) · lastlight-code-facts unchanged (0.8.1) · lastlight-workflow-engine unchanged (0.11.0) · agentic-pi unchanged (0.7.0)

Full changelog: v0.38.1...v0.38.2

v0.38.1 — dependency refresh

Choose a tag to compare

@cliftonc cliftonc released this 02 Oct 05:02
v0.38.1
183f472

Highlights

  • Dependency refresh (#432). Folds the open Dependabot PRs (#411–#419) into one change with a single lockfile regeneration.
    • vitest 5 in every package (the CLI was already on it).
    • Runtime: zod 4.6, hono 4.13.9, OpenTelemetry 0.222 / 2.11, react 19.3, astro 7.3, @clack/prompts 1.8, drizzle-orm 0.45.3, @slack/bolt 5.1, @ff-labs/pi-fff 0.11, plus patch bumps.
    • Tooling: vite 8.3, wrangler 4.139, biome 2.5.14, drizzle-kit 0.31.11, pglite 0.5.8, tsx, @types/node.
  • Held back: @libsql/client 0.18. It changes connection-release behaviour under the #421 SQLite write lock, so it will ship in its own PR after a closer look.

No behaviour changes intended.

Packages

  • lastlight 0.38.1 · lastlight-core 0.38.1
  • lastlight-code-facts 0.8.1 · lastlight-shared 0.11.1
  • lastlight-evals 0.18.1
  • lastlight-workflow-engine unchanged (0.11.0; dev-only bump) · agentic-pi npm unchanged (0.7.0; the sandbox image vendors the workspace copy, so it picks up the new deps anyway)

Full changelog: v0.38.0...v0.38.1

v0.38.0 — re-reviews converge

Choose a tag to compare

@cliftonc cliftonc released this 01 Oct 10:48
v0.38.0
a5c157b

Highlights

  • Re-reviews converge (#429, #431). With review.analysis on, pr-review now remembers what it reviewed and re-reviews only what changed.
    • A review ledger per PR. Each review records the units it covered, per-file line hashes, and every finding (posted or withheld). The next review is dispatched with it. Finding status (open / withheld / addressed / resolved) comes only from structured signals — the quoted code gone at the new head, our thread resolved — never from reply text. No table or migration: the ledger rides the run's scratch.
    • Delta-scoped sites. Units get a stable identity and a delta against the last review (new / changed / affected / unchanged). Rows of unchanged units are carried, not investigated, so a small push gets a few sites and a push that changed nothing the last review covered plans none.
    • A per-line convergence gate. A re-review finding whose lines were all already there at the last review is a late discovery: it is withheld as converged unless it is must-fix on non-low-risk code, which posts labelled "Missed in an earlier review". A re-found open finding is withheld as already-raised, and the summary opens with what was addressed and what is still open.
  • Risk tiers (review.risk.rules). First-match globs — a repo's .lastlight/ rules, then the operator's, then built-ins (docs/tests/generated low; migrations/schema/auth/CI high) — raised one tier by a security/state obligation or high fan-in. They weigh site ranking, gate strictness and coverage; they never decide whether a changed unit is surveyed.
  • Review coverage. Each run writes per-unit surveyed vs investigated, risk-weighted, and the in-scope units nobody investigated. The run detail page gets a Review tab showing coverage and the ledger.
  • Evals: chained re-review cases. A pr-review case can list rounds: [{ head_commit }] to replay a PR's review history through the real workflow, carrying the ledger forward. Each round is diffed from its own merge base, and seeded human discussion can carry from_round so an early round never reads the review a later one is graded on. The scorecard and dashboard show late discoveries, converged / already-raised withholds, cumulative gold recall and per-round coverage and cost.

Measured on real skillspro review chains: a push that changed nothing reviewed costs about a third of a first review (~35 s, no investigators); a 1-of-34-unit push investigates one file; across both repeats 1 of 8 later-round comments landed on already-reviewed lines.

Packages

  • lastlight 0.38.0 · lastlight-core 0.38.0
  • lastlight-code-facts 0.8.0 · lastlight-shared 0.11.0
  • lastlight-evals 0.18.0
  • lastlight-workflow-engine unchanged (0.11.0) · agentic-pi unchanged (0.7.0)

Full changelog: v0.37.0...v0.38.0

v0.37.0 — dynamic fan-out

Choose a tag to compare

@cliftonc cliftonc released this 01 Oct 03:20
v0.37.0
7c8c95d

Highlights

  • Fan-outs can take their branch list from run-time data (#423). A type: fanout phase can declare branches_from: { file, max } and a branch: template instead of a fixed branches: list.

    • The template is rendered once per item of a JSON manifest an earlier phase wrote into the workspace, with each item available as {{item.*}}.
    • max caps the spend.
    • An empty manifest is a no-op.
    • A malformed manifest fails that phase, loudly, rather than the run.
  • pr-review runs one site investigator per real site (#423). site-review no longer declares sixteen fixed slots and pre-closes the unused ones. site-plan writes the list of sites it actually formed, so a PR with two sites runs two investigators. With models.review-site-pair set, each site's second investigator is site-00N-b.

  • Compact fan-out view in the dashboard (#423). A fan-out with five or more branches is drawn as one block:

    • status counts and one chip per branch, including chips for branches that haven't started yet;
    • clicking a chip opens that branch;
    • an expand toggle brings back the full cards.

    The workflow diagram labels dynamic fan-outs fan-out · dynamic ≤ N.

  • Evals record the right model for every fan-out branch (#423). Branch results now carry the model template they ran under, so pair investigators are scored on review-site-pair. Stored runs in the old sixteen-slot format still replay.

Packages

  • lastlight 0.37.0 · lastlight-core 0.37.0
  • lastlight-workflow-engine 0.11.0 · lastlight-code-facts 0.7.0 · lastlight-shared 0.10.3
  • lastlight-evals 0.17.0
  • agentic-pi unchanged (0.7.0)

Full changelog: v0.36.0...v0.37.0

v0.36.0 — superseded reviews stop cleanly, kinder re-reviews, range-anchored comments

Choose a tag to compare

@cliftonc cliftonc released this 30 Sep 08:24
v0.36.0
b6eb528

Highlights

  • A superseded review stops cleanly (#426). When a new push supersedes an in-flight pr-review, the cancel is now final. Before, the killed phase's failure could flip the run from cancelled to failed, and the dead run carried on through its remaining phases. A fan-out also stops launching branches, gates and gate re-runs into a killed sandbox, and the phase shows superseded: … instead of a misleading Sandbox agent failed (exit 137).
  • No model call for an empty selection (#426). When the site investigators find nothing, select is skipped: its only correct answer is fixed. This removes the "empty completion" failure it used to hit. Workflow skip_if gains an anchored startsWith(...) expression.
  • Brief but kind review summaries (#427). A clean first review reads "Looks good — no issues to raise." A clean re-review thanks the author and says it's good to merge, but only when nothing from an earlier review is still open: an open ledger point, an already-raised finding, or an unresolved inline thread of ours. An @bot review of an unchanged head isn't treated as a re-review. The default soul gains a "brief, not curt" line.
  • Review comments can highlight a line range (#428). Site investigators can give an optional startLine, so a comment highlights the defect's own stretch of code rather than a single line with unrelated context above it. A range that doesn't fit the diff falls back to exactly the placement it would have had without one, so asking for a range never costs a finding its inline comment.
  • micro-select eval (#424). A new lastlight-evals script replays pr-review's select phase over preserved runs, to compare selection models on identical input, with gold judging and a dashboard view.

Packages

  • lastlight 0.36.0 · lastlight-core 0.36.0
  • lastlight-workflow-engine 0.10.0 · lastlight-code-facts 0.6.0 · lastlight-shared 0.10.2
  • lastlight-evals 0.16.0
  • agentic-pi unchanged (0.7.0)

Full changelog: v0.35.2...v0.36.0

v0.35.2 — resumed runs keep their context

Choose a tag to compare

@cliftonc cliftonc released this 29 Sep 15:54
v0.35.2
34b7c3f

Highlights

  • A resumed run keeps its dispatch's template context (#425). resumeSimpleRun rebuilt only the base template fields and dropped everything the dispatch had rendered onto the run: the PR snapshot, analysisEnabled, triageEnabled, and a comment's commentBody. Every re-entry goes through resume: boot recovery after a restart, admission of a run queued at the concurrency cap, and Retry. So a pr-review on a busy instance, or one in flight during a deploy, skipped all its analysis phases and silently posted the light single-pass review, or failed at post-review with no findings. Resume now restores the persisted context underneath the fields it owns.

Packages

  • lastlight 0.35.2, lastlight-core 0.35.2, lastlight-evals 0.15.2

Full changelog: v0.35.1...v0.35.2

v0.35.1 — SQLite write lock, review supersede, sweep grace window

Choose a tag to compare

@cliftonc cliftonc released this 29 Sep 14:36
v0.35.1
0022b7f

Highlights

  • SQLite write lock (#421). Every SQLite write now holds one in-process lock. Before, a plain write racing an open transaction could fail with SQLITE_BUSY: database is locked, and the next transaction then failed to commit with cannot commit transaction - SQL statements in progress. The worst case was the boot resume sweep re-dispatching several orphaned runs at once, which failed every one of them.
  • A newer commit replaces an in-flight review (#422). When a review of an older commit is still running and the PR's current commit is due a review, the old run is cancelled: its containers are killed, and the new review waits for it to stop before using the shared workspace. A review of the same commit is never replaced, so an @bot review sent mid-flight gets the "already working" reply.
  • The review sweep no longer reviews mid-CI (#422). New review.sweepPendingGraceMinutes (default 60): the 30-minute sweep reviews a PR whose checks are still pending only once they have been pending that long. It is still the safety net for checks that never finish, but it no longer fires a minute after a push.
  • review.placeholderCheck (#422). Set it to false (with postsCheck: true) and the last-light/review check appears only once a review actually starts, with no queued placeholder while CI runs.
  • Workspace guard (#422). A run now refuses to reset a per-PR workspace that another still-running run owns (WorkspaceBusyError). If two runs collide, one fails instead of both.

Packages

  • lastlight 0.35.1, lastlight-core 0.35.1, lastlight-shared 0.10.1, lastlight-evals 0.15.1

Full changelog: v0.35.0...v0.35.1

v0.35.0 — pr-review: units + sites, paired investigators

Choose a tag to compare

@cliftonc cliftonc released this 29 Sep 11:27
v0.35.0
e17034d

Highlights

  • One pr-review analysis path: units + sites. With review.analysis.enabled, the review now runs facts → seed → a unit survey (one bounded model call per changed function or module region, in parallel, cached) → sites → reconcile → post-review. The survey's rows are treated as a signal of where to look: rows cluster into sites, and each top site gets one investigator that reads the code, runs probes, and writes 1–3 grounded findings, or a none it has to back with an executed probe. One select pass then merges duplicates, sets importance and writes the comments. The five-branch agent survey, the adjudicator, dossier and jev-classify are gone.
  • Paired investigators. Set models.review-site-pair and every site gets a second investigator on a different model. Two models miss different things, so the pair finds more than either one, and select merges what both report.
  • Coverage and context. review.analysis.siteTop (1–8, default 5) sets how many sites get an investigator. Test-file sites now fill slots the code sites leave free instead of being skipped. Investigators read the PR's title, description and closed issues as a claim to check. select reads the PR's earlier reviews, threads and comments, so a point someone already raised is recorded rather than posted again.

On the Martian held-out set (18 PRs across 5 repos, 54 known defects, two repeats), judged on the posted review: sites with gpt-6-luna investigators matched 24/19 defects at $0.37/PR, deepseek-v4-flash 26/28 at $0.61, and the luna + deepseek pair 28 in both repeats at $0.70. The earlier Haiku-investigated arm matched 13 of 44. The full record of what was tried is in docs/plans/pr-review-units-sites.md.

Fixes

  • A single fan-out branch that failed (for example a provider's transient 404) no longer marks a posted review as failed. A failed run left the head unassessed, so the review sweep would re-dispatch it and post again. Branches that hit a provider error are also retried once.
  • A light re-review now runs review and posts, instead of posting nothing.
  • On the in-process backends, a timed-out bash phase now kills its whole process group, and bash phases no longer block the event loop.

Deploy notes

Bump deploy.version to v0.35.0 in the overlay. This changes what pr-review runs when review.analysis.enabled is on:

  • Pin the investigator model. models.review-site is unset by default and falls back to models.review-survey. Measured, cheap models beat Haiku here: openai/gpt-6-luna (with variants.site-review: medium) or opencode/deepseek-v4-flash. models.review-select falls back to models.review, so pin it too if review points at a model you don't want writing the comments.
  • Removed keys are ignored with a warning: review.analysis.surveyEngine, reviewEngine, independentReview, adjudicate, jevModel, admit, jevTimeoutSeconds, surveyPasses, and models.review-adjudicate. Delete them from overlays.
  • surveyConcurrency is now siteConcurrency (the old name is still read). With the pair, a PR runs up to 10 investigators, and the docker backend runs at most 6 at once.
  • probes is inert for now. Falsify is not yet attached to the sites engine; the investigators run their own probes.
  • Kubernetes: review.analysis.enabled is refused at startup on the kubernetes backend.

Packages

lastlight 0.35.0 · lastlight-core 0.35.0 · lastlight-evals 0.15.0 · lastlight-code-facts 0.5.0 · lastlight-shared 0.10.0 · lastlight-workflow-engine 0.9.0 · agentic-pi 0.7.0 (unchanged)

Full changelog: v0.34.1...v0.35.0

v0.34.1 — sandbox CPU and memory per phase

Choose a tag to compare

@cliftonc cliftonc released this 27 Sep 08:30
v0.34.1
bbaf663

Highlights

  • Sandbox CPU and memory, per phase and per workflow. Docker and Kubernetes sandboxes now record the CPU time and memory high-water mark of every phase, read from the sandbox's own cgroup. The home page gains a CPU card and a Sandbox CPU time chart, with the day's largest memory peak as a line. Recent Workflows shows each run's CPU (and its peak memory in the tooltip), and phase detail shows CPU Time and Peak Memory against the limit. A fan-out's shared container is recorded once rather than split across its branches. The in-process backends (gondolin / none) report nothing (#409).
  • Token stats count cache writes. The Tokens total and token chart now include cache-write tokens, shown as a cache write bar. Anthropic reports the uncached prompt as a cache write while OpenAI-compatible providers (OpenCode Zen) report it as input, so without them a switch to open models looked like a jump in input tokens. Historical totals rise by the Anthropic cache-write volume. That's a correction, not new usage (#409).

Fixes

  • PR review no longer puts prose in "Apply suggestion" blocks (prompt-level). The adjudicator had lost the definition of suggestion and was filling it with a description of the fix, which pressing Apply would have committed into the source file. The field is now defined everywhere it's written as exact replacement code, or omitted (#410).

Deploy notes

Bump deploy.version to v0.34.1 in the overlay. The release adds three nullable executions columns (additive migration on SQLite and Postgres, applied at boot). No config changes.

On Kubernetes, the usage reading travels on the pod's log stream. A pod that dies early (OOM-kill, deadline) records none, and it's a best-effort metric rather than a security boundary. See spec/09-sandbox.md → Resource usage.

Packages

lastlight 0.34.1 · lastlight-core 0.34.1 · lastlight-evals 0.14.1 · lastlight-shared 0.9.1 · lastlight-workflow-engine 0.8.1 · lastlight-code-facts 0.4.0 · agentic-pi 0.7.0

Full changelog: v0.34.0...v0.34.1

v0.34.0 — PR review on open models, command policy, fewer and better review comments

Choose a tag to compare

@cliftonc cliftonc released this 26 Sep 06:57
fc98c95

Highlights

  • PR review on open models. pi 0.87 and OpenCode Zen (opencode/… models, OPENCODE_API_KEY): the whole pr-review pipeline can now run without an Anthropic key (#402).
  • Command policy per phase. Workflow phases can block or log installs, test runs, and (new) any bash call that reaches outside the agent's workspace (host: whole-disk find, ~/.nvm, global node_modules, PATH pointing outside). pr-review blocks installs and test runs in the phases that must not run code, and blocks host on survey, review and adjudicate (#403, #404).
  • Fewer, better review comments (#405):
    • Impact rule: style, locale, "no tests", dead-code and convention findings are recorded but no longer posted. A finding adjudicate calls a defect is never demoted this way.
    • Verdict hygiene: falsify's reproduced now means the scenario was executed; a grep or file read is corroborated.
    • Derived severity: severity is computed from the evidence, not the adjudicator's label, and the posting caps rank on it, with a tie-break computed in code within each severity.
    • Summary after the caps: the review's opening summary is written after the caps decide what posts, so it never mentions a finding that was held back.
  • Calibration harness for pr-review findings, and the two fixes it found (#398).

On the Martian cal.com set (3 cases × 3 repeats, all-open stack): posted comments per case are down about 20%, precision .220 → .246, cost at or below the previous release.

Fixes

  • {{artifactUrl}} links no longer 400 on docker/kubernetes with buildAssets.location: server (#400).
  • The survey gate rewrites hypothesis ids to their canonical form, which removes a mismatch that sent adjudicate into a second full pass in about half of cases.

Deploy notes

Bump deploy.version to v0.34.0 in the overlay. Config changes that can matter:

  • review.analysis.maxInlineComments default 10 → 5. Combined with maxBodyComments: 0, a review now posts at most 5 comments. Pin maxInlineComments in the overlay to keep the old cap.
  • review.analysis.thresholds and internalFloor are removed. They're accepted and ignored with a warning; delete them from overlays.
  • review.analysis.probes takes three values (off | static | full). true still reads as static, and false or unset as off.
  • New model key models.review-summary (the post-cap summary; one tool-free call). Unset, it uses models.default. Pin it if default points at a provider this deployment has no key for.
  • New review.analysis keys: adjudicate (legacy | dossier | jev, default legacy), jevModel, falsifyTimeoutSeconds (600), jevTimeoutSeconds (120). All have safe defaults.
  • To run pr-review on open models, set OPENCODE_API_KEY in the host .env, then recreate the agent container (lastlight server start agent); a restart doesn't pick up a changed secret.

Packages

lastlight 0.34.0 · lastlight-core 0.34.0 · lastlight-evals 0.14.0 · lastlight-code-facts 0.4.0 · lastlight-shared 0.9.0 · lastlight-workflow-engine 0.8.0 · agentic-pi 0.7.0

Full changelog: v0.33.2...v0.34.0