Skip to content

Covel v0.0.49

Latest

Choose a tag to compare

@github-actions github-actions released this 08 Oct 16:58
· 267 commits to main since this release
5151f15

[0.0.49] - 2026-10-09

This release fixes session recovery, background-task status, streaming retries, plugin and storage boundaries, and desktop layout issues. SQLite now uses Node's built-in driver while retaining the existing database format. macOS Apple Silicon and Windows x64 binaries are unsigned; the macOS build is not notarized. Read the upgrade notes before installing.

Added

  • Recorded scripted sessions replay in one command, on any machine. pnpm e2e:replay plays each session under tests/llm-replay/ against its recording: it starts the replay proxy, a test server with a new database, its own home directory, the session's random seed and UTC, then scripts/e2e-plugin-verify.ts, and stops them. The server reads nothing from .env, ~/.covel or the shell's COVEL_* variables, so another checkout sends the same requests; any request missing from the recording fails the run. Recordings are committed, one answer per request named after its digest. --record --upstream <origin> records a session again: answers already recorded are reused, only new requests reach the model, and answers the session no longer asks for are removed. pnpm llm:replay stays for ad-hoc recordings under debugs/. The lantern-barrow session is defined but not recorded yet: record it with pnpm e2e:replay --record --upstream https://api.deepseek.com lantern-barrow. Verified with a fake model on a session that stopped at character creation (8 requests, all answered in a copy of the repository at another path, with another HOME, time zone and LANG); not yet verified on a recorded three-turn session. See docs/guide/e2e-plugin-verify.md.

Changed

  • SQLite runs on Node's built-in node:sqlite; better-sqlite3 is gone. The native package needed node-gyp or a matching prebuild on every install, and an ABI match for the desktop sidecar. SqliteStore, the SQLite media store and the sqlite-vec capability now use DatabaseSync through a small Drizzle driver in packages/store/src/sqlite/drizzle-node-sqlite.ts, on Node 26 and on the Node 24 inside Electron 44. Existing covel.db files open unchanged. The unused pg peer that drizzle had resolved left the lockfile with it.
  • The Playwright suite runs against a production build. pnpm e2e builds the web app into dist/web-e2e and serves it with vite preview next to the Vite dev server. A fresh browser context fetched the dev server's several hundred modules on every page, all from one Vite process, so more workers did not make the suite faster. The 23 tests whose specs import the app's own modules in the page (import("/src/...")) still run on the dev server, in the chromium-dev project; the configuration finds them by reading the specs. On one 4-core machine the whole suite went from 6 min 5 s to 4 min 26 s with the default workers, and to about 4 min with 4. The build is made on every run, so a change to the front end needs no manual build. See docs/guide/e2e-testing.md.
  • A plugin outside the host range it declares is no longer loaded. covel: ">=0.0.45 <0.1.0" in PLUGIN.md was checked only when a package was installed, so a package copied into the plugin directory, or installed before the app was updated, still ran on a host it does not support. Startup and reload now apply the same check: the package is listed as failed with both version numbers, and none of its code runs. A package without covel loads as before.
  • pnpm create-plugin writes the host range of a standalone plugin. The generated PLUGIN.md carries covel: ">=<version>" with the version of the checkout that made it, so a host older than the one the plugin was written on refuses the package with both version numbers. --with-tools, which scaffolds a bundled plugin under plugins/, writes no range.
  • Inside the app, the interface language is in Settings only. The top bar of the session and debug pages no longer has a language select; Settings → General → Interface Language is the same setting. The landing page, which has no way into Settings, keeps the select, and so does the menu of a narrow window.
  • The narrators' character sheets show a number attribute with its range. characterSheetSegments takes schema (pass ctx.world.characterSchema); a number attribute whose schema declares min/max renders as "might": "2/5" in the player sheet and the NPC profiles, so the model can tell a maxed strongman from a feeble 5/100. narrator and chat-mode-narrator pass it.
  • The dice-check roller pre-computes the modifier of each ranged number attribute. Without a conversion rule the narrative used the raw value as the modifier (a 2/5 became +2), which made checks fail far more often than the DC table intends. The injected checkContext now lists each of the player character's ranged number attributes with its converted modifier ((value - min) ÷ range × 10: a 2/5 is +4), and the narrative reads the modifier from the list instead of estimating it. Attributes without a declared range are still converted by the narrative.

Fixed

  • A filled opening form no longer repeats a full stop supplied by both the field and its template. Narrative interpolation omits a template . or 。 immediately after a value already ending with that character. Stored player input and punctuation within the value remain unchanged.
  • Completed background tasks no longer accumulate as failed turns below the story. Restoring a worker's model reasoning created a separate foreground task that was then marked interrupted, even when its durable job had committed successfully. The timeline now joins worker traces to the job's source turn using backgroundTurnId, preserves the durable outcome, and deduplicates reasoning across refreshes and either hydration order. Actual job failures remain visible under their source task.
  • A suspension can be resumed again after the server stopped in the middle of resuming it. The claim a resume takes was never released when the process ended first, so every later resume answered 409 and the player could only abandon it. The claim is now released at the next server start (SQLite) or after one hour (PostgreSQL). DataStore gains releaseStaleSuspensionClaims.
  • A settled tabletop check reaches the narrative without row IDs. The check text the narrators read carried the form submission ID and the turn ID of the stored receipt. It now carries the check only: action, attribute, modifier, die, difficulty, total and outcome.
  • A media download drops a redirect's body unread. ctx.media.ingestUrl read the whole body of each 3xx response, which maxBytes did not limit.
  • The model database refresh times out while the body downloads. The 30-second limit ended when the response headers arrived, so a stalled body kept the refresh running.
  • A detached runtime's proposal is checked again after a PreStateCommit hook replaces its payload. The declared-effects check ran only before the hook, so a hook could move the write to a namespace the runtime did not declare.
  • Operator Access says when the browser refused to save or clear the token. The pane showed the change as done and reloaded the app.
  • The dice-check roller resolves a translated attribute name like the rest of the app. An attribute name with no entry for the session locale fell back to the first entry instead of the same language or English.
  • pregame reads the world through ctx.world. It called a store method that only bundled plugins receive.
  • World text. Emberback's Chinese opening no longer has an untranslated word with a wrong gloss; Chihiro Onodera's English character card uses "she" throughout.
  • A concealed runtime's output is no longer readable through the runtime-outputs endpoints. GET /api/sessions/:id/runtime-outputs and /runtime-outputs/:outputId returned the stored row of a runtime that declares io.concealed, with its output text and tool arguments and results, while /turns, traces and the live stream hid them. Both endpoints now return such a row with empty results and with input / output of each tool call set to null.
  • The plugin-data _index endpoint no longer lists hidden world data. It returned the namespace names and keys of _hidden.* data and of kernel bookkeeping, which every other plugin-data read leaves out.
  • Desktop: a plain-HTTP link to a host whose name starts with 127. asks for confirmation. http://127.0.0.1.example.com/ was treated as a loopback address and opened in the system browser without the plain-HTTP dialog. Only 127.0.0.1, localhost and [::1] open directly now.
  • A Latin lorebook key matches when it touches Chinese, Japanese or Korean text. The whole-word rule counted a character of those scripts as a letter, so the key Alice did not match "我去找Alice谈谈" and the entry was left out of the prompt. A character of those scripts now ends a word; art is still not found in start.
  • Keyword recall finds a Latin name written next to Chinese text. "Alice去了北方的灯塔" was indexed as one term, so a search for Alice or 灯塔 did not match the message by term, and the substring fallback compared a lowercased query with the original text. Latin words and runs of Chinese, Japanese or Korean are now separate terms, and the fallback ignores case. The archival search uses the same terms.
  • A retried model call streams its text. Only the first attempt of a streamed call sent its text on, so after a first attempt that failed before any output the narration of the retry appeared all at once when it ended. A retry now streams unless an earlier attempt already sent text.
  • A stream that breaks while the model is still reasoning is retried. Reasoning is never shown to the player, but it counted as output the player had read, so a connection lost during a long thinking phase failed the turn instead of repeating the call. Only text that was shown ends the call now.
  • A background job that ends right after an action no longer stays "running" in the page. The event stream pauses during an action and resumes from the last event ID it received. A stream that had received no event yet (a new session, or right after a reset) had no ID to resume from, so the server replayed nothing and an event emitted in between was lost: the step stayed pending, and a session kept in the browser did not save that result. system.connected and system.reset now carry cursor, the ID of the newest event, and the client resumes from it.
  • Align Memory, SQLite and PostgreSQL store contracts: reject duplicate ordinary message inserts and repeated checkpoint message IDs before replacement; refresh snapshot capture timestamps; and index SQLite turn-result append positions by session, timestamp and sequence. MemoryStore checks message identity through a global index that follows rollback, savepoints and session deletion.
  • A client that stops reading its action stream no longer holds the session lock. Action and subscription SSE writes go through one bounded queue per connection (256 frames, 2 seconds per write and for the closing drain); execution and commit callbacks only enqueue. An overflow, a write error or a missed deadline closes the connection without changing the committed result, and the client recovers through the read-only execution and session endpoints.
  • A provider error that refuses the request is not sent a second time. When a stream failed before any output, the agent loop repeated the call without streaming, also after a 401 or 400 that fails the same way again. The repeat is kept for failures a plain call can get past.
  • A failed model response keeps its token usage in the trace. A response cut at the output limit, or rejected after it arrived, was recorded with zero tokens, so the Debug cost panel undercounted retries.
  • A model that streams a long tool call is not timed out while it writes. Fragments of a tool call's arguments count as output for the first-output and idle limits on the OpenAI Chat, Responses, Anthropic and Gemini wires. The call is still released only when it is whole, and a shown stream that breaks before any text is retried.
  • A session kept in the browser starts with its world's character schema. A session created on the server takes metadata.characterSchema of its world as its schema; one created in the browser-private profile started with none, so the world's preset characters had fields without a schema. The browser now writes the same schema and refuses an invalid preset before it saves the session.
  • MemoryStore keeps a row's identity on update, as SQLite and PostgreSQL do. Updating a state entry gave it the ID of the new record, and updating a world, a lorebook entry or a suspension replaced its creation time (for a suspension also its turn and runtime). The browser-private profile and tests run on MemoryStore, so they saw IDs and times the SQL backends never produce.
  • The scene stage reads its cast through the public plugin API. Its cast and direction runtimes listed characters with a store method that only bundled plugins receive, so the same package installed as a community plugin chose speakers from an empty roster. They now read ctx.world.characters, and no bundled plugin refers to an internal package type.
  • A media write that outlives its runtime is no longer claimed for the session. ctx.media.put and ctx.media.ingestUrl recorded ownership and a reference after the bytes were stored, also when the runtime had timed out or been aborted while they were written. The call now rejects and the bytes stay unowned for media cleanup.
  • The browser's execution-timeline cache stops growing. Each save merged into every row ever saved for the session and rewrote the whole record. The cache now keeps the newest 1000 runtime executions of a session.
  • The request budget counts a response schema twice. The schema is written into the system prompt and, on the Responses and Gemini wires, sent again as a native field; the budget counted one copy.
  • World ZIP activation rejects worldData errors with 400 before creating a record and removes only the new directory, allowing a repaired same-ID retry; warnings and startup seed tolerance remain unchanged.
  • Active EventBus subscriptions preserve replay epochs and sequences beyond the pinned-state eviction budget; released sessions remain eligible for normal LRU/TTL eviction.
  • Character sync overlays only new batch proposals onto world snapshots, preserving same-batch deduplication and sequential updates. Missing-character updates fail as tools, preventing partial function-runtime commits and false agent completion. Function tool calls preserve the execution's frozen logicalTurn.
  • Unexpected resume, snapshot and fork errors use standard production sanitization after cleanup; resume lock timeouts return 503 session_busy. A resume whose runtime or commit failed still says why, as a turn does, without returning the runtime result. World translation streams sanitize unexpected errors without changing their open HTTP status, and still report why a translation was refused.
  • A revised world stays in its own language. A revision removes the other editions of a world and writes no locale files, but the manifest the model returned could declare more editions (a request such as "add an English edition" does this), and a session in that language then got the original text. The revised manifest is now set back to the world's own edition.
  • Request-body limits return standard 413 envelopes, not 500, for declared-length and streaming requests at the existing 1/20/64 MiB limits.
  • Development reload recomputes plugin layout diagnostics so repaired quarantined community packages recover static declarations without importing unapproved entry, handler or guard modules.
  • Plugin CLI/ZIP validation checks committed input schema files and same-package producers and rejects committed+select before installation writes; normal turn select and optional external contracts remain supported.
  • Theme application, import and save-as await persistence without clearing overrides or names or reporting false success on failure. Failed appearance edits, resets and ignored token submissions restore saved previews while protecting later edits and scheme switches.
  • Plugin catalog switches have accessible names and native keyboard activation without submitting forms; inputs, textareas and selects associate localized labels with stable unique IDs while retaining the panel interaction lock.
  • English and Russian model-role connectivity hints direct players to the current Providers & Models page instead of the retired API Keys pane.
  • Disabled Lorebook entries are excluded from archival keyword/vector retrieval as well as automatic prompt injection, including existing vectors before ingestion runs. Re-enabling restores search; ingestion removes disabled indexes and rebuilds them when enabled again without affecting other owners or characters.
  • Desktop key saving while the sidecar is unavailable now writes a private temporary file and atomically replaces the saved keys; write, permission-setting and rename failures preserve the existing keys.
  • Execution status reads reject a session deleted or recreated during an authorized lookup, rather than returning the replacement session’s retry input; polling remains non-blocking.
  • Newly generated worlds declare only the requested content-language edition. Generation and revision validate the complete character roster against the world’s character schema before accepting it, preserving supplemental retry/discard and whole-revision failure behavior. Existing revision language declarations are not clamped.
  • Read-only session recovery replaces terminal streaming fragments with durable narrative, bridges non-overlapping message windows without dropping loaded history, and prevents stale plugin catalogue reads from overwriting a newer visit or observation. Healthy POST streams remain authoritative.
  • A session reloaded in Stage keeps its world, its view and its decision. After a reload the session could be restored before the world list had loaded, and it then had no world: no world name or art, and a world whose default view is Stage opened as text. The session now takes its world from the list when the list arrives. The view a world asks for also applies when the world arrives after the page, until the player picks a view. And a story that the restore delivers after Stage is on screen is no longer replayed in the dialog with the decision hidden until the player clicks through it: the restore counts as running a turn until it has checked the session, and a story that arrives then is history. Each depended on the page loading before the restore finished, which the production build does (the desktop app, a served build) and the dev server rarely did. Found by running the browser tests against the production build.
  • The Windows window buttons no longer cover the right end of the top bar. The desktop shell draws minimize, maximize and close over the top-right corner of the page, and the top bar reserved no room for them, so the control at its right end (the language select) was hidden. The top bar now leaves that width free. A layout with the icon rail has no top bar, so the first row of each page was under the buttons; on Windows it now gets a strip of the buttons' height above the page.
  • The Windows window buttons take the colours of the page under them. They were drawn on near-black with light symbols in every theme, a dark box in the corner of a light one. The app now reports the background and text colour of the top bar (of the page, in a layout with the icon rail) when the theme or the colour scheme changes, and the shell draws the buttons in them. Over the startup screen the buttons follow the system colour scheme, as that screen does.
  • The schemas of story-events and scene-stage no longer make Ajv log strict mode warnings in the server log and in test output. They accept the same values as before.
  • A corrected agent is required to call a tool. A runtime that must use a tool (completion.require: tool-use, or a completing tool) and answered in prose got one correction, but the next request still left the choice to the model, which answered in prose again and failed. The corrected request now sends toolChoice: "required" (a tool the manifest names stays the choice; a thinking model keeps automatic choice), and the choice is the model's again once it calls a tool.
  • A rejected tool call names the fields in its error line. Invalid parameters for tool …: Invalid input: expected array, received undefined did not say which field; the line now reads fields: Invalid input: …, as the details already did, and a model fixes the call on its next step instead of repeating it.
  • Reasoning a model writes inline as <think>…</think> no longer reaches the reply. Qwen3, DeepSeek-R1 and their distills reason in the reply itself, and a gateway may pass that through as content; a story then showed the reasoning to the player and a JSON reply did not parse. The Chat and Responses adapters move those blocks to the reasoning channel, streamed replies included, and do not send them back on the next request.
  • A bookkeeping agent retries a model call that broke off mid-stream or ran into the output limit. Every agent runtime streams, and a stream that failed after its first output was never retried, so maxRetries: 3 on guide, affinity and the other bookkeeping agents covered only failures before any output. Only a story runtime shows its stream to the player; every other runtime now drops the partial output and retries. A response cut at the output limit is retried once: a model that looped, or one that echoed its tool call into the reply as escaped text (grok-4.7-build-fast behind some gateways), usually finishes on a fresh attempt, and a second cut fails with the message to raise the limit. The time a failed stream spent writing is credited back, as a finished stream's already was, so the retry still has time to run.
  • The memory extraction of a turn is no longer lost when the player answers quickly. A background job's maxExecutionMs also counted the time its finished result waited for the session lock to commit. When a job outlasted maxSettleWaitMs, the next player turn went ahead and held that lock for its whole run, so a memory/extract that finished in 90 seconds was timed out at 120 and not retried. The limit now ends when the runtime returns; a job that was cancelled or timed out is still refused at commit. The log no longer warns completion follow-up failed for a job that timed out.
  • pnpm e2e:verify runs without .env and .env.llm. It required both files although the harness reads no provider key (the server does), so it could not start where the keys are environment variables. It now loads each file only when present, like pnpm i18n. The plugin-testing guide's example passed --plugins, which the harness rejects; it now passes --plugin.
  • A missing user worlds directory takes one line in the server log. On a fresh install $COVEL_HOME/worlds does not exist until the first world is installed, and the seeder and the file watcher each printed an ENOENT stack trace for it. The directory still counts as an unscanned source, so stale worlds are still kept until a complete scan.

Upgrade notes for v0.0.49

  • Upgrading from v0.0.48 does not require deleting SQLite data. Existing covel.db files remain readable after the switch to node:sqlite; this release adds no data migration. Keep a backup before upgrading. Upgrades from versions before v0.0.48 still need the v0.0.48 breaking-change notes, and versions before v0.0.42 also need the v0.0.42 guide.
  • Custom storage backends must implement DataStore.releaseStaleSuspensionClaims(olderThanIso): Promise<number>. Built-in backends already release abandoned resume claims according to their startup and expiry policy.
  • Plugin host ranges are enforced at startup and reload. A plugin whose declared covel range excludes 0.0.49 will not run; authors should test the package and update its range when supported. Plugins without a range keep their existing behavior.
  • Desktop downloads are unsigned. macOS builds are also not notarized, so the operating system may require first-launch confirmation. There is no signing or notarization claim for these artifacts.

What's Changed

  • fix(plugins): remove Ajv strict-mode warnings from two plugin schemas by @ackness in #161
  • test: remove Ajv strict-mode warnings from schema fixtures by @ackness in #162
  • fix(desktop): keep page controls clear of the Windows window buttons and draw them in the page's colours by @ackness in #163
  • feat(plugin-sdk): show attribute ranges in character sheets; fix(dice-check): pre-compute modifiers from ranges by @charlesli1989 in #164
  • docs(skills): add the read-only whole-repo audit skill by @ackness in #165
  • chore: gate Oxlint and strict index access, enforce the plugin host range at load, add a security policy by @ackness in #166
  • test: give every Vitest run a temp directory of its own by @ackness in #167
  • ci: shorten the path from push to a green CI run by @ackness in #168
  • Run SQLite on node:sqlite and make agent model calls survive unreliable models by @ackness in #169
  • Replay recorded sessions in one command, run Playwright on a production build, fix Stage after reload by @ackness in #170
  • fix: land the reviewed fixes from the whole-repo audit by @ackness in #171
  • fix: nine issues from the 2026-10-08 read-only audit by @ackness in #172
  • fix: ten issues from the second 2026-10-08 read-only audit by @ackness in #173
  • fix: resolve pre-release timeline and form display defects by @ackness in #174
  • chore(release): prepare v0.0.49 by @ackness in #175

Full Changelog: v0.0.48...v0.0.49