v0.5.31
0.5.31 carries 17 merged pull requests. Six of them do one thing: a managed agent turn against a model you host yourself now finishes. Six changes close the separate ways such a turn used to be re-sent, abandoned, un-cached or silently settled as "done". Beside that, chat gains an organization-wide Audio transcription model setting and stops warning about audio nobody asked it to transcribe, a UI evaluation round's findings are fixed across dialogs, settings and governance, two paginated lists stop hiding rows the caller can read, Approve can ask before it decides something outside the task, and an administrator can delegate the notification export without handing out an admin seat. No migration, no image in the stop-gated tier: the upgrade is tale update followed by a plain tale deploy.
Highlights
A managed agent turn against a self-hosted model finishes (#3387, #3396, #3397, #3398, #3400, #3401)
A Claude Code or Codex turn is run by a real CLI inside the sandbox, talking to Tale's LLM gateway. Both ends carry their own patience, their own idea of how big the model's context is, and their own retry rule — and none of them had been told what the other was doing. On a fast hosted model that never showed. On a model you host yourself, whose prefill can take minutes, every one of those mismatches turned into a turn that ran twice, ran forever, or ended as a success that changed nothing. Six changes close them, each with the measurement that found it.
-
Two ends, one budget (#3387). The
claude-codeharness set noCLAUDE_STREAM_IDLE_TIMEOUT_MS, so the pinned CLI gave up on a stream after its own 300-second default without a chunk and sent the same request again while the gateway — which waitsSANDBOX_LLM_GATEWAY_STREAM_IDLE_TIMEOUT_SECONDS, 600 seconds by default — was still serving the first one. Any model whose prefill outlasts 300 seconds ran the turn twice, visible as two identical runs on the same task.gatewayStreamIdleTimeoutSeconds()is now the single reader of that budget, feeding both the gateway's provider config and the harness credential, and Codex gets the same number throughmodel_providers.tale.stream_idle_timeout_ms— it runs its own idle watchdog and resends on the same rule. BYO-key runs keep the CLI defaults. -
The wait before the first byte (#3396). An idle-stream budget covers the gap between chunks, not the gap before the first one, and a slow prefill sends nothing at all until its first token. Two further bounds apply there: Bun's own fetch timeout, about 300 seconds, which the CLI lifts only for
api.anthropic.comor whenAPI_FORCE_IDLE_TIMEOUTis explicitly false; and the SDK's request timeoutAPI_TIMEOUT_MS, 600 seconds by default. Past either one the CLI re-sent the request while the gateway kept serving the abandoned copy — its server does not cancel upstream when the client leaves. Against a model that serves one request at a time the copies queue behind each other and each resend waits behind them: a livelock, not a slow turn. Seen live on two deployments — one turn's request re-sent at 16:20:48, 16:25:49, 16:30:51 and 16:35:56, the original answering 200 after 1,298 seconds to a client that had long gone. The managed env now setsAPI_FORCE_IDLE_TIMEOUT=0and tiesAPI_TIMEOUT_MSto the gateway's budget, so the gateway is the single bound on a silent upstream. In a controlled run of the real CLI against an endpoint that stays silent for 420 seconds, the default env re-sent at t=303.7 s and never finished; with the fix, one request, answered at t=423 s,result: success. -
The gateway's own ceiling (#3397). Raising the idle budget was not enough, because the gateway's per-request timeout was pinned at 600 seconds. Its streaming client has no read timeout, but when a stream ends before its first event Claude Code falls back to a non-streaming request that carries the whole prefill — and that one is bounded by the request timeout. On a deployment whose operator had raised the idle budget to 1800 seconds, a restarted model server produced a 504 after exactly 600,049 ms. The gateway's request timeout is now
max(600, the stream idle budget), the turn gets it asAPI_TIMEOUT_MS, and a lowered idle budget still leaves the request timeout at its 600-second floor. -
Compacting inside the window the model actually serves (#3398). A managed Claude Code turn never learned its model's context window. For a model it does not know, the CLI assumes 200,000 tokens and compacts only near 167,000 prompt tokens. On a self-hosted model whose catalog says 32,768, a production turn grew to about 140,000 tokens; its cold prefill then outlasted the CLI's own 30-minute stream watchdog — which the CLI caps there whatever the environment says — so the turn could never be answered and looped until it was cancelled. A new harness placeholder
${model.contextWindow}resolves the turn model's effective window (the connector catalog's window, narrowed by the organization's context limit) exactly as the chat lane does, and the harness passes it asCLAUDE_CODE_AUTO_COMPACT_WINDOW, gated to windows below 200,000 so a Claude model's own larger window is never cut down. All five managed exec builds resolve it: automation start, answered-ask resume, task start including a fresh relaunch, and task steer restart. When the window is unknown the value is empty, which the CLI ignores; a turn never fails over it. -
A foreign model's prompt cache survives the turn (#3400). The pinned CLI opens every system prompt with an attribution line whose checksum changes on every request. Anthropic's API reads that line; any other server sees plain text near the start of the prompt, so a local model's prefix cache ends right there. Measured on a split model behind a managed desk turn: every turn reused exactly the 16,276 tokens of tool definitions the chat template renders first and then prefilled the whole conversation again — 74 to 93 seconds per turn at about 28K tokens, extrapolating to roughly 15 minutes per turn near 100K. Diffing consecutive requests in the gateway's log store showed the only difference was that line, at character 79 of 85. For a model that is not Claude, the exec now also sets
CLAUDE_CODE_ATTRIBUTION_HEADER=0, in the same branch that already turns thinking off for such models. Claude models keep the line, and so does the subscription lane, which only serves Claude. -
A turn that answered nothing is a failure (#3401). A serving cluster that fails mid-prefill can answer a model call with an empty 200 — stream open,
end_turn, stream closed, zero usage. The CLI then ends the turn cleanly with no assistant message, and the classifier, which only failed on an error or a missing turn end, settled the nodeokwith empty text: the workflow went on as though the agent had deliberately changed nothing. This was observed live. A completed turn that got nothing from its model now fails with "The model returned an empty answer, so the agent did nothing this turn." — "nothing" meaning no text, no timeline part and no output tokens. A turn that only called tools, only reasoned, or only reported tokens still counts as an answer. An automation node settles as a retryableharness_errorso the stepper re-kicks it in place; past the deadline it still settles asdeadline. The task lane keeps its retryable path and its conversation handle, and a resumed start whose conversation answered nothing is no longer mistaken for a dead handle. The same change fixes the usage ledger: a Claude Code turn's totals now come from the result's ownusage, where a reasoning-only turn used to book zero.
Chat gets an audio transcription model, and stops warning about audio you never sent (#3399)
Server-side transcription had no setting: it took whatever the organization's credentials happened to offer, and the composer carried a standing warning about audio configuration whether or not anyone had tried to send audio.
Settings › Governance › Models now carries Audio transcription model, with Automatic — choose an available transcription model from the organization's active default provider credentials — or an exact provider and model pin. One resolver answers every caller: the settings status, the composer's capability read, server dictation, audio and video attachments, and a video link's audio fallback. An explicit pin that is unavailable reports an error and is never silently substituted; an unsupported audio upload is refused before the file is transferred. OpenRouter's dedicated speech-to-text catalog is discovered separately from its chat listing, so transcription models it omits from the default list — including entries with zero token context — are selectable, while pure speech-to-text entries stay out of chat and harness selection. Transcript caching is scoped to the serving provider, model, endpoint, response format and the file's bytes, so changing the target re-transcribes rather than answering from a stale entry. Existing completed attachments keep their transcripts.
Ordinary chat is quiet again. Instead of the standing composer warning there is one shared recovery dialog, shown only after an attempted server transcription is actually refused — for dictation, for an audio or video attachment, and for a retry. It offers the settings action the reader is permitted to take, or tells them whom to ask. Known unavailability blocks microphone capture and media transfer before they begin, while ordinary files and the browser's own speech recognition — still the preferred dictation path — remain usable. Dismissing it preserves the draft and returns focus, and a late failure from another conversation or another organization cannot open or overwrite the dialog you are looking at.
The policy is organization-scoped file configuration seeded through the existing config catalog. An absent file, or {}, means Automatic. There is no SQL migration and no backfill.
The UI evaluation round's findings are fixed (#3399)
The same change closes a round of confirmed UI defects that left dialogs unresponsive, navigation and settings stale, governance actions incomplete, and task completion able to publish after cancellation.
- Menus and dialogs hand off cleanly. A dropdown menu can hand off to a modal while its own exit animation is still mounted, and two modal layers left Radix's outside-pointer lock behind when the second overlay closed — a dialog that no longer answered the pointer.
@tale/ui'sDropdownMenunow leaves modality to the dialog, andDataTableActionMenutakes atriggerRefso the toolbar button, not the menu item that vanished, is where focus returns. Keyboard access, notification destinations, navigation context, responsive layouts and accessible editor labels are restored alongside it, together with organization-access failures, managed organization-creation capabilities, and the return navigation after a forced password change. - Actions that quietly did nothing now work: bulk chat archive and delete, retained notification preferences, the team list's refresh, trash row identity, URL-backed usage filters, erasure approval actions, legal-hold release and refusal detail, contact import validation, indexed-page feedback, and the OAuth error page's way back into the app.
- Task completion and cancellation are fenced in one transaction with consistent lock ordering, so a cancelled or superseded agent completion cannot publish comments, outputs or reviews. Remote effects that already completed before the cancellation are outside that guarantee.
- Automation authoring is server-validated: saves and deployments run the same acceptance checks and test gates the server owns, editor metadata survives a save, the settings and trigger dialogs close on success, copied run JSON is the run's input, and a diagnostic trace that hits its cap no longer takes the functional output with it.
- From a goal is gone. Goal-based automation creation and the private app builder flow behind it are retired; its route answers 404. Blank creation, package upload and shared MCP authoring are unchanged.
A list now pages the rows you can actually see (#3392, #3393)
Two doors carried the same defect in the same shape: read a page of rows, cut it at LIMIT, then drop the rows the caller may not read. A row the caller cannot see still spent a slot, so the page came back short — and when a whole page of them sat newer than the caller's own rows, it came back empty.
- Inbox (#3392). A member whose conversations sat behind a page of unassigned, admin-triage rows saw an empty tab under a badge reading 2. An empty first page is terminal in the UI: there is no row to scroll, so nothing loads more, and the tab stayed blank however many matching rows sat behind it. The assignment predicate moves into the list query's
WHERE, and oneresolveAssignmentScopefragment is now interpolated by the list and both count doors, so the list and its badge cannot disagree again. - REST documents (#3393).
GET /v1/documentsnow returns every hub document the caller can read. With?limit=1and one team-scoped document newer than the caller's own, the page came back empty whilecontinueCursorpointed past it — and a client that stops on an empty page stops there. The page query carries the samehubAccessClausethe in-app listing and hub search already run, the post-cut filter is gone, and the cursor is now the last row of the page the caller was actually given.
Neither door ever returned a row the caller was not entitled to; both returned too few. Two bounded readers in the same documents file still filter after their cut and report an honest truncated, which is left as a separate decision.
Approve can ask before it decides something outside the task (#3385, #3394)
Approve on an automation-owned task writes Done in one click. That fits most automations, but not one whose Done means something outside Tale — a relay, a filing, a message sent on the organization's behalf. During a test round a stray click attested an action nobody had taken.
The task subject contract gains review.approve.confirm: a sentence the automation declares, localised under i18n with the same fallback chain as every other declared text — exact tag, then base language, then the English sentence. When an automation declares it, Approve opens a confirmation showing exactly that sentence; without it, Approve stays a one-click close. An empty sentence, one over 500 characters, or a malformed locale tag is refused.
#3394 is why it works on a real deployment. The task modal resolves its owning contract through resolveTaskSubjectContract, which rebuilt the resolved entry field by field and left approveConfirmation out — so a deployment whose automation declared the sentence, and whose listing returned it in all three languages, still closed the task on the first click. The narrowing now drops only the ownership tag, so every field the entry resolves reaches the task, including any field added later.
Delegate the notification export without handing out an admin seat (#3388)
A worker that only needs to read notifications was refused GET /api/v1/notifications/sync with 403 ROLE_FORBIDDEN on a Developer seat, and the only way to unblock it was promotion to admin — which also hands over member management, SSO and SCIM administration, and password resets.
An administrator can now delegate that one right through the existing competence register, with no new table and no migration: grants and revocations use the current governance routes. tale: is a reserved namespace holding a closed set of platform capabilities; a grant naming any other tale: slug, in any casing, is refused with 400 COMPETENCE_CAPABILITY_UNKNOWN and audited as competence_grant_denied. Removing a membership stamps that member's live tale: grants revoked, so a re-added member starts without the right. The endpoint passes an owner or administrator by role, or a member holding one live, unexpired tale:notifications.export grant and an enabled seat — checked before the query is parsed or any member row is read — and GET /api/v1/me answers capabilities.notificationExport so a client can tell before it polls.
A managed deploy says why it failed, and checks packs before it pulls images (#3395)
A managed deploy ran silently for 38 minutes and then printed only "Deployment failed. Review the source pins, paths, permissions and retained recovery receipts." The cause was a pinned CLI older than the packs it was deploying: its embedded manifest schema stripped a field the pack declared, the release builder correctly refused the pack — and deploy replaced that authored refusal with its own fixed line. The 38 minutes were the runtime image pull, which ran before the configuration was ever validated.
All three parts are fixed. One failure rule now covers every configuration and deployment command: deliberate, bounded errors keep their words, while anything unexpected still collapses to the caller's fixed summary so no credential, registry or Docker output reaches a log — deploy had a stricter private copy of that rule, and that copy is what hid the cause. The refusal names the fields it could not read (…normalization changes release semantics at subjects.task.review.approve; this Tale CLI does not read those fields as written, so use a Tale CLI at least as new as the Tale the pack targets), listing paths only, never values, capped at five. And preparation now checks configurations before it pulls the runtime, logging each phase and image as it goes, so a pack this CLI cannot read fails within seconds instead of after the pull.
The first managed deploy after a snapshot no longer fails (#3389)
Every first managed deploy after a snapshot failed with "Docker refused managed Compose startup", while the rerun — which skips the snapshot — passed. A snapshot pauses the containers using each volume; Docker marks a paused container unhealthy, and unpausing keeps that status until the next probe. The deploy read the stale status as a runtime change and ran docker compose up, which refused immediately because a dependency read unhealthy.
The snapshot now reads each container's health before pausing and, after unpausing, waits for every container that was healthy or still starting to report healthy again — each with its own window of retries × (interval + timeout) + 5 s using Docker's defaults for unset settings, which is as long as Docker itself would take to call it unhealthy. A container that was healthy and does not recover fails the call by name and health status; one that was already unhealthy or has no health check is not awaited. Stale runtime health is now waited out rather than handed to Compose, the proxy is paused once for both of its snapshot volumes instead of twice, and a managed-Compose failure names the unhealthy services instead of the generic refusal. The config-volume migration gets the same guarantee.
Behaviour changes
- A managed Claude Code turn now carries
CLAUDE_STREAM_IDLE_TIMEOUT_MS,API_FORCE_IDLE_TIMEOUT=0,API_TIMEOUT_MS,CLAUDE_CODE_AUTO_COMPACT_WINDOW(only when the model's effective window is below 200,000) and, for a model that is not Claude,CLAUDE_CODE_ATTRIBUTION_HEADER=0. A managed Codex turn carriesmodel_providers.tale.stream_idle_timeout_ms. BYO-key runs keep each CLI's own defaults. - The gateway's per-request timeout is
max(600, SANDBOX_LLM_GATEWAY_STREAM_IDLE_TIMEOUT_SECONDS)rather than a fixed 600 seconds. Raising the idle budget therefore also lengthens how long a stalled non-streaming answer is held before it is cut. - An agent turn whose model answered nothing settles as a retryable failure instead of a success. An automation node re-kicks in place; past its deadline it settles as
deadline. - A Claude Code turn's usage totals are read from the result's
usage, so a reasoning-only turn is no longer booked as zero tokens. - Chat shows no audio setup warning until a server transcription is actually refused. An explicit transcription pin that is unavailable errors rather than falling back to another model.
- Approve on an automation-owned task opens a confirmation when the deployed automation declares
subjects.task.review.approve.confirm, and stays a one-click close otherwise. - Goal-based automation creation is removed; its route answers 404. Blank creation, package upload and MCP authoring are unchanged.
- The Inbox list and
GET /api/v1/documentspage over rows the caller can read, so a page that used to come back short or empty now carries them. The cursor of a documents page is the last row of that page. - A member holding a live
tale:notifications.exportgrant passesGET /api/v1/notifications/syncwithout an admin seat. - The managed proxy answers
GET /api/app/organizations/capabilitieswith{"canCreate": false}, so the app hides organization creation and points at the operator instead of failing the attempt.
API contract changes
- The OpenAPI document moves 1.13.0 → 1.14.0, one additive step. The surface is unchanged at 80 paths, 127 operations and 58 schemas, and the
Error.codeenum keeps its 150 values. GET /api/v1/megainscapabilities.notificationExport: true when the key holder may export members' notifications, by role or through a livetale:notifications.exportgrant.GET /api/v1/notifications/syncdocuments the delegated path; its 403 keepsROLE_FORBIDDENand now names the capability.COMPETENCE_CAPABILITY_UNKNOWNis on the app's governance route, not this surface.- No operation was added, removed or renamed, and no existing response shape changed.
Security
- hono 4.12.34 → 4.13.5 (CVE-2026-84363, #3300). Hono's query parsing did not stop at a URL fragment, so a
?after a#was read as a query string. Every other consumer of a URL — browsers,new URL(), reverse proxies — ignores everything from the first#, which makes this an interpretation differential: anything in front of the application that inspects the query string sees no parameters while the application reads and acts on them. Hono is the platform backend's HTTP framework, so this one ships in the running product. - sharp 0.35.3 → 0.35.4 (GHSA-rgj7-g3m4-5g8c, #3298), which picks up fixes for libheif vulnerabilities, two rated critical, that can lead to remote code execution on glibc-based Linux under certain conditions. In this repository
sharpis a development dependency only — image optimisation for the marketing site and the documentation screenshot pipeline — so no running Tale service decodes user-supplied images with it. - vitest 4.1.10 → 4.1.11 (#3299), a path-traversal / arbitrary-file-read fix in
@vitest/mocker. Test tooling only; it is in no shipped image. - Least privilege for the notification export (#3388). The previous answer to a refused export was to promote the caller to administrator, which also granted member management, SSO and SCIM administration and password resets. The delegated capability grants exactly the one right, is organization-scoped, optionally expiring, revocable, and revoked when the membership ends; the
tale:namespace is closed, so an unknown slug cannot be granted by typo and the attempt is audited. - An approval that decides more than the task now says so (#3385, #3394). A one-click Approve on an automation whose Done relays a decision outside Tale is a real hazard — a stray click attested an action nobody had taken during a test round. The confirmation is declared by the automation, so it can only appear where the automation's author said it should.
- A cancelled agent run cannot publish (#3399). Task completion and cancellation are fenced in one transaction with consistent lock ordering, so a superseded or cancelled run cannot write comments, outputs or reviews after the fact.
- Not a disclosure. The two list fixes (#3392, #3393) made short pages whole; neither ever returned a row the caller was not entitled to read. The scoping rule is unchanged — only the place it is applied.
Known issues
- A managed deployment needs a new bundle for the organization-creation capability. The proxy image does not change in this release, but the managed proxy policy the CLI renders does: it gains the
GET /api/app/organizations/capabilitiesresponse that lets the app hide organization creation. A managed deployment gains it only when a newly prepared bundle is applied; updating the platform image alone leaves the retained proxy policy as it was. Self-hosted and workspace deployments get the answer from the platform backend and need nothing. - The transcription setting is only as good as the organization's credentials. Automatic selects from the active default provider credentials that offer a compatible model; with none configured, the capability read answers
NO_TRANSCRIPTION_MODELand server dictation is unavailable. Browser speech recognition and published video captions keep working regardless. The OpenRouter smoke test covered a 1.78-second recording, which says nothing about long-recording limits. - The agent-turn fixes are bounded by the CLI they were read from. The values, the gates and the env variable names come from the pinned Claude Code build; a future pin can change them, and each of the six is proved by fixtures and one or two controlled runs rather than by a long soak on a slow model.
- The 0.5.29 proxy change is exercised live in one TLS mode only. The hosted fleet proved the rendered
trusted_proxiesblock and the removal of theX-Forwarded-Proto {scheme}pins inTLS_MODE=letsencrypt. No deployment onTLS_MODE=externalhas exercised them; an operator there should still verify sign-in callbacks, secure cookies, uploads and streaming through the full path. - The web tier's backend-URL default lives in the image, not in the generated compose (0.5.30). A workspace deployed with the CLI behaves correctly once its platform container runs a 0.5.30-or-later image, but the compose file the CLI writes still names no
TALE_BACKEND_URL. Any deployment still on an older platform image needs the explicit variable. - The scheduled-pack fix does not reach an existing install (#3381, 0.5.29). Provisioning skips an automation an organization already has, so an upgraded instance keeps the version whose input schema refuses its own scheduler. Edit that automation's
inputsto admittriggerandfiredAtand deploy a new version; a new organization is seeded correctly. - A budget hold covers a turn's first round. A turn that calls tools runs up to five model rounds, each billing its full prompt again, and only the first round's worst case is held while it runs. Concurrent sends can no longer each pass a cap with room for one, but a long multi-round turn can still settle above the cap it was admitted under.
- A run still carries no usage or cost. #3401 fixes the Claude Code turn totals that fed the usage ledger; it does not add a
usageblock toGET …/runs/{runId}, which remains contract debt with its design recorded. - Nothing backfills a task timeline (#3379, 0.5.29). Edits made before that release wrote audit rows only and do not appear; a label deleted from the catalog renders as its raw id rather than dropping the row.
- Unchanged from v0.5.20, where each is described in full: the
es/co-ccColombian cédula detector still ships switched off and a locale-agnostic PII toggle still widens national-ID matching to every locale; thinking-block replay on the native Anthropic connector is not done and the live Max-plus-tool-call check is still owed;rag_searchembedding calls inside a harness turn are unmetered; the product edit dialog cannot clear a field; the app's skill editor still carries the retiredprivatevisibility. - The
x-tale-paginationextension is a declaration on the OpenAPI document; generated clients that do not read vendor extensions still branch on the two cursor names untilcursoris retired. - Cloud sync, left for later: there is still no Sync now action — the cadence is the fifteen-minute scan, so a reconnected account waits for the next run. A config whose owner leaves the organization is still deactivated silently by a different door, and a source-deleted item is still a status stamp with no bell.
- Documents indexed before 0.5.27 keep one vector per repeated passage until they are re-indexed; the content hash is unchanged, so only an explicit
retry-indexing(or a content change) re-embeds them. A site that has not been scanned since 0.5.27 has no stored robots rules until its next scan. - The rail's navigation memory has had part of its manual round: the R5 round drove six EN/DE/FR desktop and phone cases covering parts of
NAV-F16–NAV-F19; the remaining section, the second-account cases andNAV-B6–NAV-B9are still unrun. - A reply-language directive is a directive: a model may still answer in the prompt's language and nothing on the wire marks a slip.
- No image input on the REST chat send. A
visionmodel reads an image over REST only on a thread the app continued with an image attachment; the design of anattachmentsfield on the send is recorded as contract debt. - No REST door authors or deploys an automation —
POST /automationsanswers 405 by design. Build and deploy in the app, or over the MCP endpoint'ssave_automationanddeploy_automation; the REST key lists, reads, runs and wires triggers. - The app's zip upload of a skill bundle rewrites the bundle and moves
updatedAteven when the zip is byte-identical, wherePUT /skills/{slug}writes nothing. - A tool call the reply cap cut keeps
input: {}on the storedtool-callpart; the raw text the model emitted is still not on the transcript. - Folder names written before 0.5.24 keep their bytes; a sync engine's hub-path lookup can create an NFC twin beside a legacy NFD folder. No backfill ships.
- Two bounded document readers still filter after their cut (#3393's scope note); both report an honest
truncated, so a caller can tell the answer was cut. - Behind a Docker-published port, every IPv6 client arrives as the bridge gateway's address and shares one per-address rate-limit bucket and one audit address until the daemon runs with
ip6tablesand the reverse proxy's network is IPv6-enabled — an operator item, documented on the Own Compose page. - Recorded as contract debt, each with its design in the ledger: a queued send is invisible on the message list until a worker opens it; a webhook delivery the deployed
inputsschema refuses moves no trigger stamp; the MCPrun_deployedtool keys its idempotency apart fromstart_runand REST;robots.txt$end-anchors andAllow:lines are not honoured (prefix and*rules are), and a page is fetched three to four times per scan; a cancelled run answerstrace: nullandeffects: nullwhere a failed run answers both; approvals and asks have no REST twins; a task cannot be archived or deleted over REST; a webhook bind does not say whether the deployedinputsschema admits a delivery; an exhaustedrepeatUntilis only a trace note;Websitecarries noscanStartedAtand the crawler has no page cap, path filter or stop verb of the caller's; website search has no dense leg and its substring fallback stampsscore: 0silently; noIdempotency-Keyon the task start; no queue position on a queued send; a corrupt Office document still fails asindexer_errorand is retried five times where a PDF landsmalformed; no/.well-known/security.txt; no changelog feed on tale.dev; no SDK, collection or per-code table beyond theError.codeenum;GET /notificationsrows carrytypeas a free string and nothing pushes them to a machine caller; a skill keeps no version history on the machine door; the per-task circuit breaker is not built; the messages a conversation snapshot applied are readable only in the app.
Migration notes
- No migration. The application database stays at 0106 and the knowledge database is unchanged, so nothing runs at boot beyond the usual convergence check.
- No new environment variable.
SANDBOX_LLM_GATEWAY_STREAM_IDLE_TIMEOUT_SECONDSis not new — it has shipped since 0.2.93 — but this release documents it in.env.exampleand the environment reference for the first time, and widens what it does: above 600 it now also raises the gateway's per-request timeout, which bounds a whole non-streaming answer. Recreate the backend services after changing it. - One new configuration file.
governance/transcription-model.ymljoins the seeded per-organization catalog. An absent file, or{}, means Automatic, so an upgraded organization needs nothing; a pin is two fields,providerSlugandmodelId. The shippedclaude-codeandcodexharness definitions change too; those are system configuration and travel inside the platform image. - No image in the stop-gated tier changes. The
proxyanddbimages carry no source change in this range and the object store runs its pinned third-party image, so a plaintale deployis the whole upgrade — no--stop, no downtime window. The managed proxy policy does change (see Known issues): a managed deployment picks it up with its next prepared bundle. - The platform image (all six agent-turn changes, the transcription lane, the UI round, the two list fixes, the approval confirmation, the delegated capability, hono) and the docs image (15 pages in each of en, de and fr — 45 files — plus a regenerated Models screenshot) carry source changes; ui-docs gains the
DataTableguide'striggerRefparagraph, and web only a development dependency. Thedb,proxy,sandbox,sandbox-runtime,sandbox-buildkitd,sandbox-egressandsandbox-llm-gatewayimages carry no source change — the gateway's README moved, its binary did not. - The CLI has source changes in this range (#3389, #3395, and the managed proxy policy), so a managed deployment should move its pinned CLI reference as well as its platform reference. The release executables report 0.5.31.
@tale/uiand@tale/marketing-uiare pinned by this release as theui-v0.5.31andmarketing-ui-v0.5.31tags on their snapshot branches; a consumer outside the monorepo installs"@tale/ui": "github:tale-project/tale#ui-v0.5.31". Unlike the 0.5.30 tags, these are not content-identical to their predecessors:@tale/uicarries the dropdown-menu modality change andDataTableActionMenu'striggerRef, and a consumer that relies on the menu owning pointer modality should re-check its own dialogs.
Upgrading
-
On the 0.5 line (0.5.0 – 0.5.30):
tale update tale deploy
Nothing in this release needs
--stop. A deployment crossing from a version older than 0.5.29 should read that release's notes, which do: itsproxyimage change is only applied by a--stopdeploy. -
Managed deployments move by pinning the CLI and the runtime to this release's commit, preparing a new bundle and applying it with the pinned CLI — see Managed deployments on the CLI install page. Pin the CLI reference too: this range changes it, and #3395 means a CLI older than the packs it deploys now says so within seconds instead of after the image pull. The bundle's backend-local phases run under the interpreted CLI (
cli/tale.mjs) that thesetup-cliaction andbun run --filter @tale/cli buildproduce beside the executable; the executable from the release page has no interpreted bundle beside it and cannot prepare a managed bundle. On a Linux x64 host whose CPU lacks AVX2, passlinux-baseline: 'true'to thesetup-cliaction so the bundle embeds the baseline executable. -
New install:
curl -fsSL https://raw.githubusercontent.com/tale-project/tale/main/scripts/install-cli.sh | bash mkdir tale-05 && cd tale-05 tale init tale deploy
On a CPU without AVX2 the downloaded executable aborts with
Illegal instruction; build it from source withbun run build:linux-baselineintools/cliinstead. -
Running agents against a self-hosted model? Nothing is required, but two knobs are worth a look.
SANDBOX_LLM_GATEWAY_STREAM_IDLE_TIMEOUT_SECONDSis now the single patience budget both the gateway and the agent CLI obey, so raise it to cover your model's worst cold prefill rather than leaving a turn to be re-sent. And make sure the connector catalog states your model's real context window:CLAUDE_CODE_AUTO_COMPACT_WINDOWis derived from it, and an unknown window leaves the CLI assuming 200,000.
What's Changed
- fix(cli): wait for paused containers to report healthy again by @yannickmonney in #3389
- fix(platform): stop Claude Code and Codex re-sending a turn on a slow prefill by @yannickmonney in #3387
- feat(platform): ask before an approval the automation says decides more by @yannickmonney in #3385
- feat(platform): delegate the notification export as a capability by @yannickmonney in #3388
- fix(platform): scope the inbox list in SQL, not after the page by @Israeltheminer in #3392
- fix(platform): scope the REST documents page in SQL, not after the cut by @Israeltheminer in #3393
- fix(platform): carry the approve confirmation to the task modal by @yannickmonney in #3394
- fix(cli): say why a deploy failed, and check packs before images by @yannickmonney in #3395
- fix(platform): stop Claude Code resending a request before its first byte by @yannickmonney in #3396
- fix(platform): let the gateway wait a raised budget for a whole answer by @yannickmonney in #3397
- fix(platform): let Claude Code compact inside the serving model's window by @yannickmonney in #3398
- fix(platform): keep a foreign model's prompt cache across Claude Code turns by @yannickmonney in #3400
- fix(platform): fail an agent turn whose model answered nothing by @yannickmonney in #3401
- fix(platform): resolve UI findings and add audio model settings by @larryro in #3399
- chore(deps): update dependency sharp to v0.35.4 [security] by @renovate in #3298
- chore(deps): update dependency vitest to v4.1.11 [security] by @renovate in #3299
- fix(deps): update dependency hono to v4.13.5 [security] by @renovate in #3300
Full Changelog: v0.5.30...v0.5.31