Skip to content

Tale v0.5.41

Choose a tag to compare

@larryro larryro released this 20 Sep 15:15
3f7f6e0

0.5.41 carries 7 merged pull requests — every change since 0.5.39. A v0.5.40 tag was pushed on 2026-09-20 at the commit of the sixth of them, but its pipeline stopped at the container test gate — the image-validation step reached the job's twenty-minute limit while rebuilding the platform image on the runner — so no multi-architecture manifest, no latest tag, no GitHub release and no executables were ever published under that number; latest stayed on 0.5.39, and the tag remains, so the number is skipped. This release is the first that carries those six pull requests, with a seventh on top. Its themes: an organization's own inference endpoint — a vLLM or Ollama box, an internal gateway — is connected from the Add credential dialog instead of a file an operator writes; the model a person pinned in the chat composer remembers which provider served it; spend is booked at the vendor's prompt-cache hit price instead of the plain input rate, which had overstated a cache-heavy day of DeepSeek traffic about eleven-fold and "reached" a spend cap that was not, and a turn the gateway refuses for spend ends there instead of being retried three times on one-cent keys; a Claude Code automation turn on DeepSeek, Moonshot, the Vercel AI Gateway or OpenRouter rides the vendor's native Anthropic door, which closes the two 400s that killed an agent run on the down-converted wire; the chat transcript's send-snap wins against the thread-open hold and a trackpad's momentum, and Scroll to bottom follows a streaming reply; a reply that failed before its first token stops "thinking"; the organization's embedding model is picked from the provider's catalog, and a shipped provider whose catalog lists none is refused up front instead of failing at index time; and DeepSeek's catalog names the vendor's current Flash model. No contract change — 1.19.0 stays — one migration (0112) adds a nullable column, no environment variable changes, and no image in the stop-gated tier changes, so the upgrade is tale update followed by a plain tale deploy.

Highlights

Define a custom AI provider from the settings UI (#3431)

An organization-owned provider — a self-hosted model server such as vLLM or Ollama, or an internal gateway that speaks the OpenAI or Anthropic API — existed only as a file under TALE_CONFIG_DIR/<org>/providers/ that an operator wrote by hand or through the managed-configuration lane; from the app it looked impossible. Settings > AI providers > Add credential now pins a Custom provider entry under the catalog (it survives a search that matches nothing), and its setup step takes the provider's facts beside the key: a Provider name (it names the provider and the credential; the identifier is derived from it and numbered past an existing one), the API format (Chat Completions for vLLM, Ollama, LiteLLM and most gateways; Messages for Anthropic-compatible endpoints), the Base URL, and how the endpoint's Models are known — Discover from the endpoint, which reads its /models with this key, or Enter model IDs for a server that cannot list them. One submit writes the definition file through the existing definition door and then the credential; a credential the server refuses rolls the definition back. The row and the picker carry a Custom badge; the row menu's Check models lists the endpoint afresh with the organization's key; Edit credential edits the provider's facts against the definition's loaded hash; and deleting the provider's last credential retires the provider itself — the dialog says so beforehand, the exact previous file is archived under .history/<name>/ like every save, and the audit log records provider_definition.deleted. Behind it, DELETE /api/app/providers/definitions/{name} (admin or developer, an optional expectedHash compare-and-set) is refused with 409 PROVIDER_IN_USE while any credential still names the slug, so keys are retired deliberately and never orphaned. Two listing defects surfaced with the first real custom provider and are fixed: a live /models is fetched with the organization's default key for the provider — most hosted OpenAI-compatible endpoints refuse an anonymous listing, which is what "catalog fetch returned HTTP 401" was, and a remembered anonymous refusal no longer holds back the first keyed attempt — and a bare OpenAI-shape listing (id, object, created, owned_by, the shape of api.openai.com, DashScope, DeepSeek, vLLM and Ollama) no longer normalizes to nothing: such an entry is admitted under the same assumed 128,000-token window an allowlist entry carries and reads as tool-capable, while a catalog that publishes its windows (OpenRouter) stays strict. A private or loopback address still needs the deployment's own opt-in, which an operator sets; the dialog names it. The providers page gains Define a custom provider in English, German and French, and the self-hosted providers page and the local-provider tutorial point at it.

The sticky model pick remembers its provider (#3431)

The composer's sticky pick stored the model id alone, and a new chat was seeded by id: a model listed by a shipped provider and by a custom provider on another endpoint of the same vendor landed on the shipped copy — whose key the person never meant — and failed with that provider's refusal. Migration 0112 adds chat_model_provider_slug to the preferences row; the preference saves and serves the (provider, id) pair, the seed prefers the exact pair and falls back to whichever provider serves the id for a pick saved before providers were part of it, the thread-title lane names the thread on the pick's own connector, and the picker's section headers show each provider's display name, so an organization-defined provider reads by the name its admin gave it rather than by its slug.

Spend is booked at the vendor's cache-hit price, and a spend refusal ends the turn (#3433)

An automation agent run ended with "the setup assistant could not run". Its record: the node's gateway key had been minted at the $1.51 the organization's monthly cap had left, the gateway refused the forty-third call with 402 Model-level budget exceeded, and the auto-retry then resumed the same transcript three times on one-cent keys, each dead on its second call. The cap had been reached on overstated data: for that day the vendor console showed ¥6.50 for 33 million tokens, the platform had booked $10.20 for the same traffic — 87 % of it prompt-cache hits — because the gateway had been told only an input and an output rate, so it billed every cached token at the input rate (its documented fallback), and always at the peak rate. Three fixes. Cache hits are priced: a catalog entry's pricing gains optional cacheReadCentsPerMillion and cacheWriteCentsPerMillion; every shipped catalog carries the vendor's cache prices as published on 2026-09-20 (49 of the 59 shipped entries; the file headers cite the listings), a live listing's pricing.input_cache_read and input_cache_write (OpenRouter and the Vercel gateway spell them the same) are read into the entry, a hit price above the input price is dropped as a listing glitch, the pricing patch pushed to the sandbox gateway carries cache_read_input_token_cost and cache_creation_input_token_cost, a stored patch that lacks a cache rate the catalog now has is rewritten rather than trusted, and the chat lane's one cost formula bills the reported cached share at the hit price (the ledger entry carries cachedInputTokens). Base prices moved where the vendor page had: claude-sonnet-5 is priced at last ($2 / $10 per million), gpt-5.6-sol $5 / $30 → $4 / $20, kimi-k2.6 and kimi-k2.7-code → $0.95 / $4, glm-5.1 and glm-5.2 → $1.40 / $4.40, and deepseek-v4-pro → $1.32 / $3.96 (see the DeepSeek highlight). A mid-turn 402 is a spend refusal, not a provider hiccup: both agent hosts settle a harness result carrying API status 402 as budget_exceeded, with a reason that names the exhausted allowance and keeps the gateway's line, and neither lane's retry gate re-kicks it — a resumed retry replayed the whole transcript into a key sized from the same balance. Settlement reads the key's live spend: the gateway meters in memory and dumps its counters to the store every ten seconds, and the plain read served the store, so a settle seconds after a turn's last call missed that call (a 156-cent turn booked 150; a one-call retry booked 0); the read now asks for the live counters the gateway's own budget gate meters and falls back to the stored row only when the live index lacks the key. The REST thread's model view keeps its documented input/output pair; the cache prices are billing inputs, not part of the contract.

Claude Code automation turns ride the connector's native Anthropic door (#3432)

An automation agent node running Claude Code on DeepSeek rode the connector's OpenAI gateway record although the shipped definition declares the vendor's native Anthropic endpoint and the serving resolver said so: the workflow agent host never read the flag on the kick's routing, the scheduled start, the key mint or the answered-ask resume, where the task lane has threaded it since 0.5.19. The down-conversion still breaks on the shipped gateway: both agent nodes of a 2026-09-20 run lost their first attempt within ten seconds to 400 The reasoning_content in the thinking mode must be passed back to the API on a fresh session's second call, and the setup node then read two PDFs, which Claude Code attached as Anthropic document blocks, the gateway turned into a file part, and DeepSeek refused with 400 … file must have a file_id or file_data on every resumed retry. The lane now reaches all four places, so the exec model, the scheduled start's arguments, the minted key's allowed-model reference and the provisioned record all name the connector's …__anthropic record. The shipped connectors were checked against vendor documentation for a native door: Moonshot (api.moonshot.ai/anthropic), the Vercel AI Gateway (ai-gateway.vercel.sh) and OpenRouter (openrouter.ai/api) now declare one beside DeepSeek's — from the vendor docs, not live-probed from this platform — and a gateway-standard connector on the Claude Code lane (OpenRouter) takes an organization-scoped record of its own, because the gateway's built-in OpenRouter implementation speaks only the OpenAI wire to the vendor: a Claude model then passes through natively in both directions instead of being down-converted here and converted back there. Z.ai keeps no door on purpose (its Anthropic endpoint silently drops image blocks for every model, which a harness endpoint cannot yet declare), and Qwen's door is a workspace-specific host that a shared definition cannot name — an organization declares it on its own connector.

The chat send-snap outranks the thread-open hold and a trackpad's momentum (#3433)

"Sending sometimes does not scroll" had two reproduced causes. Opening a thread arms a two-second position hold, and a content tick under a live hold returned before it consumed the send intent, so a send within that window never scrolled once the reply settled inside it. And any wheel or touch movement over the transcript cancelled the snap, direction ignored, so a trackpad's momentum tail after scrolling to the bottom — or a one-pixel downward wheel with the pointer resting over the messages — killed the glide mid-flight. A send or edit intent now wins over any live hold on the next content tick; only an upward wheel turn or a finger travelling down the screen counts as taking over; and Scroll to bottom pressed while a reply streams engages a follow latch that keeps the view at the growing bottom until the person scrolls up, sends, or opens another thread — it survives the client-side reveal that keeps draining text after the server settled, where a one-shot jump had left the button back on screen a second later. The button no longer flashes during the send glide, and a glide that lands late on slow frames gets a fresh 250 ms settle. Underneath, the transcript is one list with a spacer after it instead of three lists around the last user message: a send used to move the previous turn's rows between lists, which remounted them, and a freshly inserted lazily-rasterized row spends its first frames at a 200-pixel placeholder, so the content above the new message jumped on every send and, after a long reply, clamped the scroll position out from under the glide. Every row now keeps its rendered height as it becomes history, and the spacer is written synchronously from a mutation observer, so the scroll height no longer bounces on every streamed chunk — the wobbling scrollbar thumb is gone. A real-Chromium browser test covers the hold-plus-send, the wheel directions, row identity and the follow latch.

A reply that failed before its first token stops thinking (#3431)

A turn that failed before any text — a refused key, a stop before the first token — kept "Thinking · Ns" ticking under its error: the failed settle drained the row like any other, and with no text nothing ever painted a first glyph, so the pre-answer shell stayed. The shell now drops on a terminal row (failed, stopped, refused), and the thread view never presents a terminal row as streaming, not even under a generation row that outlived its settle; a failure mid-stream settles what the server wrote, never a tail it refused.

The embedding model is picked from the provider's catalog (#3435)

Settings > Data residency > Embedding model took the model as a free-text tag that failed only at index time. The Model row now reads the organization's provider catalogs — the same listing the AI-providers and governance pages use — and lets the catalog decide: a provider whose catalog lists embedding models gets a closed select over them, and a pick fills Vector width from the catalog; an organization-defined provider whose listing tags embedding models gets the same select plus Other model…, which opens a Model tag field, because a bare /models listing tags nothing as an embedding model; a shipped provider whose catalog lists none is refused — "The {provider} catalog lists no embedding model. Choose a provider that serves one." — with no field and the shared Save off; a shipped catalog that could not be loaded is refused naming the remedy; and only a provider with no listing to consult (an organization-defined one without an entry, Azure's catalog: none deployments) takes a typed tag, with a hint that says why. Three shipped catalogs list an embedding model today: OpenAI (text-embedding-3-small), OpenRouter (qwen/qwen3-embedding-8b) and Z.ai (embedding-3), each at width 1536. The @tale/ui Select gains errorMessage, rendered under the control the way Input renders its own — an error routed through description landed in the label column of a settings row — and the form's three selects use it. The data-residency page says so in English, German and French.

DeepSeek's catalog names the vendor's current Flash model (#3429)

DeepSeek released DeepSeek-V4.1-Flash on 2026-09-10 under the API id deepseek-flash — there is no deepseek-v4.1-flash on the vendor API — and retired deepseek-v4-flash, which is "temporarily routed" to the new model; the shipped catalog still offered the retired id, and the Pro price had been stale since the vendor's 2026-08-16 pricing change. deepseek-flash replaces deepseek-v4-flash: native image input, a 1,048,576-token window, 384,000 output tokens, the same effort knob, and the official peak rates ($0.30 / $1.20 per million, cache hit $0.006) — DeepSeek bills half of that off-peak and the one input/output pair the schema carries cannot express a time of day, so the conservative figure ships; deepseek-v4-pro moves to $1.32 / $3.96. On OpenRouter, deepseek/deepseek-v4.1-flash joins the curated defaults at the aggregator's price and deepseek/deepseek-v4-flash stays, because third-party hosts still serve it there. Auto's draft band prefers deepseek-flash, then deepseek-v4.1-flash (the aggregator's spelling), in place of the retired id. Shipped catalogs change only with a release (Refresh catalogs skips source: static), so the card shows the new default after the deploy.

A sandbox vision batch can switch the model's thinking off (#3434)

The sandbox runtime's tale-vision tool — the batch that reads images for a text-only agent — can spend a page's per-image deadline on the model's reasoning before it returns a transcription. tale-vision --thinking disabled sends the standard disabled-thinking request for that batch; the default (provider, or omission) leaves the provider's behaviour and the historical cache keys unchanged, and an explicit override caches under its own key. Nothing in the platform passes the flag yet; it is there for a deployment that calls the tool directly. Image bytes, model selection, output limits, deadlines and the ordinary Read fallback are untouched. The runtime's README documents it.

bun dev builds the sandbox runtime image when it is missing (#3430)

For contributors: the host development loop brought up every Compose backing service but never the sandbox runtime image, which is neither a Compose service nor a registry image, so on a fresh checkout — or after a local image cleanup — the first agent session died with Unable to find image 'tale-sandbox-runtime:latest' locally. After the backing services and the gateway wait, bun dev probes for the image and, when it is missing, runs the same one-time build docker:dev uses (several minutes, labelled as such); a failed build degrades to a warning that carries the exact retry command, and the fleet keeps booting with only sandbox sessions unavailable. The contributor-setup page says so.

Behaviour changes

  • Settings > AI providers > Add credential pins a Custom provider entry under the catalog; its setup step takes Provider name, API format, Base URL, Models (Discover from the endpoint or Enter model IDs) and the key. Organization-defined providers carry a Custom badge in the row and in the picker; the row menu gains Check models; Edit credential edits the provider's facts; deleting the last credential retires the provider after a warning. The shipped providers and the file lane are unchanged.
  • A live model listing is fetched with the organization's default key for the provider (models-endpoint catalogs only); a remembered anonymous refusal no longer blocks the first keyed attempt. A bare OpenAI-shape /models entry is admitted under a 128,000-token window and reads as tool-capable; OpenRouter's listing stays strict.
  • The composer's sticky pick saves and seeds the (provider, model) pair; a pick saved before this release resolves by id, as before. Picking Auto clears both. The picker's section headers are the providers' display names. The thread-title lane names a thread on the pick's own connector.
  • A reply that ended without text shows its notice without the dots or the ticking timer; a stopped-before-the-first-token row renders Generation stopped.
  • Chat scrolling: a send or edit snaps the new message to the top even inside the thread-open window; a downward wheel or a sideways swipe over the transcript no longer cancels the glide, an upward one still does; a finger travelling down the screen escapes, one travelling up does not; Scroll to bottom during a stream keeps following the reply until an upward scroll, a send or a thread change; the button stays hidden during the send glide; the previous turn's rows keep their height when a send demotes them; the scroll height no longer bounces per streamed chunk.
  • Spend: the usage ledger and the per-message cost bill a provider-reported cached share at the catalog's cache-hit price (the input rate when the catalog has none); the gateway's per-model pricing carries the cache pair; agent-turn settlement books the key's live counters. Historical rows are not rewritten.
  • A managed agent turn that the gateway refuses with 402 — the key's budget, sized from the organization's remaining cap, is spent — settles as budget_exceeded with a reason naming the exhausted allowance; neither the automation nor the task auto-retry resumes it (was: harness_error, retried three times on one-cent keys).
  • A Claude Code automation agent turn on DeepSeek, Moonshot, the Vercel AI Gateway or OpenRouter rides the connector's native Anthropic endpoint through an organization-scoped …__anthropic gateway record on the kick, the scheduled start, the key mint and the answered-ask resume; the vision model's routing is unchanged; every other harness keeps the record it used. Z.ai and Qwen keep the OpenAI base.
  • Shipped catalogs: deepseek-flash (vision) replaces deepseek-v4-flash; deepseek/deepseek-v4.1-flash is added on OpenRouter; 49 entries carry cache prices; the base prices of claude-sonnet-5, gpt-5.6-sol, kimi-k2.6, kimi-k2.7-code, glm-5.1, glm-5.2 and deepseek-v4-pro change as listed above. Auto's draft band prefers deepseek-flash.
  • Settings > Data residency > Embedding model: Model is a select over the provider's catalog where the catalog lists embedding models (a pick fills Vector width), a select plus Other model… for an organization-defined provider whose listing tags them, a refusal for a shipped provider whose catalog lists none or could not be loaded, and a typed Model tag only where no listing can tell. A stored tag a shipped catalog does not list shows as its own entry, so nothing configured is hidden. Switching the provider empties the model.
  • @tale/ui Select takes errorMessage, rendered under the control with role="alert" and announced with the trigger; it implies the invalid state.
  • tale-vision --thinking disabled in the sandbox runtime sends thinking: {type: "disabled"} for that batch and caches apart from the provider default; --thinking provider and omission are unchanged.
  • bun dev builds tale-sandbox-runtime:latest from source when it is missing (contributor loop only).
  • Documentation (English, German and French): Define a custom provider on the providers page with the self-hosted providers page and the local-provider tutorial pointing at it (and the listing's assumed window "when the listing publishes neither"); the chat basics say the pick keeps its provider; the data-residency page describes the catalog-driven Model row; the contributor-setup page describes the runtime-image build; the @tale/ui docs say Select takes errorMessage.

API contract changes

  • None. The contract stays at 1.19.0: 86 paths, 134 operations, 63 schemas, 167 Error.code values; X-Tale-Api-Version answers 1.19.0, and generate:openapi on the release commit reproduces the shipped document byte for byte. The REST thread's model view keeps its documented pricing pair; the catalog's cache prices are deliberately not on the wire.
  • App doors, not the machine contract: DELETE /api/app/providers/definitions/{name} (admin or developer; optional expectedHash; 409 PROVIDER_IN_USE, 404 PROVIDER_NOT_FOUND); DELETE /api/app/provider-credentials/{id}?retireUnusedCustomProvider=1; GET /api/app/providers/catalogs rows carry origin: "shipped" | "organization"; the chat-model preference write takes providerSlug beside modelId and the preferences read answers chatModelProviderSlug beside chatModelId; the composer's model options carry providerLabel.
  • MCP: unchanged.

Security

  • Custom provider definitions are written and deleted only by an admin or developer (#3431). The delete door refuses while any credential names the provider (409 PROVIDER_IN_USE), takes an optional compare-and-set on the definition's reviewed hash, archives the exact previous file under .history/<name>/, and writes a security-category audit row (provider_definition.deleted). A base URL on a private or loopback host is refused unless the deployment opted in (TALE_ALLOW_PRIVATE_PROVIDER_HOSTS), and a public host needs https; cloud-metadata addresses stay blocked. A name that belongs to a shipped provider is refused as reserved.
  • The organization's provider key now travels to the provider's own /models listing (#3431) — the same endpoint the key already serves chat calls to — as a bearer on the listing request. It is never logged and never part of the catalog cache key; a listing is the provider's catalog whoever fetched it. A provider with no usable default credential lists anonymously, as before.
  • A spend refusal is final (#3433). A turn the gateway refuses with 402 is settled as budget_exceeded and not re-kicked, so an exhausted cap no longer funds three more transcript replays on one-cent keys; and a turn's last calls are booked from the gateway's live counters instead of a stale row.
  • The minted key follows the record the session calls (#3432). On the Claude Code lane the key's allowed-model reference binds to the connector's …__anthropic record — the same routing the exec model and the provision use — and the pricing override is provisioned on that record, so a session on the native door is metered and capped like one on the OpenAI record.
  • The sticky pick's provider slug is validated (lowercase letters, digits and single hyphens, at most 120 characters) before it is stored (#3431).
  • No dependency changes in this range, and no advisory is fixed. The @better-auth/oauth-provider advisory noted in 0.5.33 (CVE-2026-67332 / GHSA-p2fr-6hmx-4528, medium) remains open with its workaround in place; the 1.7.0 upgrade is still a separate dependency pull request.

Known issues

  • 0.5.40 was never published. The v0.5.40 tag exists at the commit before this release's last pull request; its per-architecture images (0.5.40-amd64, 0.5.40-arm64) reached the registry, but no 0.5.40 manifest, no release and no executables did, so a deployment cannot pin that number. Move to 0.5.41.
  • Spend history is not rewritten. Ledger rows and per-message costs booked before this release keep the overstated figures; cost figures from this release on are lower for cache-heavy traffic, so a month-to-date total spans two pricing rules. A budget that was "reached" on the old figures frees up only as new, correctly priced usage replaces the window.
  • Off-peak pricing is not modelled. The shipped gateway cannot express a time-of-day rate, so DeepSeek's catalog carries the peak rate all day and books double the vendor's off-peak charge; nor are long-context tiers (xAI at 200,000 tokens and above, OpenAI beyond 272,000) — one pair per model.
  • Cache prices are not surfaced in the model popover, over REST or in the OpenAPI document; they are billing inputs only.
  • A pin to deepseek-v4-flash — a sticky pick, an agent's supportedModels, a governance default, an automation llm node, a REST caller — answers CHAT_MODEL_UNKNOWN after the upgrade and needs re-picking, although DeepSeek still routes the old name for now; a catalog alias the resolver honours is not built.
  • The Moonshot, Vercel and OpenRouter harness doors are declared from vendor documentation and were not live-probed from this platform; a Claude Code automation on one of them exercises the door for the first time.
  • Three follow-ups from the lane fix are open: the Read hook's PDF guard arms only for text-only models (deepseek-flash has native vision, so a native PDF read reaches the wire and the vendor's Anthropic door drops the document silently — the model is told it was unsupported); a resumed retry that fails with the same invalid-request 4xx replays the rejected transcript instead of re-kicking a fresh conversation; and Z.ai's Anthropic door stays unused until a harness endpoint can declare that it drops images.
  • A shipped provider whose catalog carries no curated embedding entry (DeepSeek, Anthropic, xAI, Moonshot, Qwen, Gemini, Nous) is refused as an embedding provider, including the ones whose vendor does serve embeddings — the remedy is curating the entry, width included, in a release, not a typed tag. OpenAI offers text-embedding-3-small only. Azure keeps the free field although it is shipped: its deployments carry the admin's own names and there is no listing to consult.
  • The custom-provider form authors api-key and env credentials, models-endpoint or no catalog, and the two API formats; a subscription auth entry or an embedding declaration on an existing definition is preserved on edit, never authored — that stays the file lane. A definition that survives a failed retirement (the credential is gone either way) stays visible in the picker until its next credential is deleted.
  • A bare listing's assumptions are assumptions: 128,000 tokens of context and tool support. A server that refuses tools fails loudly on the wire, never silently; a server that serves less context than assumed is corrected by a lower context limit in governance.
  • tale-vision --thinking disabled is validated against a controlled upstream and the runtime's own tests; whether a given vendor honours the disabled mode, and what it does to transcription quality and latency, remain deployment checks. Nothing in the platform passes the flag.
  • The chat scroll rounds CHAT-F39 and CHAT-F40 are browser-tested for the wheel and the follow latch; trackpad momentum and touch gestures on a real device stay manual. The deferred-send path (attachments still processing) sets no scroll intent, by design.
  • The bun dev runtime-image step wraps the same recipe docker:dev uses; its own rendering during a missing-image boot was not observed live, because a second orchestrator cannot share the running fleet's ports.
  • Unchanged from v0.5.39, where each is described in full: a knowledge entry the REST door wrote before that release keeps source: "manual"; a scan reuses a robots verdict up to a minute old and a row stored earlier carries no sitemap list; the ask retraction is best-effort and forward-only; the app's own archive stays unfenced; GET /api/v1/teams is read-only and a complete set; the chat content cap counts UTF-16 code units; the nullable on a oneOf branch is the OAS 3.0 spelling; 0.5.38 shipped with generated notes and its four pull requests (#3424–#3427) have no known-issues record.
  • Unchanged from v0.5.37, where each is described in full: ledger rows booked before that release keep the subject they were booked under and a project agent can appear twice in Top assistants across the upgrade; app.usage_events is write-retired, not dropped; nothing in the schema forbids a door string in usage_ledger.user_id; the run list labels a keyed start Started by api-key:…; the GOV-F20 round is manual; llm nodes are unmetered and a run carries no usage or cost.
  • Unchanged from v0.5.36, where each is described in full: automation files: mounts and workflow document.* steps do not apply the team audience; a single-sign-on sign-in with an empty group list revokes nothing and a SCIM group replace overwrites hand-added members silently; the legacy team mirror columns stay; the three GIN indexes of migration 0109 were built without CONCURRENTLY; the team rounds NAV-F6, SET-F18, SET-F19, SET-F42, KNOW-F20, PROJ-F23, PROJ-F24 and CONV-F12 are manual; a team skill's teams list is validated only when it changes; REST Document.teamId stays as the deprecated single-team spelling.
  • Unchanged from v0.5.35, where each is described in full: a frame carries the signed-in session only from a same-site host page and the shell's embedding policy is the union across organizations; revoking a trusted-header key or turning the card off ends no session; the AUTH-F21–AUTH-F24, AUTH-B10 and SET-F41 rounds are manual; approvals have no REST twin; moving a folder has no door and documents already at the root stay there; the auto-retry resumes only a turn that announced its conversation handle; the Google Drive row counts a deployment app from either lane.
  • Unchanged from v0.5.34, where each is described in full: a managed deployment gets the organization-creator behaviour only once its specification declares organizations.creators and a new bundle is applied; the AUTH-B9 and AUTH-F20 rounds are manual; the creator list is matched against sign-in addresses.
  • Unchanged from v0.5.33, where each is described in full: the sign-up gate's first-boot race; the boot catch-up that marks provisioned accounts verified asks nobody; the break-glass administrator's password-rotation, single-sign-on-link and memory-adapter limits; the cross-scope webhook guard governs deliveries from that release on; a site's robots policy upgrades at its next scan; a scan waiting on render capacity takes longer by design; the governance pickers list only providers with an active credential; one dependency advisory is open.
  • Unchanged from v0.5.32, where each is described in full: the embedding pacing is proved against a controlled server, its bound is per Tale process, and minTokensPerSecond is a statement nothing verifies; the Kubernetes page's verified scope is one kind cluster, config-data needs RWX or a single node, and Tale ships no Helm chart.
  • Unchanged from v0.5.31, where each is described in full: a managed deployment picks up that release's proxy policy only when a newly prepared bundle is applied; the transcription setting is only as good as the organization's credentials; the six agent-turn fixes are bounded by the pinned Claude Code build they were read from; the 0.5.29 proxy change has been exercised live in TLS_MODE=letsencrypt only; the web tier's backend-URL default lives in the image, not in the generated compose; the scheduled-pack fix does not reach an automation an organization already has; a budget hold covers a turn's first round only; nothing backfills a task timeline.
  • Unchanged from v0.5.20, where each is described in full: the es/co-cc Colombian cédula detector still ships switched off and a locale-agnostic PII toggle still widens national-ID matching to every locale; thinking-block replay on the native Anthropic connector is not done and the live Max-plus-tool-call check is still owed; rag_search embedding calls inside a harness turn are unmetered; the product edit dialog cannot clear a field; the app's skill editor still carries the retired private visibility.
  • Cloud sync, left for later: there is still no Sync now action — the cadence is the fifteen-minute scan, so a reconnected account waits for the next run. A config whose owner leaves the organization is still deactivated silently by a different door, and a source-deleted item is still a status stamp with no bell.
  • Documents indexed before 0.5.27 keep one vector per repeated passage until they are re-indexed; the content hash is unchanged, so only an explicit retry-indexing (or a content change) re-embeds them.
  • The rail's navigation memory has had part of its manual round: the R5 round drove six EN/DE/FR desktop and phone cases covering parts of NAV-F16–NAV-F19; the remaining section, the second-account cases and NAV-B6–NAV-B9 are still unrun.
  • A reply-language directive is a directive: a model may still answer in the prompt's language and nothing on the wire marks a slip.
  • No image input on the REST chat send. A vision model reads an image over REST only on a thread the app continued with an image attachment; the design of an attachments field on the send is recorded as contract debt.
  • No REST door authors or deploys an automation — POST /automations answers 405 by design. Build and deploy in the app, or over the MCP endpoint's save_automation and deploy_automation; the REST key lists, reads, runs, answers asks and wires triggers.
  • The x-tale-pagination extension is a declaration on the OpenAPI document; generated clients that do not read vendor extensions still branch on the two cursor names until cursor is retired.
  • The app's zip upload of a skill bundle rewrites the bundle and moves updatedAt even when the zip is byte-identical, where PUT /skills/{slug} writes nothing.
  • A tool call the reply cap cut keeps input: {} on the stored tool-call part; the raw text the model emitted is still not on the transcript.
  • Folder names written before 0.5.24 keep their bytes; a sync engine's hub-path lookup can create an NFC twin beside a legacy NFD folder. No backfill ships.
  • Two bounded document readers still filter after their cut; both report an honest truncated, so a caller can tell the answer was cut.
  • Behind a Docker-published port, every IPv6 client arrives as the bridge gateway's address and shares one per-address rate-limit bucket and one audit address until the daemon runs with ip6tables and the reverse proxy's network is IPv6-enabled — an operator item, documented on the Own Compose page.
  • Recorded as contract debt, each with its design in the ledger: a queued send is invisible on the message list until a worker opens it; a webhook delivery the deployed inputs schema refuses moves no trigger stamp; the MCP run_deployed tool keys its idempotency apart from start_run and REST; a page is fetched three to four times per scan; a cancelled run answers trace: null and effects: null where a failed run answers both; approvals have no REST twin; a task cannot be archived or deleted over REST; a webhook bind does not say whether the deployed inputs schema admits a delivery; an exhausted repeatUntil is only a trace note; Website carries no scanStartedAt and the crawler has no page cap, path filter or stop verb of the caller's; website search has no dense leg and its substring fallback stamps score: 0; no Idempotency-Key on the task start; no queue position on a queued send; a corrupt Office document still fails as indexer_error and is retried five times where a PDF lands malformed; no /.well-known/security.txt; no changelog feed on tale.dev; no SDK, collection or per-code table beyond the Error.code enum; GET /notifications rows carry type as a free string and nothing pushes them to a machine caller; a skill keeps no version history on the machine door; the per-task circuit breaker is not built; the messages a conversation snapshot applied are readable only in the app; a run carries no usage or cost.

Migration notes

  • One migration, 0112 (0112_user_preferences_chat_model_provider.sql): ALTER TABLE app.user_preferences ADD COLUMN IF NOT EXISTS chat_model_provider_slug text, plus a column comment. Nullable, no default, no backfill, no index, no constraint — a metadata-only change that takes the table lock for an instant. The file is safe to re-run. It is rolling-deploy safe in both directions: a pick saved before the column exists keeps resolving by id alone, exactly as before, and the previous image ignores the column. The application database moves from 0111 to 0112; the knowledge database is unchanged, and Better Auth adds no column.
  • Migrations are tracked by file name. 0.5.38 shipped a second file with the 0108 prefix (0108_approvals_one_pending_conversation_draft.sql, one partial unique index); the boot applies every .sql file in name order once and records each by name, so a deployment crossing 0.5.38 runs that file at boot beside this release's 0112.
  • No environment variable is added or removed; .env.example is unchanged. No scheduled job is added or retired. One new audit action, provider_definition.deleted (category security); two app-door error codes, PROVIDER_IN_USE (409) and PROVIDER_NOT_FOUND (404 on the new delete door); no machine-contract code changes.
  • The shipped catalogs change with this release: nine model catalogs (cache prices, the DeepSeek and OpenRouter entries, the base prices listed above) and six provider definitions (the three new harness doors; comments on DeepSeek, Z.ai and Qwen). They are read-only image inputs, replaced on upgrade; Refresh catalogs skips them. An organization's own provider files are untouched, though an admin can now author, edit and delete them from the app; the app's deletions archive under TALE_CONFIG_DIR/<org>/providers/.history/<name>/ like every save.
  • The sandbox gateway's per-model pricing is re-pushed on the next session that names a model whose stored patch lacks a cache rate the catalog now has; the shipped gateway image (sandbox-llm-gateway) itself is unchanged and already accepts the cache fields. A Claude Code automation on DeepSeek, Moonshot, the Vercel gateway or OpenRouter provisions the connector's …__anthropic record and its pricing on its first turn after the upgrade.
  • No image in the stop-gated tier changes. The proxy and db images carry no source change, and the managed proxy policy the CLI renders is unchanged. A plain tale deploy is the whole upgrade: no --stop, no downtime window.
  • The platform image (the settings dialog, the composer and transcript, the billing, both agent hosts, the routing, the catalogs), the docs image (six edited pages, each in English, German and French; no new page), the ui-docs image (the Select sentence and the changed @tale/ui component) and the sandbox-runtime image (tale-vision) carry source changes. The web, db, proxy, sandbox, sandbox-buildkitd, sandbox-egress and sandbox-llm-gateway images carry no source change. A deploy pulls the new runtime image and retags it as the spawner's tale-sandbox-runtime:latest; a session container that is already running keeps the image it was started from until the spawner recreates it.
  • The CLI has no change of its own in this range, but the reference tree it embeds — the platform's core and shared modules, where the cost formula, the catalog normalizer, the model choice and the shared provider schema live — does, so the release executables are rebuilt and differ from 0.5.39; they report 0.5.41. The shared provider schema's objects are strict, so a CLI older than this release refuses an organization catalog entry that names the new cacheReadCentsPerMillion or cacheWriteCentsPerMillion keys; nothing the app writes today does. A managed deployment should move its pinned CLI reference together with its platform reference, as always.
  • @tale/ui and @tale/marketing-ui are pinned by this release as the ui-v0.5.41 and marketing-ui-v0.5.41 tags on their snapshot branches; a consumer outside the monorepo installs "@tale/ui": "github:tale-project/tale#ui-v0.5.41". @tale/ui changes in this range (Select.errorMessage), so ui-v0.5.41 differs from ui-v0.5.40 and ui-v0.5.39, which are content-identical to each other; @tale/marketing-ui does not change, so marketing-ui-v0.5.41 is content-identical to its predecessors.

Upgrading

  • On the 0.5 line (0.5.0 – 0.5.39; there is no 0.5.40 deployment to be on):

    tale update
    tale deploy

    Migration 0112 is applied at boot. Nothing in this release needs --stop. A deployment crossing 0.5.39 runs migration 0111 at boot as well; one crossing 0.5.38 runs that release's 0108_approvals_one_pending_conversation_draft.sql, one crossing 0.5.37 runs migration 0110, one crossing 0.5.36 runs migration 0109, one crossing 0.5.35 runs migration 0108 (0108_trusted_header_keys.sql) and Better Auth's session column, and one crossing 0.5.33 runs migration 0107. A deployment crossing from a version older than 0.5.29 should read that release's notes, which do: its proxy image change is only applied by a --stop deploy.

  • Before you upgrade, check what depends on the catalogs and the spend figures. Anything pinned to deepseek-v4-flash — a chat pick, an agent's supportedModels, a governance default, an automation llm node, a REST caller — must be re-pointed at deepseek-flash (or deepseek/deepseek-v4.1-flash on OpenRouter); it answers CHAT_MODEL_UNKNOWN after the deploy. A budget rule that was tuned to the old, overstated figures books less from now on; a report that reads usage_ledger cost figures sees two pricing rules across the upgrade. A Claude Code automation on DeepSeek, Moonshot, the Vercel gateway or OpenRouter rides a different gateway record after the deploy. A custom provider whose /models refused an anonymous listing lists correctly once the organization holds a default key for it.

  • Managed deployments move by pinning the CLI and the runtime to this release's commit, preparing a new bundle and applying it with the pinned CLI — see Managed deployments on the CLI install page. The bundle's backend-local phases run under the interpreted CLI (cli/tale.mjs) that the setup-cli action and bun run --filter @tale/cli build produce beside the executable; the executable from the release page has no interpreted bundle beside it and cannot prepare a managed bundle. On a Linux x64 host whose CPU lacks AVX2, pass linux-baseline: 'true' to the setup-cli action so the bundle embeds the baseline executable.

  • New install:

    curl -fsSL https://raw.githubusercontent.com/tale-project/tale/main/scripts/install-cli.sh | bash
    mkdir tale-05 && cd tale-05
    tale init
    tale deploy

    On a CPU without AVX2 the downloaded executable aborts with Illegal instruction; build it from source with bun run build:linux-baseline in tools/cli.

What's Changed

  • fix(platform): list DeepSeek V4.1 Flash as deepseek-flash by @larryro in #3429
  • fix(platform): build the sandbox runtime image in bun dev when missing by @larryro in #3430
  • feat(platform): define custom AI providers from the settings UI by @larryro in #3431
  • fix(platform): ride the Anthropic harness lane in workflow agent turns by @larryro in #3432
  • fix(platform): bill cache hits at vendor rates and stop retrying 402 by @larryro in #3433
  • feat(sandbox): add per-request vision thinking control by @yannickmonney in #3434
  • feat(platform): pick the embedding model from the provider's catalog by @larryro in #3435

Full Changelog: v0.5.39...v0.5.41