Tale v0.5.32
0.5.32 carries 4 merged pull requests. The large one teaches Tale to pace its embedding work to what an embedding server can actually take: how many requests it may have in flight, and how long a request may take once the server's own queue is counted — both stated by the operator, next to the model. Beside it, self-hosting gains a Deploy on Kubernetes guide with five apply-ready manifests, a code file the browser names no type for can be attached and sent again, and the chat budget banner becomes an announced, readable Alert in the composer's own column. No migration, no contract change, and no image in the stop-gated tier: the upgrade is tale update followed by a plain tale deploy.
Highlights
Embedding requests are paced to the server's stated capacity (#3402)
Knowledge indexing sent a document's 64-text batches with up to three requests in flight per Embedder, a flat 60-second timeout per request and three attempts. Every indexing job, website scan and search builds its own Embedder, so concurrent jobs multiplied that bound: five jobs put fifteen requests on a server sized for three.
On a hosted embedding API that never showed. On a server you run yourself — one that computes a request at a time and queues the rest — a request queued behind three full batches is answered after about two minutes. Tale abandoned it at 60 seconds and sent it again into the same queue, so large documents failed to index, and the pile-up starved whatever else shared those GPUs.
Two optional settings in embedding.json now describe the server, and everything else follows from them.
-
maxConcurrentRequests(1–64, default 3) is how many requests may be in flight per organization, endpoint and model, shared by every embedder in the process rather than counted per embedder. Requests wait in arrival order, except that a search query — someone is waiting on it — takes the next free slot ahead of queued batches, and never interrupts one in flight. A lowered limit applies at once; a raised one applies once the requests started under the old limit have finished. A failed, timed-out or abandoned request always frees its slot, and the slot is held across a request's own retries so a retrying batch does not rejoin the queue at its end. The organization is part of the key because a lane is a queue callers wait in, and nothing one organization does may hold up another's. -
minTokensPerSecondis the rate the server sustains for this model under its usual load — measured with the other work on that hardware running, but not counting time a request waits behind other requests, because Tale allows for that itself. Stated, it sizes each request's timeout as queue wait plus compute: the request's own tokens, estimated generously from its characters, plus2 × maxConcurrentRequests − 1other requests each counted as at least a full batch of 64 texts of 1,024 tokens (and at least this request's own size), divided by the rate and multiplied by 1.5. Never under 60 seconds, never over the ceiling.
The ceilings are the budgets that already existed. A batch gets at most 15 minutes — the budget of one rag.index_file attempt. A search query gets at most 5 minutes, for each request and for the whole search including its wait for a slot and any pause between attempts: a query runs inside a chat turn's tool call, which does not move the turn's heartbeat, and the chat generation watchdog takes a turn for dead after ten minutes of stillness. A guard test holds the search ceiling at or below half that window so the turn has time to report the search's own error instead of being swept as "interrupted by a restart". Without a stated rate, nothing says how long the server's queue may take, so a request may use its whole ceiling.
Retries no longer pile work onto a busy server. A timed-out request is not sent again: the server held it the whole time, and a repeat only lengthens its queue. A refused connection, a rate limit or a server error is retried after a growing pause, or after the pause the server asked for with Retry-After when that is longer; a server that asks for more than a minute is left alone and the caller's own slower retry decides when to come back. The OpenAI SDK's internal retries stay off, so this module's loop is the one retry policy.
One failed batch stops its siblings. The batches of a call run on a signal of their own, combined with the caller's through AbortSignal.any. The first batch to fail cancels the running ones and leaves the queued ones unsent, so no slot is held for vectors nobody will store, and that first error is the one the caller sees — the aborts it caused never replace it.
An expired indexing job stops instead of racing its own retry. The worker now hands each handler pg-boss's job signal, which pg-boss aborts when a job outlives its queue's budget or the worker shuts down. rag.index_file stops before its next slice and cancels the request in flight. It writes no failed status — the row keeps its progress — so the retry resumes after the stored slices rather than running beside the attempt it replaces; a retry that never comes is settled by the RAG watchdog as before.
The CLI speaks the same three ways. knowledge-embedding declares both fields and checks convergence per setting: a value sets it, an omitted key leaves whatever the file holds, and null clears it. A change confined to minSimilarity, maxConcurrentRequests or minTokensPerSecond leaves every stored vector valid, so it skips the empty-corpus refusal that a model change still triggers.
Both settings are documented with a worked example under Pace requests to a self-hosted embedding server on the Data residency and knowledge storage page, and in the knowledge-embedding section of the CLI install page.
Deploy on Kubernetes (#3404)
A new self-hosted install page, Deploy on Kubernetes, in English, German and French. It carries five apply-ready manifest files for a single namespace, installed and upgraded by one envsubst '${VERSION}' | kubectl apply loop:
00-namespace.yaml— namespace, shared Secret, theconfig-dataclaim10-stores.yaml— the Postgres StatefulSet behind thedbandknowledge-dbServices, and MinIO20-application.yaml— API, worker, web tier, video-token provider, and the backend egress NetworkPolicy30-proxy.yaml— Caddy onhostPort80/443 with its certificate claim40-sandbox.yaml— the egress proxy, the LLM gateway with both Services, and the spawner's ServiceAccount, Role, RoleBinding and Deployment
Around them: the prerequisites (a NetworkPolicy-enforcing CNI, a StorageClass, RWX or a single node for config-data, node sizing, the entry point, image access, a RuntimeClass for nested Docker), the namespace layout, the Service names, the probe table, what the spawner enforces, and the commands that prove a deployment — the migration count, both policies, health, fence probes from a session Pod, and a rolling restart with two replicas.
The page exists because the whole 0.5.31 stack was run on a kind cluster — ten services, the spawner on SANDBOX_BACKEND=kubernetes, first owner, an agent task with a deliverable, session lifecycle including idle stop and resume, cross-replica access and a rolling restart. Five contract gaps came out of that run and are the load-bearing warnings on the page:
enableServiceLinks: falseis mandatory. A Service namedsandboxinjectsSANDBOX_PORT=tcp://…and the spawner exits parsing it;DB_PORTcorrupts the URLenv.shderives.- The platform image's
NET_ADMINiptables fence self-severs on a Pod network — only scope-link routes are accepted, so even CoreDNS is rejected (getaddrinfo EAI_AGAIN db). The page usesTALE_SKIP_SSRF_FIREWALL=1together with a backend egress NetworkPolicy, which is the fence on this path. - Session Pods must reach
backend-apiinside the spawner policy's namespace allowance, so the application roles live in the sandbox namespace. - That same allowance exposes the stores' Service ports to session Pods. Credentials are still required, and it is documented as a known limit.
- Caddy listens on the port named in
SITE_URL, so an address with a non-standard port also needs that port as the proxy'scontainerPortandhostPort.
Run Compose yourself keeps the service contract — names, volumes, probes, environment — and now points at the new page instead of carrying the Kubernetes tables itself; the install index links both.
A code file the browser names no type for can be sent again (#3403)
Attaching a source or config file the browser reports no MIME type for — render.cjs, and the same for .mjs, .go, .rs, .sh — made the send fail. resolveFileType answered '', the composer staged fileType: '', and both chat send doors refuse an empty type (z.string().min(1) on POST …/messages and on its deferred-send twin), so the request came back 400 {"error":"invalid body"} in about ten milliseconds. Indexing was fine throughout; only the send was refused.
resolveFileType never answers an empty content type now. Six files had already patched the hole with their own || 'application/octet-stream' — eight copies in all, three of them in the upload hook alone, including the metadata write three lines above the staged attachment, which is why the stored row said application/octet-stream while the wire said ''. The fallback lives in the resolver, the copies are gone, and the store and the wire carry one value. It widens nothing: application/octet-stream is in no allowlist, and the text, RAG and size gates key on the extension.
The second half is what the failure looked like. The refusal branch and the start-throw catch already cleared the optimistic send, but the rejection branch restored the composer text and left the bubble and its Thinking · Ns shell on screen, so a refused send read as a turn that never answers — until a reload showed the message gone. A rejected turn now drops its overlay with the text it gives back.
The chat budget banner is announced, and readable (#3405)
The "Usage limit reached" banner above the chat composer is a @tale/ui Alert in the composer's own column instead of a full-width strip with a bare bottom border.
The strip was designed for the top of the chat pane and kept those classes when it was moved above the composer, so it spanned the pane while the composer and the deferred-send tray sit in a centred column. Worse, the exceeded state painted the whole line in the destructive colour: 3.3:1 on its own tint in light mode, below the WCAG AA floor of 4.5:1. The Alert doctrine puts the accent on the fill, the border and the glyph and keeps the copy at foreground contrast. And the banner had no live region at all, so the hard block — the state in which sending is refused — was never announced; Alert supplies role="alert" with aria-live="polite".
The period was also interpolated raw, which produced "2,000 of 10,000 token left this monthly", "setzt sich monthly zurück" and "se réinitialise monthly". The four period strings in English, German and French now use an ICU select: today / this week / this month for what is left, and daily / weekly / monthly for when it resets. No key was added or removed.
A menu reopens on the click right after a pick (#3404)
Shipping with the Kubernetes page, because it was what kept the end-to-end lane red: Radix keeps a closed dropdown menu mounted while it animates out, and during that window the old content is still a dismissable layer whose own trigger counts as "outside". Picking an item and clicking the trigger again inside the exit animation toggled the menu open and the layer's outside-dismiss closed it again — the click was lost, and nothing opened. The end-to-end language spec did it reliably fast and failed on two of three runs; a person does it whenever they are quick.
@tale/ui's shared DropdownMenu wrapper now ignores a left-button pointer-down that lands on the menu's own trigger, so the trigger owns that click. The regression test drives real Chromium with the exit animation pinned: uncontrolled and controlled menus reopen when clicked right after an item was picked, and a plain trigger click still closes an open menu.
Behaviour changes
- Embedding requests to one organization's model share one in-flight bound per endpoint and model across every indexing job, website scan and search in a Tale process. Unset, that bound is 3 — the value each embedder used before, now shared instead of multiplied — so an organization that states nothing sends fewer concurrent requests than it did on 0.5.31, not more.
- An embedding request that runs out of time is no longer retried. Connection refusals, rate limits and server errors still are, and a
Retry-Afterlonger than one minute ends the request instead of waiting it out. - When one batch of a multi-batch embedding call fails, the rest are cancelled or never sent, and the first error is the one reported.
- An indexing job that reaches its 15-minute attempt budget, or whose worker is shutting down, now stops and cancels its request in flight. It records no
failedstatus; the retry resumes after the slices already stored. - Without
minTokensPerSecond, a single embedding request may now take up to 15 minutes (5 for a search query) instead of being cut at 60 seconds. This is the intended change — the flat minute was what abandoned work the server was about to finish — but a genuinely unreachable server is now noticed later than it was. - A file whose name and browser report give no MIME type is stored and sent as
application/octet-streamrather than an empty string. Nothing is admitted that was not admitted before: that type is in no upload allowlist, and the text, RAG and size gates decide on the file extension. - A chat turn whose send is rejected clears its optimistic message and thinking shell instead of leaving them on screen.
- The chat budget banner renders in the composer's column as an
Alertwithrole="alert", and its period phrasing is selected per locale rather than interpolated. - A
@tale/uiDropdownMenureopens when its trigger is clicked during the menu's exit animation. Any consumer of the package inherits this; inside this repository only the platform renders the component.
API contract changes
- None. The OpenAPI document stays at 1.14.0 with 80 paths, 127 operations and 58 schemas, and the
Error.codeenum keeps its 150 values. No operation was added, removed or renamed, and no request or response shape changed — regenerating the specification from this commit reproduces the committed file and its contract fingerprint exactly. - The two new embedding settings are organization configuration in
embedding.json, reached through the app's knowledge administration and the CLI'sknowledge-embeddingresource. They are not on the/api/v1surface.
Security
- No security advisory is fixed in this release, and no dependency in this range carries one. Two packages are added:
p-queue9.3.0 as a platform dependency, andp-timeout7.0.1 — which p-queue depends on — pinned through the rootoverridesbecause 7.0.2 is younger than Renovate's 90-day minimum release age. - The MIME fallback widens nothing (#3403).
application/octet-streambelongs to no upload allowlist, and the text, RAG and size gates key on the file extension. What changed is that the stored row and the wire now carry the same value instead of disagreeing. - The Kubernetes page turns off one fence and names its replacement (#3404).
TALE_SKIP_SSRF_FIREWALL=1is required there because the platform image's in-container iptables fence self-severs on a Pod network; the backend egress NetworkPolicy in20-application.yamlis what constrains backend egress instead, and the page says so rather than leaving the variable unexplained. The page also states, as a known limit, that the spawner's namespace allowance lets session Pods reach the stores' Service ports — credentials are still required, and a two-namespace layout is the recorded follow-up. - An announced hard block (#3405). The budget banner's exceeded state is the one in which sending is refused. It had no live region, so a screen-reader user met a disabled composer with no explanation; it is now a
role="alert"region whose copy also meets AA contrast.
Known issues
- The embedding pacing is proved against a controlled server, not a long soak. The lane, the priority, the timeout formula, the retry rules and the abort behaviour are covered by 58 automated cases plus observed runs of the production Node runtime against a local HTTP server — two embedders sharing a limit of 2 never exceeding it at the server, an abandoned request's connection closing, a search overtaking queued batches, a
Retry-After: 2honoured at 2,006 ms, and a failed batch stopping its siblings. What a real self-hosted model server does under a day of production load remains a deployment check. - The bound is per Tale process, not per deployment. The API and every worker replica can each reach
maxConcurrentRequests. Size the value for one process and count your replicas; the timeout formula already allows for one more process with the same bound, and an operator who shares a server more widely than that should state a lower rate. minTokensPerSecondis a statement, not a measurement. Nothing verifies it. Set it too high and requests are cut before the server finishes; too low and a genuinely dead server is noticed late. Its worked example on the data-residency page is the recommended starting point.- The Kubernetes page's verified scope is one cluster. A fresh single-node kind cluster on Kubernetes 1.36, kube-network-policies, the local-path StorageClass, and Tale 0.5.31 — every Pod ready, 104 migrations counted, both policies present, the edge at 200 with HTTP→HTTPS at 308, the full session lifecycle and an agent task with a deliverable. A managed distribution, a different CNI or a multi-node cluster will need its own run of the page's own verification commands, and the page lists the scope it was proved at.
config-dataneeds RWX or a single node. The shared configuration volume takes writes and locks from more than one role; the page says so, and a cluster without a suitable StorageClass cannot follow it as written.- Tale still ships no Helm chart. The page is manifests and
envsubst, deliberately, and it is not an operator: the CLI's Docker rollout coordination does not run a Kubernetes deployment for you. - Unchanged from v0.5.31, where each is described in full: a managed deployment picks up the proxy policy added in that release only when a newly prepared bundle is applied; the transcription setting is only as good as the organization's credentials; the six agent-turn fixes are bounded by the pinned Claude Code build they were read from; the 0.5.29 proxy change has been exercised live in
TLS_MODE=letsencryptonly; the web tier's backend-URL default lives in the image, not in the generated compose; the scheduled-pack fix does not reach an automation an organization already has; a budget hold covers a turn's first round only; a run still carries no usage or cost; nothing backfills a task timeline. - Unchanged from v0.5.20, where each is described in full: the
es/co-ccColombian cédula detector still ships switched off and a locale-agnostic PII toggle still widens national-ID matching to every locale; thinking-block replay on the native Anthropic connector is not done and the live Max-plus-tool-call check is still owed;rag_searchembedding calls inside a harness turn are unmetered; the product edit dialog cannot clear a field; the app's skill editor still carries the retiredprivatevisibility. - Cloud sync, left for later: there is still no Sync now action — the cadence is the fifteen-minute scan, so a reconnected account waits for the next run. A config whose owner leaves the organization is still deactivated silently by a different door, and a source-deleted item is still a status stamp with no bell.
- Documents indexed before 0.5.27 keep one vector per repeated passage until they are re-indexed; the content hash is unchanged, so only an explicit
retry-indexing(or a content change) re-embeds them. A site that has not been scanned since 0.5.27 has no stored robots rules until its next scan. - The rail's navigation memory has had part of its manual round: the R5 round drove six EN/DE/FR desktop and phone cases covering parts of
NAV-F16–NAV-F19; the remaining section, the second-account cases andNAV-B6–NAV-B9are still unrun. - A reply-language directive is a directive: a model may still answer in the prompt's language and nothing on the wire marks a slip.
- No image input on the REST chat send. A
visionmodel reads an image over REST only on a thread the app continued with an image attachment; the design of anattachmentsfield on the send is recorded as contract debt. - No REST door authors or deploys an automation —
POST /automationsanswers 405 by design. Build and deploy in the app, or over the MCP endpoint'ssave_automationanddeploy_automation; the REST key lists, reads, runs and wires triggers. - The
x-tale-paginationextension is a declaration on the OpenAPI document; generated clients that do not read vendor extensions still branch on the two cursor names untilcursoris retired. - The app's zip upload of a skill bundle rewrites the bundle and moves
updatedAteven when the zip is byte-identical, wherePUT /skills/{slug}writes nothing. - A tool call the reply cap cut keeps
input: {}on the storedtool-callpart; the raw text the model emitted is still not on the transcript. - Folder names written before 0.5.24 keep their bytes; a sync engine's hub-path lookup can create an NFC twin beside a legacy NFD folder. No backfill ships.
- Two bounded document readers still filter after their cut; both report an honest
truncated, so a caller can tell the answer was cut. - Behind a Docker-published port, every IPv6 client arrives as the bridge gateway's address and shares one per-address rate-limit bucket and one audit address until the daemon runs with
ip6tablesand the reverse proxy's network is IPv6-enabled — an operator item, documented on the Own Compose page. - Recorded as contract debt, each with its design in the ledger: a queued send is invisible on the message list until a worker opens it; a webhook delivery the deployed
inputsschema refuses moves no trigger stamp; the MCPrun_deployedtool keys its idempotency apart fromstart_runand REST;robots.txt$end-anchors andAllow:lines are not honoured (prefix and*rules are), and a page is fetched three to four times per scan; a cancelled run answerstrace: nullandeffects: nullwhere a failed run answers both; approvals and asks have no REST twins; a task cannot be archived or deleted over REST; a webhook bind does not say whether the deployedinputsschema admits a delivery; an exhaustedrepeatUntilis only a trace note;Websitecarries noscanStartedAtand the crawler has no page cap, path filter or stop verb of the caller's; website search has no dense leg and its substring fallback stampsscore: 0silently; noIdempotency-Keyon the task start; no queue position on a queued send; a corrupt Office document still fails asindexer_errorand is retried five times where a PDF landsmalformed; no/.well-known/security.txt; no changelog feed on tale.dev; no SDK, collection or per-code table beyond theError.codeenum;GET /notificationsrows carrytypeas a free string and nothing pushes them to a machine caller; a skill keeps no version history on the machine door; the per-task circuit breaker is not built; the messages a conversation snapshot applied are readable only in the app.
Migration notes
- No migration. The application database stays at 0106 and the knowledge database is unchanged, so nothing runs at boot beyond the usual convergence check.
- No new environment variable.
.env.exampleand the environment reference are unchanged.TALE_SKIP_SSRF_FIREWALL, which the Kubernetes page uses, has shipped since long before this release; the page documents it, it is not new. - No new configuration file, but an organization's existing
embedding.jsonaccepts two more optional keys,maxConcurrentRequestsandminTokensPerSecond. A file that states neither behaves as before, except that the default bound of 3 is now shared across the process rather than counted per embedder. Both follow the platform's preserve-or-clear rule: a value sets it, an omitted key leaves what the file holds, andnullclears it. - No image in the stop-gated tier changes. The
proxyanddbimages carry no source change in this range, and neither does the managed proxy policy the CLI renders — unlike 0.5.31, nothing about the proxy needs a newly prepared bundle. A plaintale deployis the whole upgrade: no--stop, no downtime window. - The platform image (the embedding lane, the job signal, the MIME fallback, the chat surface and budget banner) and the docs image (five pages in each of English, German and French — the new Kubernetes guide, the install index, Run Compose yourself, the CLI install page and data residency — plus the navigation and search index) carry source changes. web and ui-docs rebuild only because
@tale/uimoved; neither renders aDropdownMenu, so nothing in them behaves differently. Thedb,proxy,sandbox,sandbox-runtime,sandbox-buildkitd,sandbox-egressandsandbox-llm-gatewayimages carry no source change. - The CLI has source changes in this range (#3402's
knowledge-embeddingschema and convergence), so a managed deployment should move its pinned CLI reference as well as its platform reference. The release executables report 0.5.32. @tale/uiand@tale/marketing-uiare pinned by this release as theui-v0.5.32andmarketing-ui-v0.5.32tags on their snapshot branches; a consumer outside the monorepo installs"@tale/ui": "github:tale-project/tale#ui-v0.5.32".@tale/uicarries the dropdown-menu trigger fix, so it is not content-identical toui-v0.5.31;@tale/marketing-uihas no source change in this range and its tag is content-identical to its predecessor.
Upgrading
-
On the 0.5 line (0.5.0 – 0.5.31):
tale update tale deploy
Nothing in this release needs
--stop. A deployment crossing from a version older than 0.5.29 should read that release's notes, which do: itsproxyimage change is only applied by a--stopdeploy. -
Managed deployments move by pinning the CLI and the runtime to this release's commit, preparing a new bundle and applying it with the pinned CLI — see Managed deployments on the CLI install page. Pin the CLI reference too: this range changes it. The bundle's backend-local phases run under the interpreted CLI (
cli/tale.mjs) that thesetup-cliaction andbun run --filter @tale/cli buildproduce beside the executable; the executable from the release page has no interpreted bundle beside it and cannot prepare a managed bundle. On a Linux x64 host whose CPU lacks AVX2, passlinux-baseline: 'true'to thesetup-cliaction so the bundle embeds the baseline executable. -
New install:
curl -fsSL https://raw.githubusercontent.com/tale-project/tale/main/scripts/install-cli.sh | bash mkdir tale-05 && cd tale-05 tale init tale deploy
On a CPU without AVX2 the downloaded executable aborts with
Illegal instruction; build it from source withbun run build:linux-baselineintools/cli. -
Running your own embedding server? Nothing is required — an organization that states nothing keeps working, with a bound of 3 now shared across the process instead of multiplied by it. But if indexing has been failing on large documents, this is the release to state the two facts about that server in
embedding.json: how many requests it can take at once, and the rate it sustains. Pace requests to a self-hosted embedding server, on the Data residency and knowledge storage page, works an example through. -
Deploying on Kubernetes? The Deploy on Kubernetes page is new in this release, under Self-hosted › Install. Read its prerequisites before applying anything — a NetworkPolicy-enforcing CNI, a StorageClass, and RWX or a single node for
config-dataare requirements, not recommendations — and run its verification commands before admitting users.
What's Changed
- fix(platform): pace embedding requests to the server's stated capacity by @yannickmonney in #3402
- fix(platform): send a chat attachment the browser gave no MIME type by @larryro in #3403
- docs(docs): add the Kubernetes deployment guide by @larryro in #3404
- fix(platform): restyle the chat budget banner to the composer column by @larryro in #3405
Full Changelog: v0.5.31...v0.5.32