Repository navigation
v0.62.0
Cotal 0.62.0 is out!
Grab it: npm i cotal-ai@0.62.0
Changes in this release
-
Add a delegated user intent: SPEC §13.16, a design record, and the host decisions it needs. A signed-in user admits one launch or one retirement on the host's authenticated route, and a platform control holder consumes that intent once from its current registration, epoch, assignment revision and lifecycle.
@cotal-ai/coreadds the closed request types and parsers for the user's intent and the holder's execution,parseRemoteDelegatedUserIntentExecutionResult, andresolveReadAcl, the read-list resolutionprovisionAgentDurablesalready used.@cotal-ai/authaddsauthorizeDelegatedUserIntentAdmissionandauthorizeDelegatedUserIntentExecution, the record, pin and flight types,delegatedUserIntentHoldsAliasandjoinOrStartDelegatedUserIntent,AuthServiceHandle.observeManagerGate, which a platform composition's handle carries so the host reads the holder's gate over the context's own connection, andAuthServiceHandle.activateManagedLifecycle, which activates a delegated launch's lifecycle at its pinned UID before any row or durable and mints nothing, so a launch the host compensates or the user retires reaches the terminal barrier even when its agent never exchanged. Admission derives nothing from the request: it takes the owner the host derived from the verified IdP subject, requiresspawnon the user's own fresh ledger row, binds the account's current platform assignment and the holder's open gate, and dry-runs the envelope walk from the user's principal over the read list the writer will provision. Execution compares the request with the record and with the fresh assignment, gate epoch and registration proof, and returns the user's owner and parent for the host's writers. The holder gains no grant, no decision readssupervise, and theplatform-controlview's same-owner rule is unchanged. Stock dispatch refuses both kinds asunimplemented; a host that owns an intent store and the enrollment and retirement writers composes the decisions on its own routes.@cotal-ai/manageraddsremoteAuthority.executeDelegatedUserIntent,StartAgentOpts.delegatedIntentandManager.retireDelegatedAgent: a delegated launch takes the hosted enrollment arm with the intent's execution and refuses material under any owner but the intent's, andretireDelegatedAgentstops a delegated agent only after the host confirmsretired: truefor its exact target and operation id. A delegated agent's name, including one whose launch failed at any step after the host's enrollment answer, stays held until then. -
The cotal config accepts a
modelPolicythat names, per role, the models (and optionally the variants) a seat in that role may launch on. The manager refuses a detached spawn, andcotal spawnrefuses a foreground one, before anything is minted when the effective role has an entry and the effective model is missing or not on its list, or its variant is missing or off a declared variants list, or the launch carries launch options (fromlaunchOptions:or--opt), which the connector applies unread after the model and which can select another model. Ids compare whole, sovendor/model-B-fastdoes not satisfyvendor/model-B. The refusal names the persona, whether the value came from its ownmodel:orvariant:field or from--modelor--variant, the value found, and the values allowed. Roles without an entry are unconstrained, a space-local entry replaces the operator-level entry for the same role, and a malformed policy (including an unsupported field of any name) fails the spawn loudly. Before this, a persona pinned to a superseded model, or to none, launched on it with no report. -
Recover an ordinary resume whose coordinator and manager were both lost after the retained agents launched. The coordinator journals
resume-activeonly after the manager's launch returns, so a crash in between leftresume-intentbehind a live agent, and the replacement manager refused it asalready live and this runtime cannot authoritatively adopt it, leaving the journal degraded and the agent owned by nobody. On a static mesh the replacement manager now reads the seat reference the lost manager recorded on that agent's slot, reaps the seat through its runtime, waits for the principal to leave presence, and launches it again. A live principal with no such record, or one that stays live after the reap, is still refused. Tmux handles now carry a reference bound to the tmux server, window and pane, and the tmux runtime can reap by it, so a successor manager also closes the window of a tmux seat it retires instead of leaving it running. The reap closes that window even when the agent's pane already exited, sinceremain-on-exitor a second pane keeps the window open. It refuses a pane that has moved out of that window, whether it still runs or has exited, and a window that has moved out of the session, including one that moves while the reap runs. -
A host stack overflow in a cotal-lang run is no longer catchable. A builtin that ran out of stack (for example
json.stringifyon an array nested a few thousand deep) used to raise a catchable L4016, so a program could branch on how much stack its host had, and a journal recorded on one host was refused with L5001 when resumed on a host with a larger stack. Both the tree-walker and the compiled engine now unwind the run through it, the same as the other faults a program cannot catch: afinallydoes not run past it, aparallel,raceor other scope it fails inside settles nothing, even when another branch failed first, the overflow itself cancels no sibling, and aconclavewhose body overflowed does not close. A resume on a host with more stack then proceeds instead of replaying a recorded scope failure. -
cotal spawn --detachrun from a managed seat's own shell on a static or open mesh now launches as that seat, so the manager records the seat as the spawner and the seat can stop the child withcotal_despawn, as it can acotal_spawnchild. Before, the CLI minted a one-shot operator instrument that no session could present again, and the seat's despawn was refused.--on <instance>keeps its pin on that path: on a static mesh the CLI mints a one-shotmanager-callerview for the seat, pinned to that instance and carrying the spawn subject only when the seat's own credential holds it. On an open mesh the seat's call keeps the TLS requirement the mesh records, and--serverwith an unregistered--spacekeeps the operator path. Without--spacethe seat's target is picked as the operator path picks it, skipping a recorded mesh that is not running. The child is now stopped when the seat exits, and on a static mesh a seat withoutcapabilities: [spawn]is refused. The seat-scoped control target thatcotal runalready used on a static mesh moves to@cotal-ai/workspaceasresolveSeatControlTarget;cotal runkeeps using it on a static mesh only. On a static mesh the seat's credential also proves its space, so a seat launched withoutCOTAL_SPACEstill acts as itself, ascotal run answerdid before; an open mesh acts as the seat only whenCOTAL_SPACEnames its space. See docs/UPGRADING.md. -
A provenance line (
→ using,→ wrote,→ removed) no longer fails or crashes a command when stderr is broken. When the stderr write throws or fails with an error such as EPIPE or ENOSPC, at once or after waiting in a full pipe whose reader goes away, the line is printed on stdout with the error and the command finishes its work. This holds for a write that throws a value that is not an Error, such asnullor an object with no string form. It also holds when the program has already ended stderr, even if it exits in the same tick. When stdout fails too, the command still finishes its work and exits 1 instead of 0, so a line no channel could carry is never lost silently. The same holds for a line still waiting in a full stderr pipe when the command exits, as when the CLI exits on a closed stdout. A program that ends itself on a stdout error, as the CLI does, still stops at that error instead of finishing its work. Node does not say which stderr bytes are still waiting, so a line that had to wait and got through just before the exit counts the same while later stderr output still waits. A connector seed run with a broken stderr used to commit every payload and then exit 1 with nothing said, or stop after the first payload and requirecotal ext seed --repair. A stderr that is closed or redirected away at launch still discards the line. -
Bare
cotal downstops a stack whose manager runs the built-inptyruntime and has never started an agent. Since theptyruntime stopped using a seat custodian, its manager published no spare capability at all, so barecotal downrefused every default stack, left the broker running and kept the space registered, and the nextcotal upof that space was refused as already in use. Ctrl-C on a foregroundcotal upnow holds the manager's stop reservation from its capability check until the manager exits, ascotal downdoes, so a concurrentcotal downcannot stop that manager or arm a reap while the Ctrl-C stop is in flight. -
Let bare
cotal downand Ctrl-C on a foregroundcotal upstop a stack whose manager runs the built-in in-processptyruntime. That manager published no spare capability, so barecotal downrefused to signal it even with no agents running, left the broker up, and kept the registry entry; a defaultManager.stop()threw on any pty seat. A default stop now stops and deprovisions the pty seats that live inside the manager process, since they cannot outlive it, and still releases every seat that can. The manager always publishes its spare capability, which now records whether its stop also stops in-process seats, anddownreports those seats as stopped instead of left running. An older CLI refuses the new record rather than misreport those seats. A stopping manager refuses new spawns and waits for the ones it already accepted, so no seat launches after a stop that reported success, anddownno longer promises that agents will be spared when it cannot list them. RepeatedManager.stop()calls share one stop, so a second call no longer reports success while the first still waits for a seat to exit. -
Stop the connector's inbox overflow valve from dropping direct messages. When a busy session's bounded inbox filled with directed mail, the valve evicted the oldest DM and, after five evictions of the same message, acknowledged it and wrote one stderr line. The sender had seen the message stored and the recipient was live and rostered, yet its inbox never carried it. An evicted direct message or role request is now never acknowledged, and the broker redelivers it after the ack wait until it finds room. A direct message stays pending on the recipient's durable, where
cotal deliver pending <name>counts it; a role request stays on its role's shared queue, which that command does not read. Evicted channel traffic is still acknowledged as before. -
A
cotal_dmreply to a peer that messaged you and has no roster row, such as a one-shotcotal send, is now stored under that sender's id in the space's DM history instead of failing withno peer "<name>" in space "<space>". The sender is matched by the exact id on the message you hold, or by its display name while the presence view is current. A name that two such senders share is refused with their ids. The receipt readsrecipient had no roster row at sendand says the DM may never reach an inbox, since a one-shot send has already exited. An operator's DM view, such as the dashboard's Direct messages lens, shows the reply. -
A platform composition can now ask whether its assigned control manager is serving. With the
platformControlinput,startAuthServicereturns a handle withplatformControlReadiness(instanceId), which answers that manager instance'sstatusreply and refuses any instance the current assignment does not name, or one whose gate another owner holds. The auth context reads over its own connection, whose grant is that instance'sdescribeandstatusand its own reply rail. The connection renews in process like the context's other connections and never leaves it, so a pooled control host no longer needs a human or operator credential, a per-read control instrument, or the manager's process id to tell whether the manager is up. -
Unpinned CLI calls to a manager no longer fail when the class queue sends the describe and the invoke to different managers. A manager that receives a call bound to another instance refuses it before running it, and the CLI's manager commands (
models,stop,spawn --detachand the rest),cotal invokeand the manager row ofcotal statusnow re-describe and re-issue after that refusal, up to 16 times, instead of printing it. Before this, a space with three managers failed about two calls in three. The refusal still surfaces once every attempt has split, and a call pinned with--onis never re-issued. -
A foreground
cotal upnow restarts the delivery daemon it started when that daemon dies while the broker is still running, and logs that it did. The daemon ends itself once it cannot reach the broker, and a starved host can make a running broker look unreachable: under heavy load the daemon loggedbroker connection unavailable past backstopand exited, nothing brought it back, and from then on every stopped seat's retirement failed on thectl.delivery-adminrail and every spawn of that name was refused as reserved pending retirement until an operator re-rancotal up. A failed restart is retried after the 30-second delivery lease TTL, the longest a dead holder's lease can block its replacement. Recovery is announced only once the replacement's responder is bound. A daemon that exits cleanly or on SIGTERM or SIGINT stays stopped, and so does one stopped withcotal down delivery, also whendownhas to kill a starved daemon or the stop lands between two restart attempts. -
cotal meshes --jsonprints one JSON object per recorded mesh per line, so a script no longer has to split the table, whose ROOT column can contain spaces. A row carriesspace,server,mode,root,defaultandorigin(up,manualorcatalog). A local or hand-registered entry also carriesoffline. A discovered entry is never probed, so it has none.tlsRequired,eventsand a discovered entry'scatalogNameappear when the record has them. An empty registry prints nothing and exits 0, and the note about a default that matches no record goes to stderr.meshes addandmeshes rmrefuse--json.The first-run connector seed now prints each
✓ addedline to stderr. Before, the seed wrote them to stdout ahead of the command that triggered it, so a firstcotal meshes --jsonstarted with seven lines that were not JSON. -
The up-resume-render-lock live smoke now also resumes a foreground
cotal upafter a secondcotal down --preserve-state. A foreground resume that skips recreating the memory-backed presence bucket now turns it red, the way a--detachresume that skips it already did. -
In a space with more than one static manager,
inspectfor a seat hosted by another manager instance now says so. The class queue hands the read to either instance, and the one that did not host the seat answerednot-found: no agent "<name>", the same reply as a name that exists nowhere. A namedcotal_despawnresolves its target through this read, so it refused a live seat with an error that read as absence whenever the lookup reached the other manager. On a live-map miss the manager already reads the durable slot; when a nonretired slot names a sibling instance as its owner it now answersfailed-preconditionwith the slot detail plusownerInstanceId, and a message that names the owning instance. A name with no slot, or only a retired one, is stillnot-found. -
A manager that starts while the delivery daemon is still binding, or while the daemon has stepped back to re-check its lease, now waits up to 60 seconds for the
ctl.delivery-adminrail to answer instead of exiting on the first unanswered request. The wait covers the three boot steps that need the daemon: the SecretStore challenge, the boot self-heal's freeze-holder liveness check, and the re-registration's verified eviction of the superseded serve family. Before,cotal upcould start a manager inside that window, the manager exited withcould not challenge the delivery daemon's SecretStoreor withre-registration could not revoke + verify-evict the superseded serve family; the gate is left frozen, and nothing restarted it until an operator rancotal upagain. Only a request that times out or finds no responder is retried. Retries come closer together as the wait runs out, so a daemon that binds in its last seconds is still asked, and the wait ends on time even while a retry is still connecting, with nothing sent after it. A daemon that answers still fails the start at once when it refuses or sends a reply the manager cannot read, and a daemon that stays silent for the whole wait still fails it with the same message as before. Renewal passes andcotal reconcile-gatekeep their single attempt. -
The manager no longer reissues a numbered spawn name. Before, once
reviewer_2despawned and its name was released, the nextreviewercollision gotreviewer_2again, so one name labeled two different agents in rosters and channel history. A numbered name is now issued once per manager process and a later collision takes the next number. Base persona names stay reusable, a hard-pinned--nameis unaffected, and a restarted manager starts its record of issued numbers empty. -
A
cotal supervise --rosterentry now takes ashare-tools:list that narrows the operator's declared MCP servers for that agent, the same selectionspawn --share-toolsmakes: an absent key keeps every declared server,[]shares none, and an undeclared name fails that entry. Before this the roster loader dropped the key, so every rostered agent launched with the full pool. A value that is not a list, or a name the flag cannot carry unchanged such asnonealone or one with a comma or surrounding spaces, fails the roster load.
What's Changed
- fix(cli): re-issue an unpinned manager call after a pre-effect bind refusal by @davidfarah2003 in #2476
- fix(connector-core): store a DM reply to a sender with no roster row by @davidfarah2003 in #2493
- fix(docs): publish the Upgrading page and check site links locally by @davidfarah2003 in #2502
- fix(manager): name the owning instance when inspect misses a sibling's seat by @davidfarah2003 in #2503
- fix(manager): wait out an unanswered delivery-admin rail at boot by @davidfarah2003 in #2507
- fix(ci): refuse a pull request whose approval does not name its head by @davidfarah2003 in #2522
- feat(cli): add --json to cotal meshes by @davidfarah2003 in #2527
- fix(cli)!: launch a detached spawn from a seat's shell as that seat by @davidfarah2003 in #2534
- feat(auth): add a read-only readiness read for the platform control host by @davidfarah2003 in #2533
- fix(manager): let bare down stop a pty manager that has never started an agent by @davidfarah2003 in #2526
- fix(manager): refuse a spawn whose model is outside the role's model policy by @davidfarah2003 in #2530
- fix(smoke): kill the stacks a suite leaves in its sandbox when it exits by @davidfarah2003 in #2525
- fix(manager): honor a roster entry's share-tools selection by @davidfarah2003 in #2536
- fix(manager): never reissue a numbered spawn name by @davidfarah2003 in #2543
- fix(workspace): move a provenance line to stdout when stderr fails by @davidfarah2003 in #2547
- test(cli): resume a foreground cotal up in the resume render lock live smoke by @davidfarah2003 in #2550
- fix(manager): stop in-process pty seats on shutdown so bare down stops the stack by @davidfarah2003 in #2557
- docs(readme): cover workflows, sign-in meshes, skills and the current packages by @davidfarah2003 in #2567
- fix(connector-core): never ack a directed message evicted from a full inbox by @davidfarah2003 in #2568
- fix(ci): end mutation reproof inside a run budget and name unfinished fixtures by @davidfarah2003 in #2570
- feat(auth): execute a signed-in user's launch or retirement intent through a platform control holder by @davidfarah2003 in #2574
- fix(cli): restart the delivery daemon when it dies under a running broker by @davidfarah2003 in #2579
- fix(lang): make host stack exhaustion uncatchable by @davidfarah2003 in #2581
- fix(manager): reclaim a retained seat a lost manager launched during resume by @davidfarah2003 in #2603
Full Changelog: v0.61.0...v0.62.0