Skip to content

v0.62.0

Choose a tag to compare

@davidfarah2003 davidfarah2003 released this 05 Oct 01:08
5d8ac75

Cotal 0.62.0 is out!

Grab it: npm i cotal-ai@0.62.0

Changes in this release

  • Add a delegated user intent: SPEC §13.16, a design record, and the host decisions it needs. A signed-in user admits one launch or one retirement on the host's authenticated route, and a platform control holder consumes that intent once from its current registration, epoch, assignment revision and lifecycle. @cotal-ai/core adds the closed request types and parsers for the user's intent and the holder's execution, parseRemoteDelegatedUserIntentExecutionResult, and resolveReadAcl, the read-list resolution provisionAgentDurables already used. @cotal-ai/auth adds authorizeDelegatedUserIntentAdmission and authorizeDelegatedUserIntentExecution, the record, pin and flight types, delegatedUserIntentHoldsAlias and joinOrStartDelegatedUserIntent, AuthServiceHandle.observeManagerGate, which a platform composition's handle carries so the host reads the holder's gate over the context's own connection, and AuthServiceHandle.activateManagedLifecycle, which activates a delegated launch's lifecycle at its pinned UID before any row or durable and mints nothing, so a launch the host compensates or the user retires reaches the terminal barrier even when its agent never exchanged. Admission derives nothing from the request: it takes the owner the host derived from the verified IdP subject, requires spawn on the user's own fresh ledger row, binds the account's current platform assignment and the holder's open gate, and dry-runs the envelope walk from the user's principal over the read list the writer will provision. Execution compares the request with the record and with the fresh assignment, gate epoch and registration proof, and returns the user's owner and parent for the host's writers. The holder gains no grant, no decision reads supervise, and the platform-control view's same-owner rule is unchanged. Stock dispatch refuses both kinds as unimplemented; a host that owns an intent store and the enrollment and retirement writers composes the decisions on its own routes. @cotal-ai/manager adds remoteAuthority.executeDelegatedUserIntent, StartAgentOpts.delegatedIntent and Manager.retireDelegatedAgent: a delegated launch takes the hosted enrollment arm with the intent's execution and refuses material under any owner but the intent's, and retireDelegatedAgent stops a delegated agent only after the host confirms retired: true for its exact target and operation id. A delegated agent's name, including one whose launch failed at any step after the host's enrollment answer, stays held until then.

  • The cotal config accepts a modelPolicy that names, per role, the models (and optionally the variants) a seat in that role may launch on. The manager refuses a detached spawn, and cotal spawn refuses a foreground one, before anything is minted when the effective role has an entry and the effective model is missing or not on its list, or its variant is missing or off a declared variants list, or the launch carries launch options (from launchOptions: or --opt), which the connector applies unread after the model and which can select another model. Ids compare whole, so vendor/model-B-fast does not satisfy vendor/model-B. The refusal names the persona, whether the value came from its own model: or variant: field or from --model or --variant, the value found, and the values allowed. Roles without an entry are unconstrained, a space-local entry replaces the operator-level entry for the same role, and a malformed policy (including an unsupported field of any name) fails the spawn loudly. Before this, a persona pinned to a superseded model, or to none, launched on it with no report.

  • Recover an ordinary resume whose coordinator and manager were both lost after the retained agents launched. The coordinator journals resume-active only after the manager's launch returns, so a crash in between left resume-intent behind a live agent, and the replacement manager refused it as already live and this runtime cannot authoritatively adopt it, leaving the journal degraded and the agent owned by nobody. On a static mesh the replacement manager now reads the seat reference the lost manager recorded on that agent's slot, reaps the seat through its runtime, waits for the principal to leave presence, and launches it again. A live principal with no such record, or one that stays live after the reap, is still refused. Tmux handles now carry a reference bound to the tmux server, window and pane, and the tmux runtime can reap by it, so a successor manager also closes the window of a tmux seat it retires instead of leaving it running. The reap closes that window even when the agent's pane already exited, since remain-on-exit or a second pane keeps the window open. It refuses a pane that has moved out of that window, whether it still runs or has exited, and a window that has moved out of the session, including one that moves while the reap runs.

  • A host stack overflow in a cotal-lang run is no longer catchable. A builtin that ran out of stack (for example json.stringify on an array nested a few thousand deep) used to raise a catchable L4016, so a program could branch on how much stack its host had, and a journal recorded on one host was refused with L5001 when resumed on a host with a larger stack. Both the tree-walker and the compiled engine now unwind the run through it, the same as the other faults a program cannot catch: a finally does not run past it, a parallel, race or other scope it fails inside settles nothing, even when another branch failed first, the overflow itself cancels no sibling, and a conclave whose body overflowed does not close. A resume on a host with more stack then proceeds instead of replaying a recorded scope failure.

  • cotal spawn --detach run from a managed seat's own shell on a static or open mesh now launches as that seat, so the manager records the seat as the spawner and the seat can stop the child with cotal_despawn, as it can a cotal_spawn child. Before, the CLI minted a one-shot operator instrument that no session could present again, and the seat's despawn was refused. --on <instance> keeps its pin on that path: on a static mesh the CLI mints a one-shot manager-caller view for the seat, pinned to that instance and carrying the spawn subject only when the seat's own credential holds it. On an open mesh the seat's call keeps the TLS requirement the mesh records, and --server with an unregistered --space keeps the operator path. Without --space the seat's target is picked as the operator path picks it, skipping a recorded mesh that is not running. The child is now stopped when the seat exits, and on a static mesh a seat without capabilities: [spawn] is refused. The seat-scoped control target that cotal run already used on a static mesh moves to @cotal-ai/workspace as resolveSeatControlTarget; cotal run keeps using it on a static mesh only. On a static mesh the seat's credential also proves its space, so a seat launched without COTAL_SPACE still acts as itself, as cotal run answer did before; an open mesh acts as the seat only when COTAL_SPACE names its space. See docs/UPGRADING.md.

  • A provenance line (→ using, → wrote, → removed) no longer fails or crashes a command when stderr is broken. When the stderr write throws or fails with an error such as EPIPE or ENOSPC, at once or after waiting in a full pipe whose reader goes away, the line is printed on stdout with the error and the command finishes its work. This holds for a write that throws a value that is not an Error, such as null or an object with no string form. It also holds when the program has already ended stderr, even if it exits in the same tick. When stdout fails too, the command still finishes its work and exits 1 instead of 0, so a line no channel could carry is never lost silently. The same holds for a line still waiting in a full stderr pipe when the command exits, as when the CLI exits on a closed stdout. A program that ends itself on a stdout error, as the CLI does, still stops at that error instead of finishing its work. Node does not say which stderr bytes are still waiting, so a line that had to wait and got through just before the exit counts the same while later stderr output still waits. A connector seed run with a broken stderr used to commit every payload and then exit 1 with nothing said, or stop after the first payload and require cotal ext seed --repair. A stderr that is closed or redirected away at launch still discards the line.

  • Bare cotal down stops a stack whose manager runs the built-in pty runtime and has never started an agent. Since the pty runtime stopped using a seat custodian, its manager published no spare capability at all, so bare cotal down refused every default stack, left the broker running and kept the space registered, and the next cotal up of that space was refused as already in use. Ctrl-C on a foreground cotal up now holds the manager's stop reservation from its capability check until the manager exits, as cotal down does, so a concurrent cotal down cannot stop that manager or arm a reap while the Ctrl-C stop is in flight.

  • Let bare cotal down and Ctrl-C on a foreground cotal up stop a stack whose manager runs the built-in in-process pty runtime. That manager published no spare capability, so bare cotal down refused to signal it even with no agents running, left the broker up, and kept the registry entry; a default Manager.stop() threw on any pty seat. A default stop now stops and deprovisions the pty seats that live inside the manager process, since they cannot outlive it, and still releases every seat that can. The manager always publishes its spare capability, which now records whether its stop also stops in-process seats, and down reports those seats as stopped instead of left running. An older CLI refuses the new record rather than misreport those seats. A stopping manager refuses new spawns and waits for the ones it already accepted, so no seat launches after a stop that reported success, and down no longer promises that agents will be spared when it cannot list them. Repeated Manager.stop() calls share one stop, so a second call no longer reports success while the first still waits for a seat to exit.

  • Stop the connector's inbox overflow valve from dropping direct messages. When a busy session's bounded inbox filled with directed mail, the valve evicted the oldest DM and, after five evictions of the same message, acknowledged it and wrote one stderr line. The sender had seen the message stored and the recipient was live and rostered, yet its inbox never carried it. An evicted direct message or role request is now never acknowledged, and the broker redelivers it after the ack wait until it finds room. A direct message stays pending on the recipient's durable, where cotal deliver pending <name> counts it; a role request stays on its role's shared queue, which that command does not read. Evicted channel traffic is still acknowledged as before.

  • A cotal_dm reply to a peer that messaged you and has no roster row, such as a one-shot cotal send, is now stored under that sender's id in the space's DM history instead of failing with no peer "<name>" in space "<space>". The sender is matched by the exact id on the message you hold, or by its display name while the presence view is current. A name that two such senders share is refused with their ids. The receipt reads recipient had no roster row at send and says the DM may never reach an inbox, since a one-shot send has already exited. An operator's DM view, such as the dashboard's Direct messages lens, shows the reply.

  • A platform composition can now ask whether its assigned control manager is serving. With the platformControl input, startAuthService returns a handle with platformControlReadiness(instanceId), which answers that manager instance's status reply and refuses any instance the current assignment does not name, or one whose gate another owner holds. The auth context reads over its own connection, whose grant is that instance's describe and status and its own reply rail. The connection renews in process like the context's other connections and never leaves it, so a pooled control host no longer needs a human or operator credential, a per-read control instrument, or the manager's process id to tell whether the manager is up.

  • Unpinned CLI calls to a manager no longer fail when the class queue sends the describe and the invoke to different managers. A manager that receives a call bound to another instance refuses it before running it, and the CLI's manager commands (models, stop, spawn --detach and the rest), cotal invoke and the manager row of cotal status now re-describe and re-issue after that refusal, up to 16 times, instead of printing it. Before this, a space with three managers failed about two calls in three. The refusal still surfaces once every attempt has split, and a call pinned with --on is never re-issued.

  • A foreground cotal up now restarts the delivery daemon it started when that daemon dies while the broker is still running, and logs that it did. The daemon ends itself once it cannot reach the broker, and a starved host can make a running broker look unreachable: under heavy load the daemon logged broker connection unavailable past backstop and exited, nothing brought it back, and from then on every stopped seat's retirement failed on the ctl.delivery-admin rail and every spawn of that name was refused as reserved pending retirement until an operator re-ran cotal up. A failed restart is retried after the 30-second delivery lease TTL, the longest a dead holder's lease can block its replacement. Recovery is announced only once the replacement's responder is bound. A daemon that exits cleanly or on SIGTERM or SIGINT stays stopped, and so does one stopped with cotal down delivery, also when down has to kill a starved daemon or the stop lands between two restart attempts.

  • cotal meshes --json prints one JSON object per recorded mesh per line, so a script no longer has to split the table, whose ROOT column can contain spaces. A row carries space, server, mode, root, default and origin (up, manual or catalog). A local or hand-registered entry also carries offline. A discovered entry is never probed, so it has none. tlsRequired, events and a discovered entry's catalogName appear when the record has them. An empty registry prints nothing and exits 0, and the note about a default that matches no record goes to stderr. meshes add and meshes rm refuse --json.

    The first-run connector seed now prints each ✓ added line to stderr. Before, the seed wrote them to stdout ahead of the command that triggered it, so a first cotal meshes --json started with seven lines that were not JSON.

  • The up-resume-render-lock live smoke now also resumes a foreground cotal up after a second cotal down --preserve-state. A foreground resume that skips recreating the memory-backed presence bucket now turns it red, the way a --detach resume that skips it already did.

  • In a space with more than one static manager, inspect for a seat hosted by another manager instance now says so. The class queue hands the read to either instance, and the one that did not host the seat answered not-found: no agent "<name>", the same reply as a name that exists nowhere. A named cotal_despawn resolves its target through this read, so it refused a live seat with an error that read as absence whenever the lookup reached the other manager. On a live-map miss the manager already reads the durable slot; when a nonretired slot names a sibling instance as its owner it now answers failed-precondition with the slot detail plus ownerInstanceId, and a message that names the owning instance. A name with no slot, or only a retired one, is still not-found.

  • A manager that starts while the delivery daemon is still binding, or while the daemon has stepped back to re-check its lease, now waits up to 60 seconds for the ctl.delivery-admin rail to answer instead of exiting on the first unanswered request. The wait covers the three boot steps that need the daemon: the SecretStore challenge, the boot self-heal's freeze-holder liveness check, and the re-registration's verified eviction of the superseded serve family. Before, cotal up could start a manager inside that window, the manager exited with could not challenge the delivery daemon's SecretStore or with re-registration could not revoke + verify-evict the superseded serve family; the gate is left frozen, and nothing restarted it until an operator ran cotal up again. Only a request that times out or finds no responder is retried. Retries come closer together as the wait runs out, so a daemon that binds in its last seconds is still asked, and the wait ends on time even while a retry is still connecting, with nothing sent after it. A daemon that answers still fails the start at once when it refuses or sends a reply the manager cannot read, and a daemon that stays silent for the whole wait still fails it with the same message as before. Renewal passes and cotal reconcile-gate keep their single attempt.

  • The manager no longer reissues a numbered spawn name. Before, once reviewer_2 despawned and its name was released, the next reviewer collision got reviewer_2 again, so one name labeled two different agents in rosters and channel history. A numbered name is now issued once per manager process and a later collision takes the next number. Base persona names stay reusable, a hard-pinned --name is unaffected, and a restarted manager starts its record of issued numbers empty.

  • A cotal supervise --roster entry now takes a share-tools: list that narrows the operator's declared MCP servers for that agent, the same selection spawn --share-tools makes: an absent key keeps every declared server, [] shares none, and an undeclared name fails that entry. Before this the roster loader dropped the key, so every rostered agent launched with the full pool. A value that is not a list, or a name the flag cannot carry unchanged such as none alone or one with a comma or surrounding spaces, fails the roster load.

What's Changed

Full Changelog: v0.61.0...v0.62.0