v0.41.0
Cotal 0.41.0 is out!
Grab it: npm i cotal-ai@0.41.0
Changes in this release
-
/api/activityreads all of chat in one stream read instead of one read per channel. The CHAT
stream already interleaves every channel into one sequence space, so the newest N across the space
is the tail of that one stream. The newCotalEndpoint.multiChannelHistory(channels, opts)reads it
with a single consumer whosefilter_subjectsare those channels, and tags each message with the
channel the broker delivered it on rather than with the payload's own claim. A channel left out of
the list is filtered by the broker and never crosses the link, so the dashboard's chat-only rule now
decides the read's filter set instead of deciding which of seventy reads to issue. No new broker
authority: a multi-filter create rides the bare$JS.API.CONSUMER.CREATE.<CHAT>row the observer
and admin profiles already hold.The multi-filter read batches its filter list. One consumer create naming every subject is a single
request, but a large enough request stops being answerable: on a seeded sweep the create succeeded
at 70, 1,000 and 5,000 channels and failed at 10,000, and what gives way is the client's request
timeout rather than the broker.MULTI_FILTER_BATCHis 1,000 andMULTI_FILTER_READ_CONCURRENCYis
4, so the list is split and at most four batches are in flight, and the pages are merged by stream
sequence before the newest N are taken. The size comes from the sweep rather than from the failure
point: a fifth of the largest count that still answered, answering in about a quarter of a second, so
a batch stays far from both the create timeout and the 8000ms response deadline. On the same corpus
the sweep goes from 43ms, 246ms, 5345ms and a timeout, to 54ms, 343ms, 1163ms and 942ms. Those walls
are single runs on one host and are shaped by it rather than being bounds; a re-run on another box
landed outside the last of them, as the paragraph below says of every span here. A space
with the 69 chat channels this work was aimed at is one batch, so the request counts and byte totals
are unchanged at that size. The walls are not claimed to be, for the reason just given.
Merging on stream sequence rather than the payload'stsis asserted in
packages/core/smoke/history-recent.smoke.tsagainst a corpus whose arrival order andtsorder are
opposite.Two limits remain and are stated rather than buried. On a 20,000 channel and 20,000 message corpus
the batched read still fails at 5,000 filters and above, because batching removes the create size
ceiling and not the cost of a batch matching a small fraction of a long stream; no unbatched
numbers were taken on that corpus, so no comparative claim is made about it. Separately, a filter
list big enough to exceed the connection'smax_payloadis refused by the client before anything
is sent, rather than by the broker and rather than timing out: the client compares the request
against the limit the broker advertised at connect and throws, so no bytes leave and the broker's
own logs show nothing. A list of 40,000 filters on this corpus's subject shape is refused that way.
Where the limit falls moves with both the advertisedmax_payloadand the length of the subjects,
so 40,000 is a length that was tried and not a threshold.Measured on the wire by a proxy between the endpoint and the broker that counts
$JS.APIrequest
lines and bytes per direction, over a seeded corpus of 69 chat channels plus 24 event channels at
limit 200. The before column is that same committed suite file, with itspackage.jsonscript,
copied onto a544a974b7checkout and run there, so it measures the old code with the new
instrumentation: 2524 broker requests and 8,015,332 to 8,016,039 bytes across four runs to return a
143,401-byte page. After: 143 requests and 908,410 to 908,422 bytes across eight runs, for the same
200 entries in the same order, and consumer creates fall from 347 to 6.The counts are the stable claim and the walls are illustrative, and the two halves of that sentence
have different evidence. On the completed no-link arms the request counts, the consumer creates and
the page size are identical in every run, while the byte totals move by tens of bytes. The
field link's two arms behave differently and the sentence has to say which. Its fan-out arm is
truncated by the deadline, so its requests, creates and bytes vary with the clock: 794 to 849
requests and 128 to 133 creates across thirteen runs. Its single read completes, and keeps the same
143 requests and 6 creates it spends with no link cost, in every run. Walls vary on one host by more than the counts do. Across an 82ms RTT / 554 KB/s
link with the shipped 8000ms deadline, the fan-out answered 16 of 70 sources in every run and timed
out, and the single read answers whole in 4304ms to 4637ms across nine runs. On the same link the pre-change tree spends 768
to 793 requests and 2,085,533 to 2,364,182 bytes across four runs to reach those 16 sources.Every span here is what its runs covered on one host. None of them is a bound. A reviewer ran the
suite four more times while this was under review and landed outside eight of the spans as they then
stood, by a byte or two on the totals and by tens of milliseconds on the walls, and the spans below
have been widened to include those runs. Expect the next run to do it again. What does not move, in
any run by anyone so far, is the request counts, the consumer creates and the page sizes of the
completed no-link arms, the 143 requests and 6 creates the single read also spends on the field
link, and the number of sources each field-link arm reaches: 2 of 2 for the single read, 16 of 70
for the truncated fan-out. The fan-out's request and create counts are not in that set, because the
deadline truncates them. 24 of 70 belongs to the pre-change tree's uncapped fan-out and to nothing
at this head, which completes at 2 of 2 uncapped on the same 143 requests. The argument rests on the
figures that do not move.pnpm smoke:web-activity-read-costreproduces the after column, and beside it a frozen copy of the
old fan-out shape as a scale-invariance control, which costs 2863 requests and 7,743,782 to
7,744,228 bytes across eight runs. That
arm runs the old shape on this build's one-page window, so it is a control and not the shipped
before; the before column is the same suite run against544a974b7.History reads now open their first window one page wide instead of four.
drainWindowdelivers
everything in the window and keeps the tail, so a four-page window moved four pages to return one
whenever the subject was most of its stream./api/dmsat limit 500 against a 2500-message backlog
moved 1,995,854 to 1,995,859 bytes across four runs to return a 346,001-byte page, at 257 requests
every run, and took 8161ms to 8753ms alone on that link. Every one of those four runs missed the
8000ms deadline with nothing else on the connection. An earlier draft published 8852ms and 7857ms
here and called the read a straddle; 7857ms is below every run I can now produce, so the straddle is
withdrawn and the read simply misses. It now moves 502,354 to 502,359 bytes with no link cost across eight runs,
and alone on the field link takes 2532ms to 2858ms across twelve runs while moving 502,361 to
502,365 bytes. Those are two different arms of the suite and the split is deliberate: the byte
figure a reader should compare against the before column is the no-link one. A sparse subject pays
one more widening step for that.multiChannelHistoryrefuses a channel name the wire layer would rewrite.chatSubjectbuilds each
filter throughtoken(), which maps an unusable character to_, trims each segment and drops
empty ones, sofoo/barwould have filtered onfoo_barand.leadonlead: the caller names one
channel and the broker returns another.isConcreteChanneldoes not catch it, because none of those
carry a wildcard. The read now runs the sameassertValidChannelthe policy path already used
against this aliasing, before it builds the subject. The dashboard passes canonicallistChannels()
rows and never reached it; the public method and its exact-filter promise did.ActivitySource(the seamactivityBackfillreads through) replaceschannelHistorywith
multiChannelHistory. The aggregation now has two sources, chat and DMs, so a partial page names
which half is missing rather than naming individual channels.The activity feed selects different messages under clock skew, and an operator can see it. The
page is the newestlimitchat messages by broker arrival, unioned with the newestlimitDMs,
ordered byts. The shape it replaces took the newestlimitper channel and then the newest
limitof that union byts. With two channels andlimit=2, channel A holding a1 (seq 1, ts 100)
and a2 (seq 2, ts 200) and channel B holding b1 (seq 3, ts 50) and b2 (seq 4, ts 60), the old rule
returns a1 and a2 and the new rule returns b1 and b2. They disagree whenever a sender's clock
disagrees with arrival order by more than the spread between thelimit-th andlimit+1-th message.
The display order is unchanged, since the page is still sorted byts; what changed is which
messages reach it. Arrival is the broker's own record andtsis a sender claim, so the new key is
the narrower one. Per-channel selection cannot be had from a single read at any window short of one
that has seenlimitmessages from every requested channel, since a quiet channel's newest can sit
arbitrarily far back in an interleaved stream. That is a property of the mechanism rather than a
measurement, and it is why the old selection was not kept. -
Distinguish an unpopulated presence watch from a current roster.
presenceView()now returns a discriminatedcurrent,unpopulated, orstalestate. An
unpopulated reconnect refill carriesfresh: false, so existing consumers fail toward unknown
instead of treating a partial roster as an authoritative absence verdict.waitForPresenceSnapshot()
now reports whether the snapshot completed or the bounded wait timed out.The web dashboard surfaces an unpopulated roster as still loading and keeps the existing stale-watch
diagnostic for a view that was populated and later went silent. -
cotal-lang DX: the conformance corpus and the language card. Every js block in the language reference is generated into a JSON artifact shipped inside @cotal-ai/lang (conformance/corpus.json) with the verdict the validator gives it, served by a new conformanceCorpus() accessor, so a second implementation can run the same claims from the file alone; pnpm gen:conformance regenerates it and smoke:lang-conformance holds the shipped bytes identical to a fresh build from the reference. The artifact states its own adjudication rule, so a reader holding only the JSON knows a refusal is checked by membership in the validator's answered codes, never by equality with a single code. docs/lang-card.md is a one-page card of the language (effects and their results, the await rule, branch keys, top refusals), validated block by block like the reference itself, carried in the connector docs bundle, and published on the docs site beside the other reference pages.
-
Pin every stack pidfile to its process's creation identity before teardown.
upwrites a sibling<pidfile>.identitycontaining the pid and process start where the OS reports one. Every stop path checks it before signalling: a reused pid or torn pin is refused and preserved, and rerunning after the process is stopped clears the stale record automatically. A live pre-pin record warns and proceeds so the first teardown after an upgrade still works; relaunching writes the pin and enables full match and mismatch protection. -
Keep agent-profile minting within one resolved mesh root.
cotal mintnow reads the persona ACL, loads the signing authority, and stores the default credential under the selected mesh root. If the current folder holds trust for a different space or account, it refuses before writing and names both roots instead of combining authority material from one root with persona policy from another. -
Resolve the persona catalog from the target mesh rather than the current directory.
cotal personaslisted the personas of whatever directory it ran in, whilecotal spawnlaunched from the mesh it resolves — so from one directory the two could name completely different sets, with neither saying anything was wrong. Everycotal personassubcommand now reads and writes the resolved mesh's catalog, which also makes--spaceand--serverreal for the listing rather than only for the live--runningoverlay: naming a mesh now moves the catalog, and an unresolvable target refuses instead of silently acting on another directory's files.cotal spawn --role/--subscribeandcotal send msg/askcomplete from that same catalog.The library functions behind this (
personasDir,listPersonas,listPersonaNames,listDeclaredChannels,listDeclaredRoles) now require an explicit root instead of defaulting to the current directory, so a caller that omits one fails to compile rather than answering about the wrong place. -
Seed personas into the catalog
cotal spawnreads, and name that directory in the output.cotal setupwrote.cotal/agents/default.mdunder the directory it ran in, whilecotal spawnloads its persona from the mesh it resolves. On a machine where those differ — a shell outside any project, plus a mesh whose root is elsewhere — setup created a file spawn would never open, sono default persona yet - run cotal setup to seed onesurvived running exactly the command it named. Setup now seeds into the resolved mesh's catalog, including when that mesh was registered against a brand-new directory with no.cotalin it yet.Every seed states its destination as an absolute path, and when the mesh's root is not the current directory both are shown, so the choice is visible rather than assumed. With no mesh running at all the current directory is still the answer — setup has to work before the first
cotal up— but it says that it fell back and why. With several meshes running and none selected it refuses and asks you to pick, instead of choosing a root on your behalf.cotal spawn's refusal now names the absolute directory it searched and the mesh that directory came from, so a persona that is missing from one catalog and present in another is diagnosable from the message itself. -
cotal statusnow names the root behind every persona row, and flags the case where the folder you are standing in is not the one a barecotal spawnwill use.Status could print
personas defaultin green under "This Folder" whilecotal spawnrefused in the same second with "no default persona yet". Both were right about their own root and neither said which root that was: the folder's catalog is<root>/.cotal/agents, while spawn loads the resolved mesh's, and the two diverge whenevercotal use, a--space, or a registry entry points elsewhere. The personas status listed and the personas spawn could launch could be completely disjoint.When the two roots differ, status now names both, says what the spawn root actually offers, and drops the green from a
defaultthat will not launch. When they agree, the output stays as short as it was.
What's Changed
- fix(cli): resolve the persona catalog, seed destination and status root from the target mesh by @davidfarah2003 in #1135
- feat(lang): ship the conformance corpus and the language card by @davidfarah2003 in #1194
- ci: give the smoke shards a real timeout budget by @davidfarah2003 in #1224
- fix(core)!: a refilling presence view must not report itself fresh by @davidfarah2003 in #1228
- perf(web): serve /api/activity from one stream read instead of one per channel by @davidfarah2003 in #1212
- docs(changeset): describe the multi-filter batching that shipped with the single read by @davidfarah2003 in #1235
- feat(workspace,cli)!: pin pidfiles to process start identity before teardown (#969) by @davidfarah2003 in #1069
Full Changelog: v0.40.0...v0.41.0