Releases: CryptoJym/plimsoll
Release list
Collector 0.7.51: dispatch bridge (#453)
Collector 0.7.51 adds the dispatch bridge.
- Dispatch bridge (#453). The collector keeps a bounded history of which work dispatches started which agent sessions, and binds captured usage to them. Adopting that history is off in this release, and it is refused against the released 0.7.48 reader, so rolling back stays safe.
Updating and rolling back:
- Rolling back to 0.7.50, 0.7.49, 0.7.48 or 0.7.47 keeps every usage row. In the named-usage rollback test, each released reader reopened a ledger sealed by 0.7.51 after its lease expired, and no usage was dropped.
Runtime CLI SHA256: 03b42a2ca1627a6c63bff7f1712b2fed27078efeb5f4b46834d3002e6cc7b7ae.
Qualification: typecheck, the proof suite and system end-to-end qualification ran on the merged commit, on our own macOS runners. Two proofs that drive the real launchd and collector supervision are local-only on those runners, because the runners run a live collector (eco-6hoxj.165.179). 0.7.51 changes no launchd, install or supervision code. The published package is the qualified artifact from that run.
Collector 0.7.50: recorded Codex tier reader (#458) and the studio5 session re-send fix (#460)
Collector 0.7.50 keeps a busy machine from re-sending every session, and gets ready to record how each Codex request was processed.
- No more full session re-sends behind one stuck session (#460). When one session couldn't be sent and had a large backlog of waiting rows, the collector looked at every row before skipping that session. On a busy machine that took longer than its time limit, so the collector fell back to re-sending every session, cycle after cycle. It now checks each waiting session once and stops at its first ready row. The same sessions are sent as before, and every waiting row is kept.
- Ready for the recorded Codex processing tier (#458). The collector now accepts and keeps a recorded Codex processing tier, whether written as
serviceTierorservice_tier, and keeps cached-input counts exactly as Codex stated them. This release doesn't write the tier yet. A later release turns that on, and this one is its safe rollback point.
Updating and rolling back:
- Rolling back to 0.7.49, 0.7.48 or 0.7.47 keeps every usage row. In the named-usage rollback test, the released 0.7.49, 0.7.48 and 0.7.47 readers each reopened a ledger sealed by 0.7.50 after its lease expired, and no usage was dropped.
Runtime CLI SHA256: 700e52f10899111b4372966a976a0e178ef57647590d41e2071a54f70c635251.
Qualification: typecheck, the proof suite and system end-to-end qualification ran on the merged commit, on our own macOS runners. Two proofs that drive the real launchd and collector supervision are local-only on those runners, because the runners run a live collector (eco-6hoxj.165.179). 0.7.50 changes no launchd, install or supervision code. The published package is the qualified artifact from that run.
Collector 0.7.49: bounded Codex capture (#454) and proof IPC fix (#455)
Collector 0.7.49 keeps Codex capture moving on machines with large Codex backlogs, and makes the proof suite steadier.
- Codex capture keeps moving on large backlogs (#454). The collector reads big Codex session files in bounded slices, so new usage keeps arriving while a large backlog is read.
- Steadier proofs (#455). Proof child processes run without the tsx command-line IPC pipe, which removes a random "address already in use" failure in CI. The runtime is unchanged by this item.
Updating and rolling back:
- Rolling back to 0.7.48 or 0.7.47 keeps every usage row. In the named-usage rollback test, the released 0.7.48 and 0.7.47 readers reopened a ledger sealed by 0.7.49 after its lease expired, and no usage was dropped.
Runtime CLI SHA256: 35313c96a999b54fb839002f073784b836abaf3f8a3db918176caaf8481185aa.
Qualification: typecheck, the proof suite and system end-to-end qualification ran on the merged commit, on our own macOS runners. Two proofs that drive the real launchd and collector supervision are local-only on those runners, because the runners run a live collector (eco-6hoxj.165.179). 0.7.49 changes no launchd, install or supervision code. The published package is the qualified artifact from that run.
Collector 0.7.48: Codex usage filed as Codex, keyless plan readings skipped, gap census, and readable commit errors
Collector 0.7.48 files Codex usage as Codex, keeps one unbound plan reading from failing its whole batch, and writes readable diagnostics for the studio3 Codex stall.
- Codex usage is filed as Codex (#449). The collector classifies Codex OTLP data by its service name, even on a machine that also holds a valid Claude credential. Before this, Codex Desktop usage on such machines was filed as Claude Code. The same Codex record sent under two credentials now makes one usage row, not two. Invalid, unknown, malformed and header-mismatched credentials are still refused.
- A plan-limit reading without an account is skipped (#449). Before, it failed its whole batch. Every other record in that batch now arrives.
- Gaps carry a census (#449). Each capture gap now includes a bounded, typed count of what was set aside. If the cloud refuses that field, the collector resends once without it.
- Readable commit errors (#449). A rollout commit exception now logs the exception class, a message hash, and file-handle hashes and offsets. It logs no content, so the next stall can be diagnosed from the log.
Updating and rolling back:
- Rolling back to 0.7.47 keeps every usage row. In the rollback test, a packaged 0.7.47 reopened a 0.7.48 ledger after its lease expired: usage was byte-identical, with zero privacy retirements and no usage dropped.
- Known limit: the studio3 live-tail exception itself was not reproduced offline. This release adds the logging that should show it.
Runtime CLI SHA256: a4ec38fe2b9f8a42941e365098a9d796711e261b7b50f70a7a6f079051217e61.
Qualification: typecheck, the proof suite and system end-to-end qualification ran on the merged commit, on our own macOS runners. Two proofs that drive the real launchd and collector supervision are local-only on those runners, because the runners run a live collector (eco-6hoxj.165.179). 0.7.48 changes no launchd, install or supervision code. The published package is the qualified artifact from that run.
Collector 0.7.47: summary repairs never insert a null session, existing triggers migrate on upgrade, and the status heartbeat stays bounded
Collector 0.7.47 replaces 0.7.46, which was rolled back after its canary. The background session-summary repair no longer fails on a session with no id, ledgers already on 0.7.45 or 0.7.46 get the corrected repair triggers when they update, and the status check stays cheap even after a failed start.
- Summary repairs never insert a null session (#446). The repair triggers from 0.7.45 grouped an OR without parentheses, so some rows reached
session_sync_summary_repairswithout a session id and maintenance failed with a NOT NULL error. That was the 0.7.46 canary failure, also seen on one 0.7.45 host. - Existing triggers migrate on upgrade (#447). An updated ledger replaces the old repair triggers instead of keeping them, so real 0.7.45 and 0.7.46 ledgers stop failing after the update. A crash mid-migration rolls back atomically, and all 23 summary triggers match a clean ledger.
- A maintenance failure stays visible (#447). The failure latch survives failed reads and heartbeats, and clears only after a successful maintenance pass.
- The status heartbeat is bounded (#447). With a status cache, a heartbeat does no SQL; Codex intake stays near 3 ms at a million pending rows.
Updating and rolling back:
- The trigger migration runs in one transaction at start. Rolling back to 0.7.45 restores that version's triggers through its own migration.
- One known limit (eco-6hoxj.163.145): if the collector fails to start before it builds its status cache, its heartbeat does a full status read until the cache exists.
Runtime CLI SHA256: 1d2f30af02426f8b9841884692e5c2230e9629bee64554132da8e0e84a15d3b9.
Qualification: typecheck, the full proof suite and system end-to-end qualification ran on the merged commit. The published package is the qualified artifact from that run.
Collector 0.7.46: install heartbeat with the computer name, and relay-aware capture checks
Collector 0.7.46 replaces 0.7.45. Setup now shows each machine by its own name and whether it is really capturing, even when it is idle. Usage from a tool's own repository finds its project. The capture checks also understand the fleet's standard OpenTelemetry relay.
- A heartbeat with the computer's name (#444).
- The collector checks in about every 15 minutes while it runs, even when idle, and once at start. Each check-in carries the collector version, the last capture time and the capture state.
- Setup shows an idle machine as healthy instead of silent.
- The machine name comes from the computer's name and is trimmed and capped at 64 characters.
- The heartbeat never delays uploads, backs off on errors, and stops quietly if the cloud has no endpoint.
- Relay-aware capture checks (#445). A Claude Code or Codex seat that reports through the fleet's OpenTelemetry relay on 127.0.0.1:4318 counts as live coverage, not as a capture gap. The relay forwards every request to this collector unchanged.
- Usage finds its project from the tool's own repository (#441). When a session has no project of its own, usage takes the repository the tool itself works in, from earlier context in the same session only.
- Plan-limit readings carry the cloud's account key (#432). Each AI account's plan-limit percent arrives under the same key the cloud's Accounts page uses.
- Startup survives a passing lsof glitch (#443). A transient process-probe failure at startup no longer rolls a fresh ledger back, and workers share the probe allowlist.
- Restore retries an inconclusive probe (#440). An inconclusive archive-handle probe is retried before a restore is refused.
- Queued summaries read only the queue (#437). Summary repair reads its queued rows instead of every event.
Updating and rolling back:
- The ledger change is additive: two new tables, created if missing. Rolling back to 0.7.45 leaves them unused.
- The cloud's install-contact endpoint is live (cloud #219). Older clouds answer 404, and the heartbeat then stops quietly until restart.
Runtime CLI SHA256: a2622864d501ccbffc154c1c810238f5df55b0c65c66262ddc29d5e8ca3abc55.
Qualification: typecheck, the full proof suite and system end-to-end qualification ran on the merged commit. The published package is the qualified artifact from that run.
Collector 0.7.45: safer capture and session totals
Collector 0.7.45 replaces 0.7.44. Setting up a machine now takes one command, Claude lane spend carries its work item, and history from folders that were never captured can be brought in. It also keeps every unsent row through retention, repairs busy session summaries in small pieces, and holds up when two parts of Plimsoll start at the same moment.
- Setup in one command (#428).
plimsoll joinregisters the agent session folders it finds, starts or safely restarts the background collector, and waits for the first upload before it says the machine is connected. - Claude lane spend carries its work item (#429). Claude Code hook and OTLP events from a dispatched lane are stamped with the lane's work item, as Codex lanes already are, so spend lines up with the work that caused it.
- Missed history can be imported (#431).
plimsoll capture-roots import-history --root <id>previews, and with--applyimports, session files from before a folder was enrolled. It only imports folders that no capture path ever covered, so nothing is counted twice. - Unsent rows survive retention (#417). Raw retention never prunes a row that still has an upload waiting.
- A fresh ledger keeps its capture roots (#426). A replaced ledger keeps each enrolled folder's saved boundary, so capture resumes where it left off.
- Busy session summaries repair in small pieces (#425). An edit to a long session repairs only its 4,096-row segment instead of rebuilding the whole summary.
- Duplicate-marked rows stop counting (#419). A row marked as a Codex duplicate after its dashboard fact exists is removed from that fact.
- Sturdier sync (#418, #420, #410, #412, #416). Sends end before the cloud's commit deadline, a crashed hook-spool unit replays exactly once, a sync pass shares one storage-retry budget, and a stuck summary recovers and says why in
plimsoll status. - Two processes can open a brand-new ledger at the same moment (#434). The daemon and a hook or CLI starting together on a new install no longer fail with "duplicate column name", and each open waits no longer than its startup deadline.
- Discovery keeps moving on a busy host (#435). A collect that has started finishes a small, bounded amount of folder discovery even when the host is slow, so new sessions are found while a large file is still being read.
Updating and rolling back:
- After the upgrade, each session that is touched rebuilds its summary once. On a copy of a Studio 5 ledger, 8 summaries were complete again in 68 seconds, with no intake loss.
- Rolling back to 0.7.44 is safe. 0.7.44 opens a ledger written by 0.7.45 and keeps its events unchanged.
Runtime CLI SHA256: cd8e8d3099a37a479ddf875465f58057c3ba5aaa3314021d38b5216ae37228ad.
Qualification: typecheck, the full proof suite and system end-to-end qualification ran on the merged commit. The published package is the qualified artifact from that run.
Collector 0.7.44: short writer holds, work links, weekly tool stats, exact Grok cost, 0.7.39 fixes, intake through updates
Collector 0.7.44 replaces 0.7.43: short writer holds, usage linked to work, weekly tool statistics, exact Grok cost, fixes for ledgers that started on 0.7.39, and intake that stays open through updates.
- Writer holds stay short (
.163.102). Periodic WAL truncation no longer waits on readers, and two new indexes shorten the dashboard's usage-authority counts. On a copy of a busy host's ledger, an 11-minute replay spooled 0 of 774 steady requests; 0.7.43 spooled 103 of 777. - Usage links to work (
.163.104). A session bound to a work item charges each eligible event to that work once, and uploads carry opaque work and run references. - Weekly tool statistics (
.165.25). The collector uploads weekly tool-use counts with bounded report sizes, and later weeks stay live. - Exact Grok cost (
.165.57). Grok's vendor cost ticks ride in the upload metadata, so the cloud prices them exactly (1,200 ticks is 120 nanodollars). - Ledgers that started on 0.7.39 (
.163.108).- The per-source admission cap rises from 600 to 3,000 requests a minute. At 600, a busy host refused more than half of its traffic.
- The session-sync delta uses an indexed seek; it used to time out on large ledgers.
- The legacy session-summary rebuild runs in bounded, resumable background passes.
- Intake stays open through updates (
.163.96). During a managed update,plimsoll lifecycle updatestarts a small listener on the collector's port that spools hook posts and OTLP exports;load-launch-agentreleases it just before the new daemon starts. The new daemon replays each spooled item once, and probe events never count as spend. An update that doesn't complete releases the listener before it exits, so the restored runtime's daemon can bind the port. In the proof, 5 of 5 sends during the stop were accepted and replayed once; 0.7.43 refused all 5. The gap with no receiver fell from 16.6 s to about 3.7 s.
Updating and rolling back:
- Rollback to 0.7.43 is supported. On a disposable home that 0.7.44 had updated and run, 0.7.43 passed snapshot list and prune, status, the daemon with hook and OTLP intake, export, and session sync (38 of 38 checks). Every event and the update receipt were kept, and 0.7.44's new tables, indexes and triggers stayed in place. 0.7.44 then reopened the same ledger and accepted intake again. 0.7.43 doesn't recognise 0.7.44's update operations, so its prune keeps both 0.7.44 snapshots; nothing is lost.
- The first start builds two indexes before readiness. On an 85 GB ledger that took about 20 minutes (1,237 s), so hosts with very large ledgers should wait for the follow-up that moves this work after readiness. Check readiness time on each host after its update.
- Only managed updates are covered by the listener. Sends during any other stop are still not recorded; token usage in session files is captured again after the restart.
Runtime CLI SHA256: 6dbca6f2e57191f8d7e59762cf80ed15aaea22373056d83069da42cab808d677.
Qualification: typecheck, the proof suite and system end-to-end qualification ran on the merged commit. The published package is the qualified artifact from that run.
Collector 0.7.43: no false red after a restart, pairing indexes built by a plain update, dispatch tagging, Codex usage filed as Codex
Collector 0.7.43 replaces 0.7.42: no false red after a restart, pairing indexes built by a plain update, dispatch tagging, and Codex usage filed as Codex.
- No false red after a restart (
.163.101). After the collector restarts, an idle Claude source reads amber while its session counts reconcile, then green. 0.7.42 could read red once during that window, about 10 minutes after the restart. A real capture gap still reads red: session files with token activity that never reach the ledger. - A plain update builds the Codex pairing indexes (
.163.100).plimsoll lifecycle updatenow runs the pairing-index step itself, after readiness and before its completion receipt. The step has one 180-second window and retries once if another process holds the ledger. If the ledger is busy, the update still completes and its receipt records the skip with a reason; a later update builds the indexes. The standalonelifecycle pairing-indexes --applystep is unchanged. - Dispatch tagging (
.165.3).plimsoll dispatch bind,closeandrestamptag a dispatched lane's session with its work item, project, attempt and role. Every usage row of that session in the window carries them, including the kept Codex response log row. The paired response span carries none. - Codex usage is filed as Codex (
.163.98).- Codex service names, including
Codex_Desktopandcodex-app-server, are now read as Codex. - A Codex service that sends with a Claude credential gets
401 source_mismatchand must use the Codex source and producer token. A named service the collector does not recognise is stored as unknown, never as Claude Code. - A response span keeps its conversation id when the span carries it, or when exactly one session-bearing event shares its trace in the same export.
- Codex service names, including
- A LaunchAgent that loads without starting is kickstarted once (
.163.91). Coverscapture-roots add,load-launch-agentandupdate. If one kickstart does not bring/statusup, the command fails with a receipt; it never loops. capture-roots discoverlists live-covered homes separately (.163.93). A home that already reports live (Codex's OTLP exporter, or Claude hooks or OTLP) is listed with the evidence by key name only.capture-roots addon such a home warns that the file path will sit beside the live path.
Updating and rolling back:
- Rollback to 0.7.42 is supported. 0.7.42 runs on a ledger and config that 0.7.43 has used: snapshot list and prune, status, the daemon with hook and OTLP intake, export and session sync all passed on a home 0.7.43 had updated (32 of 32 checks). 0.7.42 does not recognise the new pairing-index result in an update receipt, so it keeps both snapshots after such an update; nothing is lost.
- Nothing is received while the collector is stopped. Hook posts and OTLP exports during the stop are not recorded; token usage in session files is captured again after the restart. The first 0.7.43 update on a ledger without pairing indexes also builds them, within its bounded window.
- Check for
source_mismatchafter each update. A Codex producer that has been sending with a Claude credential is refused from 0.7.43 on. Grep the collector log after each host's update, and correct that producer's source header and token.
Runtime CLI SHA256: 19b0fd836967c14bad96c2bb1bb42cdcd2ce2de90f38871ab9284b5257690828.
Qualification: typecheck, the proof suite and system end-to-end qualification ran on the merged commit. The published package is the qualified artifact from that run.
Collector 0.7.42: Codex responses counted once at the source, busy-session summaries keep progress, bounded maintenance timing
Collector 0.7.42 counts each Codex response once at the source, keeps busy-session summaries moving, and bounds maintenance timing.
- Codex responses count once (
.163.95). codex-app-server reports each response twice: acodex.sse_eventlog and ahandle_responsesspan. 0.7.42 pairs the two in the local ledger before usage upload, so a response counts once. The span stays in the ledger as raw evidence and no longer counts as usage.- Pairing needs three indexes, built in a stopped-service step:
plimsoll lifecycle pairing-indexes --apply, run while the collector is stopped. Without--applythe command only reports whether the indexes are ready. - A new ledger gets the indexes when it is created. On an existing ledger the collector runs unpaired, as 0.7.41 did, until the step has run.
- On a copy of the largest fleet ledger (73 GB), the step took about 24 seconds.
- After a rollback to 0.7.41 and a later update, run the step again to refresh its historical cursor and target.
- Pairing needs three indexes, built in a stopped-service step:
- Busy-session summaries keep their progress (
.163.87). A summary rebuild keeps its progress while unread rows in the session change. A terminal receipt retarget advances the old session revision, so a 0.7.41 worker rejects stale data after a downgrade. - Possible capture losses stay visible (
.163.81). When a JSON discriminator probe saturates, the Codex and Claude tailers keep the skipped usage visible as a possible loss. A revisit queue resumes partial files without losing the capture claim gap. - Coverage walks page their saved cursors (
.163.82). A walk records a gap when a known partial file vanishes before its directory entry is reached. A file created and removed entirely between walks remains outside this claim. - Observe-only budget sampling (
.164.3).- The collector samples ledger size, process memory and CPU, outbox age and summary lag.
plimsoll statusand the CSV export show advisory targets and local history. No capture budget is enforced.- Purging a stopped collector also removes the ledger's WAL and SHM files.
- Bounded HTTP deadlines and maintenance timing (
.163.83). Proofs now cover file-backed status probes, the maintenance deadline-to-reap and absolute hook latency. The runtime keeps its bounded HTTP request behavior.
Updating and rolling back:
- Rollback to 0.7.41 is supported. 0.7.41 runs on a ledger that 0.7.42 has used, including 0.7.42's replaced dashboard and session-summary triggers. An explicit rollback keeps the live ledger.
- Nothing is received while the collector is stopped. During an update, hook posts and OTLP exports are not recorded. Token usage in session files is captured again after the restart. In managed update windows the stop has lasted 1 to 7 seconds. The first 0.7.42 update also includes the one-time pairing index build, about 24 seconds on the largest ledger.
- Keep-all retention works as before. Updates and rollbacks with
--retention keep-allbehave as in earlier releases.
Runtime CLI SHA256: 9c80da7cc66570deef7706c75efbda0f7dbb27c09bdec5c89ebb2e5fcfb83229.
Qualification: typecheck, the proof suite and system end-to-end qualification ran on the merged commit. The published package is the qualified artifact from that run.