Skip to content

Releases: plimsollmark/plimsoll

plimsoll v0.21.0

Choose a tag to compare

@carrollco carrollco released this 09 Oct 21:34

Every answer from the daemon is now tied to the request it answers, so a proxy that serves one request's answer for another can no longer make a call that ran read "nothing ran". The daemon also marks the refusals made before its own code sees a request, so they read "nothing ran" instead of "may have run", and closes a session whose open answer never reached its client at its first idle timeout instead of keeping it for its lifetime. The Go and TypeScript clients now mark a call that never left as not dispatched, as the Python client already did. Two grant-matching gaps from a security scan are closed, and the TypeScript client is published to npm by the release workflow with a provenance statement. This is protocol 3: read Upgrading before you deploy it.

Upgrading

  • Protocol 3: upgrade clients and daemons together. A v0.21.0 daemon refuses a request from an older client before reading it (unimplemented, marked not dispatched, reason protocol), and an older daemon refuses a v0.21.0 client the same way. Nothing runs in either case.
  • Pass two response headers through any proxy. Every answer carries Plimsoll-Request-Id (the client's ID, echoed) and some carry Plimsoll-Not-Dispatched. A proxy that drops response headers it does not know makes every answer read as not bound to its request: every successful call becomes a data-loss error, although the code ran.
  • A grant profile whose * widens a capped route's fixed segment no longer loads. With route_max_calls capping GET /jobs/ep1, an entry GET /jobs/* let ep1.json, ep1. or 0123 reach an upstream that serves them as the capped route, uncounted. Load fails and names both entries: move the cap to the wider route, or list the routes it should reach instead of the *.
  • A * now binds only RFC 3986's unreserved characters: letters, digits and -._~. It no longer binds $ ($batch runs other requests on OData), +, @ or =, which an upstream may read as syntax. A route segment that needs one is granted as a literal.

Answers bound to their request

Found by our failure-injection review, which put a proxy between each client and the daemon and broke the connection at every stage.

  • The request ID. Every official client (Go, Python, TypeScript) sends a fresh random Plimsoll-Request-Id, 32 lowercase hex digits, with every request. The daemon copies it onto every answer: its own refusals, the RPC library's, and an HTTP 404 included.
  • An answer without the ID, or with another one, is believed in nothing. A success becomes a data-loss error, since the code may have run (Go client.ErrAnswerNotBound, Python AnswerNotBoundError, TypeScript data_loss with answerNotBound). An error keeps its code but loses its "nothing ran" mark, its session end and its unanswered-call record. Before, a refusal captured from one request and served for a later one read "nothing ran" in all three clients while the daemon ran the later request, and placement or the code tools could then run it again. A replayed success for an identical request also passed the record check, which binds a result to the request's content, not to the request.
  • Refusals before the handler are marked. A body the RPC library cannot parse, one over its 8 MiB cap, an unsupported content type, or a procedure the daemon does not serve never reaches plimsoll's code, so nothing can have run. They carried no mark and read "may have run". The daemon now adds Plimsoll-Not-Dispatched (unsupported for a 404, request otherwise), and the clients read it when the error has no detail of its own.

A session nobody claimed is closed

Also from the failure-injection review. If the answer to an OpenSession never reached its client (a dropped connection, an answer refused for not carrying its request's ID back, a client that crashed after sending), the session lived on for its whole lifetime, 30 minutes by default: nobody could use or close it, it held a place under SANDBOX_MAX_SESSIONS and the per-caller cap, and on E2B the open's reserved paid time. The daemon now closes a session no request has named by its first idle timeout (5 minutes by default; 5 minutes too when SANDBOX_SESSION_IDLE=0), with the new session end unclaimed, instead of suspending it. A call naming the session claims it, even one refused before it ran. The daemon cannot tell a lost answer from a client that opened a session ahead of calls that did not come, so it closes both; nothing ran in either.

What it costs: a client that opens a session and makes its first call after the idle timeout gets that call refused, not dispatched, with the end unclaimed, and opens a new one; nothing ran in the closed session, so nothing is lost. CodeSandboxes does that by itself, so a warm that its conversation does not use within the idle timeout costs one extra open.

Clients

The first four are also from the failure-injection review.

  • Go and TypeScript mark a call that never left. A failure before any byte of the request was sent is marked not dispatched, reason environment, as Python's has been since v0.20.0: Go's placement then tries another daemon, and a session stays usable in all three clients. Go marks a refused or unreachable connection, a name that did not resolve, and a TLS handshake that failed (a cleartext port, an untrusted certificate). TypeScript marks only the three causes Node names with a code that was reproduced: a refused connection (ECONNREFUSED), TLS against a cleartext port (ERR_SSL_WRONG_VERSION_NUMBER) and a self-signed certificate (DEPTH_ZERO_SELF_SIGNED_CERT). Any other failure stays unmarked, since it does not prove the request stayed here.
  • Python's request timeout covers the name lookup. The daemon's host name was looked up outside both the request timeout and CancelHandle, so a resolver that did not answer held the call for as long as the resolver took. The lookup now runs in a thread of its own: at the deadline the call raises RequestTimeoutError, and a cancel returns at once, both marked not dispatched. Python cannot interrupt a lookup, so its thread runs on until the resolver answers; while 32 such lookups are still running, a new one is refused at once, marked environment. An IP address in the URL is not looked up.
  • Python refuses a base URL whose host it could never look up. A host with an empty label (plimsoll..example) or a label over 63 characters is now InvalidBaseURLError when the Client is built. Before, every call to it raised a UnicodeError, which is not a PlimsollError.
  • Python returns a chunked answer that arrived whole. The client counted an answer as complete only when it matched its Content-Length, and the daemon sends any answer over 2048 bytes in chunks without one. So if the deadline or a cancel came in the instant after such an answer arrived, the client discarded it as timed out or cancelled, and a session stopped. An answer read through its final chunk now counts as complete.
  • Go and Python say when the client stopped a session. After a session call whose answer the client could not check, the client sends nothing more on that session. Go's Session.Stopped() and Python's Session.stopped (and AsyncSession.stopped) now say why, as TypeScript's Session.stopped already did. Before, a caller found out only by sending a call and reading the refusal. Ended() and ended still report only an end the daemon stated. In Go, Ended() no longer waits for a call in flight, and Stopped() does not either.
  • @plimsollmark/client is published by the public repository's release workflow through npm's trusted publishing: no npm token exists, and npm records which repository, commit and workflow run built each version.

Contributors

  • A test server must be built by rpc.NewHandler, which wraps it in rpc.BindAnswers: the clients now refuse an answer that does not echo their request ID, so a test server built another way turns every success into a data-loss error. Every test server in the repository was moved to it.
  • make e2b-session-live also runs TestLiveE2BDaemonClosesAnUnclaimedSession, which goes through the daemon and the Go client against the live E2B service. The provider's conformance suite never reached the daemon, where the unclaimed close lives. It is paid like the rest of the target: one microVM for about its 10 s idle timeout.
  • @plimsollmark/client is published by publish-npm.yml when a GitHub release is published, with the same tag and source checks as the Python workflow; releasing.md has npm's side of the setup.

plimsoll v0.20.0

Choose a tag to compare

@carrollco carrollco released this 09 Oct 05:13

plimsoll now plugs into six agent frameworks: Agno, CrewAI and Google ADK in the Python client, and the Vercel AI SDK, Trigger.dev and Mastra in the TypeScript client. Each hands the model's code to a plimsoll daemon you run, behind the isolation tier you require (kernel by default), and answers with the run record the client checked. E2B now keeps sessions, with grants inside them, and a session on a provider billed by the second draws on the daily paid allowances. This release also carries the fixes from five security reviews and moves to Go 1.26.9. Read Upgrading before you deploy it: six changes need action.

Upgrading

  • The egress guard has a listener of its own. With E2B_GUARD_URL or SANDBOX_DOCKERCLOUD_GUARD_URL set, PLIMSOLL_GUARD_ADDR is required: the guard is served there and nowhere else, and the daemon refuses to start without it. Point the public guard URL at that address. Code in a granted sandbox may reach the guard, and on the shared listener it could also reach the RPC procedures and the health endpoints.
  • E2B_GUARD_URL must name a host, not an IP address: E2B's network rule takes a domain name and refuses an address at create.
  • Images that carry /runner.mjs must carry this release's. The runner now reports a step's output as base64 of its raw bytes, so output that is not UTF-8 survives exactly. A daemon of this release refuses a report from an older runner as a protocol error. make docker-images rebuilds the shipped images; rebuild an OpenShell image of your own the same way.
  • The session owner digest changed. Clients now key it with a value derived from the caller token, not the token itself (with a token over 64 bytes, the old digest could be tested against the clients file's stored hash). Upgrade every client of a daemon together: during a mixed rollout, one user counts as two owners.
  • A docker session that mixes cells and project calls loses its cell state at each project call. Nothing of a session's own code may run while plimsoll starts the project runner, whose input carries the key its report is authenticated with, so the interpreters are killed first. The next cell reports interpreter_started; the work directory's files are untouched. A session that runs only cells, or only project calls, is unaffected.
  • Go 1.26.9 or newer. Go 1.26.6 has eleven standard-library advisories published on 2026-10-08, among them an HTTP/2 server crash and HTTP/2 memory exhaustion in code the daemon calls. golang.org/x/net moves to v0.60.0.

Agent frameworks

  • Python: Agno, CrewAI and Google ADK. plimsoll_client.agno.PlimsollTools, plimsoll_client.crewai.PlimsollCodeTool and plimsoll_client.adk.PlimsollCodeExecutor run each call in a fresh sandbox through one shared helper, plimsoll_client.execution.CodeExecutor (Python or JavaScript, the call's files written first, a kernel floor and a 30 second budget by default). Install with the extras: pip install 'plimsoll-client[agno]', [crewai] or [adk]. Guides: the Python client README and docs/google-adk.md.
  • TypeScript: the Vercel AI SDK. @plimsollmark/client/ai-sdk gives streamText an executeCode tool, fresh by default or, with forConversation({ userId, conversationId }), one sandbox per conversation whose variables and files survive between calls on a daemon that keeps sessions. docs/ai-sdk.md.
  • Trigger.dev and Mastra add-ons now require the kernel tier unless you lower it (they accepted any tier before), cap the sandboxes one user holds (maxSessionsPerOwner, default 3, closing that user's least recently used idle one), and take the user from your own login: warm(runId, userId) on Trigger.dev, an owner option on Mastra for agent networks.
  • Whether the code ran is never left to a guess. Every tool says, before anything else, "Nothing ran (reason)" when plimsoll refused the call before it started, and "The code may have run" when an answer was lost after it might have. Nothing is ever retried automatically. A run with no step report is an error that says it may have run, not an invented exit code.
  • Cancelling a call cancels its request. In every Python tool, AsyncClient, and ADK through its guard: a request not yet sent is never sent, and one in flight has its connection cut. On docker, a cancelled run's container was gone a quarter of a second after the cancel.
  • Framework traps closed. CrewAI's tool cache, when a crew turns it on, can no longer answer for plimsoll's tool. ADK's guard (PlimsollCodeExecutorGuard, registered first on the Runner) refuses the setting that swaps plimsoll for Gemini's built-in executor, refuses a Runner with no artifact service before anything runs, and gives the model the code-block instruction ADK no longer adds; ADK's CSV preprocessing, which assumes state a fresh sandbox does not keep, is refused.

Sessions

  • E2B keeps sessions. One microVM for many calls, files and interpreter state kept between calls; a suspend is E2B's pause, so files survive it and interpreters do not. The session conformance suite passed against the live service. docs/sessions.md#e2b.
  • Grants inside E2B sessions, off unless E2B_SESSION_GRANTS chooses one of two designs: session (one guard credential for the session's life, serving only the grant of the call in progress) or call (a network rule with a fresh credential put on for each granted call and taken off after). Both passed live, before and after a suspend.
  • Paid session time is charged. On E2B and Docker Cloud a session draws on SANDBOX_PAID_SECONDS_PER_DAY and each caller's paid_seconds_per_day for the time its sandbox runs, reserving ahead what a call could cost.
  • A per-user cap in the daemon. SANDBOX_MAX_SESSIONS_PER_OWNER caps the open sessions of one user of one caller. Clients name the user in OpenSession as a digest under a key derived from the caller's token, never the user's ID. At the cap, the open closes that user's least recently used idle session, whose later calls are refused as replaced, not dispatched. An open the daemon refuses closes nothing. Describe states the per-caller and per-owner caps.
  • A session the daemon no longer knows (after a restart, say) is an end, not an error a client retries.

Grants

  • route_max_calls caps one allowed route, for a route whose call spends, such as a GPU job's submit: past the cap the broker answers 429 and the call never reaches the API. A call counts against every capped entry it matches, compared without regard to letter case, so neither a wider wildcard nor the order of allow lets calls around a cap; in a session the count spans the session. docs/capability-grants.md.

Security

Fixes from the five security reviews of this release. None was critical.

  • A project call in a docker session starts with nothing of the session's own code running. Before it, every process of the session's is killed, the interpreters it keeps included. Until a program of plimsoll's has loaded the library that makes it unreadable, a process of the same user could open its input, output and memory, and a project call's input carries the key its report is authenticated with. So a session that runs cells loses their interpreter state at its next project call, and the next cell reports interpreter_started; its files are untouched, and a session that runs only project calls keeps no process of its own and pays nothing. Pausing the interpreters instead would not hold: a paused process can arrange its own resume before it is paused. A snippet and a cell carry no key and are not protected this way; a grader whose verdict the graded code must not be able to forge uses a project call. docs/sessions.md#docker.
  • A request refused before its body is read (bad credential, missing scope, decode capacity, an unknown path, the guard's refusals) answers at once and closes the connection, instead of first waiting for a body the client holds back. A declared body over the cap is refused from its header, OPTIONS * reaches the daemon's own checks, and each listener holds at most its share of half the process's descriptor limit in open connections.
  • Every local setting, and every listener's bind, is checked before the startup smoke test, which on E2B and Docker Cloud creates a billed microVM; a listener that fails drains in-flight runs instead of exiting on the spot.
  • A sandbox whose delete gave up stays charged to the paid allowances until the provider's own lifetime for it ends, including across UTC midnight, for runs, sessions and failed session opens.
  • A sandbox ID from a provider's create answer is used, to run code or to delete, only once it is proven to be the sandbox that create made.
  • A file planted in a session's work directory can no longer hold a later call until its deadline.
  • An idle timer that fired late no longer suspends a session used since.
  • attest refuses a bundle line without its signed link as it reads it.
  • The export to this repository stages committed content only, and a paid live test needs its make target's flag, not just a key in the environment.
  • Development dependencies: the ws the Trigger.dev add-on's test environment resolved (through socket.io-client, which pins its engine.io-client exactly) moves to 8.21.3, and the example's tar 6 branch is gone, both through scoped overrides; the Python build lock moves to setuptools 84.0.0, which builds the same wheel and sdist contents, checked file by file. Th...
Read more

plimsoll v0.19.0

Choose a tag to compare

@carrollco carrollco released this 04 Oct 22:46

The Docker Cloud provider can now also speak the REST API Docker documents, kept as a backup behind one setting while the pre-launch Connect API, the only one with grants and image evidence, stays the default. The TypeScript code tool returns the digest of the run record it checked, Trigger.dev deployments have a guide and a starter repository, and the docker test suite refuses images built from other inputs. The clients are at 0.19.0.

Docker Cloud

  • The Docker Cloud provider can also speak the REST API Docker documents, as a backup. SANDBOX_DOCKERCLOUD_API chooses the wire API once, at startup: connect (the default) is Docker's pre-launch sandboxes API that earlier releases spoke, which Docker no longer documents; rest is the API Docker documents. Both passed the live suite on 2026-10-04, and the provider never switches between them at run time. Connect stays the default because it is the only one with what follows; REST is kept as the backup until it covers them, which matters if Docker withdraws the undocumented Connect endpoint. On rest:
    • No image identity is stated, because the REST API reports no booted image digest; a caller's software rule therefore fails closed, as on E2B. PLIMSOLL_HARDENED=1 refuses rest, and wants SANDBOX_DOCKERCLOUD_API=connect named rather than left to the default.
    • Grants are refused, and startup fails if SANDBOX_DOCKERCLOUD_GUARD_URL is set: after the guard's network rule is applied, the REST API declines to report the sandbox's policy, so plimsoll could not verify the rule took effect. A live test watches for that to change.
    • A run timeout above 270 seconds is refused: an exec on rest ends when its 300-second endpoint credential does.
    • SANDBOX_DOCKERCLOUD_API_URL defaults to the documented https://connect.docker.com/sandboxes; on Connect it stays required.
      Nothing changes for an existing Connect deployment. A value other than connect or rest fails startup. docs/dockercloud.md.

Clients

  • The TypeScript code tool reports the run record it checked. Every executeCode result, and every CodeSandboxes.run result, carries recordSha256, the SHA-256 of the run record the client recomputed and checked before returning. It identifies the checked record; it is not a signature. clients/typescript/README.md.
  • A guide for the Trigger.dev code tool. docs/trigger-dev.md covers what the agent receives (persistent cells, the isolation floor, the checked record), how the add-on follows Trigger.dev's code-sandbox recipe, and deploying a task that reaches a daemon.
  • @plimsollmark/client 0.19.0 on npm; the Python client is at 0.19.0 in the repository.

Docs

  • The README now says what the trial runner, run records and plimsoll-attest add up to: an exam on code an agent wrote, where the exam (the simulated system, the scenarios, the pass mark) stays with whoever runs it.
  • docs/dockercloud.md: the REST API shipped at Docker's launch on 2026-09-24, not after it; "Choosing the API" compares the two.

Contributors

  • make audit and make test run with -short, which the docker test helpers read as skip, so only make docker-suite starts containers. The plain gate takes under two minutes.
  • make docker-suite runs the docker tests that can share a host in parallel (DOCKER_PARALLEL, default 4), with the heaviest seven limited to three at a time; about 155 seconds, was about 356.
  • make docker-images stamps each image with io.plimsoll.inputs, a SHA-256 of the build context docker/ (minus the paths in docker/.dockerignore), and make docker-suite stops with "run make docker-images" when a built image's label differs, so a change under docker/ cannot pass on images built before it.
  • The Makefile includes an optional local.mk whose EXTRA_AUDIT steps make audit runs, for checks a working copy keeps to itself.

plimsoll v0.18.0

Choose a tag to compare

@carrollco carrollco released this 04 Oct 12:22

Docker guests now run as a uid no host account uses, E2B and Docker Cloud runs can draw on a daily allowance of microVM seconds, a signed bundle of run records now proves it is the whole file its harness wrote, and the efficiency advisor hands a caller only a batch route the operator has declared. Docker and OpenShell images must be rebuilt, and a bundle written by an earlier release cannot be verified or continued by this one.

Breaking changes

  • Docker guests run as uid 61000 by default, and the daemon refuses a uid the host already has. Before, every guest process ran as uid 1000. Under runc (Docker's default runtime) without user-namespace remapping, a container's uid is the host's, and 1000 is the first account most Linux machines create, often the operator's, so a container escape acted as that account. SANDBOX_GUEST_UID now sets the guest's uid and gid (default 61000), and the session identity check runs as that uid plus one (61001 by default; it was 2000). 61000 is below 65,536, so it still starts under rootless Docker and userns-remap, which map 65,536 uids, and it lies in 60706 to 61183, a range systemd leaves unallocated. Preflight (the check the daemon runs at startup and on every /readyz) refuses a guest uid, or that uid plus one, that the host's /etc/passwd or /etc/group contains; accounts from a directory service such as LDAP are not in those files, so choose the uid with them in mind. The setting accepts 1 to 65532 (the uid and the uid plus one stay below nobody, 65534) outside the bands systemd gives to accounts it never writes to /etc/passwd, and setting it for any provider other than docker fails startup. The startup smoke test checks in /proc/self/status that every uid, gid and supplementary group of a guest process is the guest uid. Inside the guest the uid has no account, so HOME is / unless the image adds a passwd line for it, a file only node can read is unreadable to a run, and os.userInfo() (Python's getpass.getuser()) fails: give package files to everyone (chmod -R a+rX), or add a passwd line if a tool needs a user name. docs/guest-dependencies.md, SECURITY.md.
  • Docker and OpenShell images must be rebuilt. The project runner (/runner.mjs, the program inside the image that writes a project's files and runs its steps) now receives the image's environment from plimsoll's plan and hands it to the steps, and its report carries version 3 of its marker. An image with an older runner would drop that environment, so it fails the startup smoke test and the daemon does not serve. The runner also kills a step that outlives its budget with SIGKILL (see Docker). Run make docker-images, and rebuild an OpenShell image on the new base image; an OpenShell image must now also carry /usr/bin/env, which plimsoll's own programs start under.
  • Docker ignores an image's ENTRYPOINT, and the image's environment reaches only guest processes. Every container command is now plimsoll's own. plimsoll's programs in the container (the runner, the smoke probes, a session's main process, the sweep, the process lister, the identity check, the relays and the interpreter launcher) start under /usr/bin/env -i with a fixed PATH; a snippet's node, a project's steps and a cell's interpreter get the environment of the verified image's configuration, handed on explicitly. Before, every process started with the image's ENV, and a list of refused variables was all that kept a variable such as NODE_OPTIONS from loading guest-written code into plimsoll's sweep. That list is gone: variables it refused, such as NODE_OPTIONS, NODE_DEBUG, or a PYTHONPATH entry inside a writable directory, are accepted now and reach guest processes only, where they can load only the guest's own code. Preflight still refuses an image that sets a variable the C library or dynamic loader reads before any program runs: any LD_ name, GLIBC_TUNABLES, LOCPATH, GCONV_PATH, and now NLSPATH, which v0.17.6 accepted. An entry that is not NAME=value is refused too.
  • A docker image tag re-pointed after startup is refused until the daemon restarts. The startup smoke test records the image IDs it proved, and a run launches only those. A tag moved to new content later (for example by make docker-images on the same host) fails Preflight, so /readyz, and runs, session opens and pool refills on it are refused before any code runs (reason environment) until a restart proves the new image. Before, the new image was launched within seconds with its lockdown, runner guard and languages unproven. After rebuilding images on a host that serves, restart plimsolld.
  • A bundle written by v0.17.6 or earlier cannot be verified or continued. A bundle (the JSON Lines file of signed run records that plimsoll-attest and the attest package keep) now carries a signed link on every line and must end with a checkpoint (see Run records and attestation). A file without them fails verify, and run refuses to append to it: start a new bundle, and keep the release that wrote an old one to verify it. Version 1 run records are refused, since no bundle with links can hold one. A Go program that uses attest.Harness directly must call Harness.Checkpoint when it finishes writing, or its bundle fails verification as incomplete.
  • Advice reaches a caller only for a batch route the operator declared. With advice: caller, v0.17.6 returned a finding to the calling agent whenever the profile granted the per-item route's collection (GET /items for GET /items/*), on the strength of the path alone. Now the profile must declare the relation in batch_of ({"GET /items": ["GET /items/*"]}, which plimsoll-specgen writes from an x-plimsoll-batch-of marker in the OpenAPI document) and grant the batch route; a path guess stays on the operator's side as a candidate. A profile that relied on the guess returns no caller advice until it declares batch_of. The audit line's grant_route now means a declared batch route the profile does not grant. docs/efficiency-advisor.md.
  • Hardened mode with E2B or Docker Cloud requires a daily allowance on every caller. With PLIMSOLL_HARDENED=1 and a provider billed by the second, the daemon refuses to start, before its smoke test creates a billed microVM, unless every caller in PLIMSOLL_CLIENTS_FILE has a paid_seconds_per_day (see Billing and spend). Outside hardened mode no allowance is set unless you set one. docs/hardened-mode.md.
  • Go embedders that implement a session provider or their own admitter have API changes. sandbox.Session gains Environments(), the environment the opened session actually runs, which the daemon checks the caller's software rule against. sandbox.RefuseCell takes the call's context. A session's Done now closes only once its sandbox is deleted, and its Err is set as it ends. An admitter other than the daemon's limiter or sandbox.WithAdmission should wrap its release with sandbox.WithCapacity, so a provider whose sandbox outlives the call (OpenShell) can hold the slot through the delete with sandbox.HoldCapacity. sandbox.SessionEnd and sandbox.SessionEndedError keep their names and behaviour.

Billing and spend

  • E2B and Docker Cloud runs can draw on a daily allowance in seconds. Each run on these providers creates a microVM (a small virtual machine) that the operator pays for by the second, and a rate limit bounds runs per minute, not seconds per day. A caller's paid_seconds_per_day in PLIMSOLL_CLIENTS_FILE and the daemon-wide SANDBOX_PAID_SECONDS_PER_DAY are allowances of microVM wall time per UTC day. Before admission, a run reserves its timeout (the provider's ceiling when the call sets none) plus the provider's teardown bound, 33 seconds on both (three delete attempts of up to 10 seconds each, with pauses of 1 and 2 seconds between them); a run that would go over either allowance is refused before any code runs (reason capacity, ResourceExhausted over RPC), and the refusal for the daemon-wide allowance states no numbers, which would reveal other callers' spend. At the end the run is charged the wall time of the provider call and the rest is given back. A run that crosses midnight UTC counts in the new day for the part it ran after midnight. Counts are per daemon and in memory: a restart forgets the day's spend, and since a spent allowance is a capacity refusal that placement retries elsewhere, a caller with the same allowance on K daemons can spend K times it. plimsoll-clients gains a limit command and -paid-seconds-per-day on create and import, and list shows allowances. Refusals are counted in plimsoll_shed_total{reason="paid_caller"} and {reason="paid_daemon"}, and the audit line states paid_seconds. Setting SANDBOX_PAID_SECONDS_PER_DAY for a provider that does not bill fails startup. docs/callers.md, docs/placement.md.
  • A run whose microVM may still exist with nothing able to delete it is charged its whole reservation. Such a microVM bills until the lifetime plimsoll requested at create runs out (the run's deadline plus 10 seconds on E2B, plus 30 on Docker Cloud), which the reservation covers, so charging only the run's wall time would undercount. This applies when the provider's delete gave up after its retries, and when a create's outcome is unknown. On E2B a create is unknown unless its error proves the request never left (a DNS or...
Read more

plimsoll v0.17.6

Choose a tag to compare

@carrollco carrollco released this 03 Oct 10:21

A session pool's memory reservation is given back exactly once, whatever the order of a failed pool start and Drain, and the docs state the small-pool warm guarantee for the sessions it covers.

Sessions

  • The memory budget stays whole when Drain races a pool start. WithAdmission reserves a session pool's memory before the provider starts the pool. If Drain finished while that start was still inside the provider, and the provider then refused the start, v0.17.4 and v0.17.5 gave the reservation back twice, so the budget admitted a pool's worth of runs more than it holds. Each reservation is now given back once, by whichever of the failed start and Drain comes first. Only an embedder that kept serving runs after Drain could be affected.
  • The small-pool guarantee is stated for the sessions it covers. A pool of 4 or fewer warms every language in every waiting container, so a session handed one finds its languages warm. Opens that find no container waiting, as a burst larger than the pool can, create their containers as they would without a pool. The v0.17.5 notes said every session. docs/sessions.md, "A warm pool".

Clients

  • @plimsollmark/client and the Python client are at 0.17.6, unchanged apart from the version.

plimsoll v0.17.5

Choose a tag to compare

@carrollco carrollco released this 03 Oct 03:03

A warm session pool of 4 or fewer containers warms every language the image runs in every container, whatever sessions hint.

Sessions

  • Small pools do not split. A pool of 4 or fewer containers (SANDBOX_SESSION_POOL) now warms every language the image runs in every waiting container, so every session handed a waiting container finds its languages warm whatever the mix of hints; opens that find none waiting, as a burst larger than the pool can, create their containers as they would without a pool. Only a pool of 5 or more divides its containers across language sets by demand. In simulation, a split pool of 1 to 3 found a language fewer than a quarter of sessions asked for, or a burst of opens, warm 75 to 91% of the time, to save at most a few idle interpreters. docs/sessions.md, "A warm pool".

Upgrading

  • A pool of 4 or fewer holds one idle interpreter more per container than a split one would when sessions ask for a single language (18 to 21 MiB under runc). Each waiting container is still charged one run against SANDBOX_TOTAL_MEMORY_MB, as in v0.17.4.

Clients

  • @plimsollmark/client and the Python client are at 0.17.5, unchanged apart from the version.

plimsoll v0.17.4

Choose a tag to compare

@carrollco carrollco released this 03 Oct 02:10

Fixes from reviews of v0.17.3: the warm pool is charged against the memory budget, a refused open no longer moves the pool's language mix, the docker provider's Drain waits for every session it found open, and the pool never rebalances away a language sessions still ask for.

Sessions

  • The warm pool counts against SANDBOX_TOTAL_MEMORY_MB. A waiting container holds memory (its limit is one run's) but no concurrency slot, and neither plimsolld's clamp nor WithAdmission charged it, so with a pool the budget no longer bounded what sandboxes hold. Each waiting container is now charged one run: plimsolld clamps SANDBOX_MAX_CONCURRENT to the budget's runs less the pool's size and refuses to start when no run would fit, and WithAdmission charges the pool when it starts and gives the charge back after Drain.
  • An open the docker provider refuses (a floor it cannot meet, say) no longer counts toward the pool's language mix; an open counts once it passes the floor. Before, a caller repeating refused opens could move the pool toward its languages. plimsolld checked the floor earlier, so this reached only embedders.
  • The docker provider's Drain reports success only once every session it found open has been removed. A session whose lifetime ran out, or whose caller closed it, at the moment Drain ran could be skipped, so Drain returned with its container still there, and could misuse a counter in a way Go reports as a panic.
  • The warm pool's rebalancing never removes a container warming a language set that at least a quarter of recent sessions ask for, since a claim will take it. In v0.17.3 a pool of 2 whose demand sat near that threshold replaced such a container about once a minute.
  • docs/sessions.md, "A warm pool", now says what a split pool does not keep warm: a language fewer than a quarter of recent sessions ask for, and concurrent opens that empty the pool, can find their interpreter cold in an already running container.

Upgrading

  • With SANDBOX_TOTAL_MEMORY_MB and SANDBOX_SESSION_POOL both set, plimsolld admits that many fewer concurrent runs, and refuses to start if the pool leaves no room for one.

Clients

  • @plimsollmark/client and the Python client are at 0.17.4, unchanged apart from the version.

plimsoll v0.17.3

Choose a tag to compare

@carrollco carrollco released this 03 Oct 01:09

The warm pool's rebalancing is bounded at one container start a minute whatever callers ask for, a small pool no longer splits its containers across languages it cannot each serve, and the docker provider's shutdown ends open sessions even when an open is still in flight.

Sessions

  • Rebalancing is a slow janitor. Claims and refills already follow the pool's split: a claim takes the waiting container closest to its hint, and the refill warms the language set furthest below its share. Rebalancing now only replaces containers of sets nobody asks for any more, which no claim would take, at most one a minute. In v0.17.2, sessions that ran six or more of one language before switching moved a container every few sessions without finding their language warm any more often. docs/sessions.md, "A warm pool".
  • A small pool does not split. A pool with fewer containers than the language sets holding at least a quarter of recent demand warms all their languages in every container. In v0.17.2 a pool of 1 serving sessions that alternated JavaScript and Python always warmed the language the previous session asked for, so no session found its own warm; a pool of 2 with three sets in turn found two in three.
  • In simulation, at one session a second, sessions alternating or rotating languages found them warm every time at pool sizes 1 to 32, and random hints found them warm 97 to 100% of the time for under 0.02 extra container starts per session.
  • The docker provider's Drain ends the sessions already open even when its context ends while another open is still in flight; before, it returned and left them to be reaped once their lifetime passed.

CI

  • The gate checks formatting: golangci-lint runs gofmt.

Clients

  • @plimsollmark/client and the Python client are at 0.17.3, unchanged apart from the version.

plimsoll v0.17.2

Choose a tag to compare

@carrollco carrollco released this 03 Oct 00:14

Fixes from a review of v0.17.1: the warm pool's split now stays put under demand that takes turns at any pool size, a failing session call is proven to come back as a result before sessions are served, and the docker provider's shutdown no longer races its pool.

Sessions

  • The warm pool's split holds at every size. v0.17.1's fix assumed small pools and two language sets: in larger pools one open still moved a set's share by up to size/16 of a container, so sessions alternating between JavaScript and Python moved a container on every open in odd pools of 9 or more, as did three sets in turn from 7 up. One open now moves a set's share by at most a quarter of a container (its weight is 1/(4 x pool size) in a pool of more than 4), and a container moves only when the move is worth more than 1.75 containers. In simulation, demand that takes turns or runs a few sessions of each moved nothing once settled at any size from 1 to 32, and random hints cost under 0.02 extra container starts per session; a lasting change still moves the pool within a few turnovers. docs/sessions.md, "A warm pool".
  • The session startup check now runs a call whose code exits 3 and serves no sessions unless it comes back as a result with exit 3, on any provider.
  • A pool whose image no longer runs a language the pool was asked for makes what it can and settles, instead of remaking the same container without pause. Embedders only.
  • The docker provider's Drain no longer waits on container removals the pool may still add to when its context ended first (a data race), and a pool started after Drain is refused. Embedders only.

Images

  • Preflight refuses an image whose ENV sets NODE_DEBUG or NODE_DEBUG_NATIVE. NODE_DEBUG=esm makes node write to stderr before a call's start marker, so since v0.17.1 every failing session snippet on such an image came back as an error instead of its exit code.

Clients

  • @plimsollmark/client and the Python client are at 0.17.2, unchanged apart from the version.

Upgrading

  • An image whose ENV sets NODE_DEBUG or NODE_DEBUG_NATIVE now fails Preflight.
  • StartSessionPool after Drain returns an error.

plimsoll v0.17.1

Choose a tag to compare

@carrollco carrollco released this 02 Oct 23:25

Fixes from a review of v0.17.0: the warm pool no longer churns when sessions alternate between languages, a docker session never reports docker's own exec failure as the call's exit, Preflight's image environment check is tighter, and the docker provider's Drain is bounded.

Sessions

  • The warm pool's split no longer churns. In a pool of 4 or fewer containers, every language set but the last one asked for was forgotten at the next open, so sessions alternating between JavaScript and Python moved the whole pool on every open and found the wrong interpreter warm. A set is now forgotten only once its weight falls below a quarter of one open's, and the pool moves a container only when the move is clearly worth it (what one set lacks and another holds beyond its share add up to more than 1.25 containers). Alternating callers now cost one container per session, the claimed one's replacement. docs/sessions.md, "A warm pool".
  • docker's exec failures are never a session call's result. docker exec exits 1 for its own failures (a daemon error, a refused exec), not docker run's 125, so v0.17.0 could return such a failure on a running container as the call's exit 1, with docker's message in its stderr. A session snippet's non-zero exit without the call's start marker is now an error.
  • Drain on the docker provider is bounded by its context while the pool is making a container, and returns when the pool's first container failed. Only an embedder that drains while StartSessionPool runs could hit this; plimsolld could not.

Images

  • Preflight compares an image's search path entries (PATH, NODE_PATH, PYTHONPATH and the others) as the directories they name, so //tmp, /usr/../tmp or /./work are refused like /tmp and /work, and it refuses NODE_COMPILE_CACHE, which makes node load cached compiled code for every module.

Clients

  • @plimsollmark/client and the Python client are at 0.17.1, unchanged from 0.17.0 apart from the version.

Upgrading

  • An image whose ENV sets NODE_COMPILE_CACHE, or a search path entry that resolves into /tmp, /work or another writable place however it is spelled, now fails Preflight.