Repository navigation
v0.3.0
Pre-releasev0.3.0 / 2026-09-01
HotCell::Client
Added
- The installed
Dockerfilestrips the setuid and setgid bits off every binary in the image as its last root step, so the strip covers anything aRUNabove it installed rather than the base image alone. The cell already runs unprivileged under--cap-drop ALLand--security-opt no-new-privileges, which makemount,suand the rest inert, so this removes escalation tools a security image should not carry rather than closing a reachable path. The scaffold also now documents what it deliberately does not do for you: build withdocker build --pull, pin a base digest and refresh it on a schedule if you need to name the exact image you shipped, and commit a platform-correctGemfile.lockand build frozen for a reproducible dependency graph. Frozen mode is off by default so a freshly installed scaffold builds before you have generated a lockfile. An upgrade leaves an existingDockerfilealone, so a cell installed before this needs the strip added by hand and the image rebuilt.
Breaking
HotCell.describe_cellsno longer warns about a client whose operation the cell does not carry, andHotCell.clientsis gone with it. The check could not tell a client the application calls from one it merely loaded, so a cell carrying a subset of what a gem ships warned on every boot. A request for an operation the cell does not carry is refused asunsupported, naming the operation, and that failure is transient and reaches the application's error reporting. An application that wants a boot check can write one over its own configuration.
Fixed
-
The installed
DockerfilesetsOMP_NUM_THREADSandOMP_THREAD_LIMITto the container'scpus. OpenMP sizes its thread pool from the host's core count, and a container'scpusquota is a CFS quota rather than an affinity mask, so libvips and ImageMagick asked a 98-core host for 98 threads however small the cell's share of it. A thread stack is 8MB of private anonymous memory, whichRLIMIT_DATAcharges, so the pool alone cleared the cell'smemorylimit — and libgomp callsexit(1)on the firstpthread_createit cannot satisfy. An upgrade leaves an existingDockerfilealone, so a cell installed before this needs both added by hand and the image rebuilt.docs/DEPLOYMENT.mdcovers why the guard has to be a test rather than a deploy to beta: the failure exists only at production's core count. -
The installed
Dockerfileapplies Debian's pending security patches with anapt-get upgradeafter theFROM.docker build --pulltakes the newest base tag, but the tag itself can sit behind an advisory already intrixie-securityuntil docker-library/ruby rebuilds it, so a clean build shipped a fixable High that no rebuild of the cell could clear. Upgrading during the build makes the image's patch level its own rather than upstream's release cadence, and stops the class rather than the one advisory. An upgrade leaves an existingDockerfilealone, so a cell installed before this needs the line added by hand and the image rebuilt.
HotCell::Core and HotCell::Server
Added
- A worker's file descriptor 2 is a pipe to the supervisor, which drains it as it runs and attaches the last 512 bytes to the
worker.killedthat reports its death, ashotcell.stderr, and to the failure the caller receives, asstderr. A C library that callsexit()raises nothing, soworker.crashedis never written and the connection carries a barecrashed; the account of what happened was on fd 2, which goes to the container runtime's log driver, where the fleet's collector drops complete non-JSON lines at ingest.libgomp: Thread creation failed: Resource temporarily unavailableis the line this exists for. The field makes a death legible rather than shipping a cell's stderr anywhere, so a worker that warns and then answers normally reports nothing. The capture is best effort and untrusted: fd 2 stays non-blocking so a warning written from inside libvips can never wait on the supervisor's scheduling, and everything a worker spawned inherits the descriptor.docs/LOGS.mdstates what the field is and is not.
HotCell::Server
Added
-
request,request.abandoned,worker.crashed,worker.killedandworker.undispatchablecarryhotcell.op, the operation the line is about. A cell runs several operations at once and they do not share limits, soworker.killed cause=fsizeon a host serving three PDF operations named none of them, and no join was available elsewhere: the response carries no operation, andhotcell_killedis taggedcellandcauseonly. The worker parses the name out of the request; the supervisor, which never reads one, learns it from the worker's report and holds it, because a killed worker cannot write its ownworker.killed. Where the name is not known the field isnullrather than an earlier request's name. See docs/LOGS.md. -
worker.undispatchableis the one line where the supervisor reads a request, since the worker died before the dispatch write and nothing else has read it. It peeks rather than reads, taking neither the bytes nor the caller's descriptors off the connection, and never waits for a request that has not arrived.
Fixed
Operation#run_toolcarriesOMP_NUM_THREADSandOMP_THREAD_LIMITfrom the cell's environment into the environment it writes for a tool. A tool sees only what its operation wrote for it, so the image's bound would otherwise have applied to in-process libvips and to nothing the cell execs.
ActiveStorage::HotCell::Server
Fixed
-
activestorage-hotcell-serverrequiresmini_magick >= 5.2.0andruby-vips >= 2.2.1. The declared floors were4.0and2.2, butMagickOperationsetsMiniMagick.restricted_env=at require time andVipsOperation'sbefore_forkguard requiresVips.block_untrusted, neither of which exists at those floors. A bundle that satisfied the gemspec could resolvemini_magick5.1.2 orruby-vips2.2.0 and fail to boot — aNoMethodErroron the magick side, a fail-closedConfigurationErroron the vips side — taking the whole conversion toolchain offline. -
MiniMagick.cli_envcarries the cell'sOMP_NUM_THREADSandOMP_THREAD_LIMIT, somagickruns under the image's bound rather than sizing its pool from the host's cores.MiniMagick.restricted_envis what had removed them.
Tooling
Changed
-
bin/conformanceverifies the container flags it claims rather than asserting them. The isolation check read the network interfaces, whether the root took a write, whether scratch wasnoexec, and an exec'd tool's environment — it never read the bounding capability set, so--cap-drop ALLcould be dropped from the run and every check still passed. The read-only check was weak the same way:File.writable?("/")is false for the cell's non-root user whether or not the root is read-only. Theisolationoperation now reports what the kernel exposes — the bounding capability set and the no-new-privileges bit from/proc/self/status, the uid, and the root mount's read-only option from/proc/self/mounts— and the battery requires each. A field the cell cannot read comes backniland fails, so a run that cannot see a flag is never mistaken for one that set it. -
bin/conformanceruns its own negative controls: it re-invokes itself once per flag withHOTCELL_CONFORMANCE_DROP, booting withoutnetwork,read-only,cap-dropandno-new-privilegesin turn, and requires each run to fail at that flag's own assertion rather than merely exiting non-zero. Thetmpfs-noexecnegative is decided by probing the runtime, so a runtime that force-mountsnoexecreportsSKIPwith the reason instead of passing vacuously. The isolation checks run first, before the timing-sensitive deadline and overload checks, so a dropped flag fails there rather than behind a flake.examples/gateis a fast container-free guard that drives the isolation check with fabricated results, proves it rejects every insecure or unreadable value, and holds the negatives' expected messages against the battery's assertions so a reworded assertion cannot rot a negative into a grep that never matches. -
bin/conformanceandbin/loadno longer pass--ulimit stack=2097152:2097152.docs/DEPLOYMENT.mdforbids a lowered stack — an overflow becomes aSIGSEGVthe supervisor reports as a transientkilled/crashed— so the helpers were measuring a shape operators are told not to deploy.examples/gateasserts neither script sets it.
Fixed
docs/DEPLOYMENT.mdno longer tells operators that conformance cannot observecap-drop,no-new-privilegesor the uid, and thatread-onlyis checked by attempting a write. All four were true before the checks above and false after.