Repository navigation
v1.0.0
First stable release. The exported API is unchanged from 0.2.0 and is now
covered by semantic versioning; what this release adds is evidence about how
it behaves on the platforms it claims to support, gathered by actually
running the suite on Linux rather than inferring from macOS. That turned up
two real defects, both invisible on macOS, and one broken CI check.
The first of the two ### Known Limitations recorded under 0.1.0 -- the
drain-timeout bound -- is resolved. The second, a group of timing-sensitive
process-group tests, is not: it is narrowed, re-diagnosed, and still skipped
on Linux.
Correctness
-
drain-timeout-secondsis now honoured on Linux. A copier thread
drained a child's stdout/stderr with a blockingread(2), and
%drain-copiersunstuck one that overran its deadline by force-closing
the stream under it. Closing a descriptor wakes a thread already parked
inread(2)on macOS/BSD, but that is a BSD courtesy rather than
anything POSIX promises, and Linux does not do it: the reader stayed
parked until whoever else held the pipe's write end let go. After
sh -c "sleep 5 & exit 0"that is a backgrounded descendant which
outlives the leader by design, so the wait was effectively unbounded.
The observed symptom on Linux was worse than 0.1.0's note described --
not merelyrunoverrunning its documented bound, butrunsignalling
process-io-error :cleanup("Copier thread did not terminate after its
stream was closed") once the force-close failed to land.The read loop now waits with
poll(2)on a bounded timeout and checks a
stop flag between turns (%await-fd-readable), so the decision to give up
is taken by the reader itself instead of being inflicted on it through the
descriptor. The bound then holds by construction on any POSIX platform
rather than by accident on some. The flag is checked before readiness,
which bounds the loop even against a child that never stops producing:
one poll interval plus one boundedread(2)per turn. Polling costs no
wakeup rate the library was not already paying, sincecommunicate's own
deadline loop already runs at+default-poll-interval+.%drain-copiers' escalation gained the cooperative stop as its first
rung -- ask, then force-close, then signal -- ordered by what each costs
when it fires, in the same continuation-passing shape as
escalate-unless-gone's SIGTERM -> SIGKILL -> give-up. Force-closing is
now a fallback for a copier parked where a flag cannot reach it (a
read-sequenceon a non-fd stream) rather than the primary mechanism.
%communicate-base's cleanup path was reordered to match: it retires the
copiers before their streams, soclose-process-streamsno longer closes
a descriptor under a live reader on every pass -- a descriptor the OS is
free to reissue the moment it is closed. -
:on-timeoutand:on-cancelare validated at the entry points that
own them. Every entry point resolves these by comparing against
:errorand treating anything else as:return, which made an
unrecognised value indistinguishable from a deliberate:returninstead
of an error.run,run-commandandrun-pipelinecompounded it: each
handscommunicatea hardcoded:on-timeout :returnand decides for
itself whether to signal, so%validate-communication-options' guard
never saw what the caller wrote. A misspelt:errrortherefore read as
:returnand silently swallowed the very timeout or cancellation the
caller had asked to have signalled.%validate-outcome-policynow guards
each policy where it is accepted --run(:on-timeout,:on-cancel),
run-command(both),run-pipeline(both), andcommunicate
(:on-cancel, which had no guard either).
Testing
-
The drain-timeout regression test ("run returns boundedly when an exited
leader leaves a pipe-holding descendant") is no longer skipped on Linux --
the fix above makes it platform-independent, and it passes on CI. -
The other seven Linux skips are re-diagnosed rather than removed. They had
been filed under the same "process-group/communicate timing is not
guaranteed identical on Linux" heading as the drain bug, which conflated
two unrelated things. They do not share its cause: each asserts that a
process group is gone within a 0.1s grace period, which a contended
shared CI runner cannot reliably deliver. (They pass on an uncontended
aarch64 Linux container and fail on GitHub's x86_64 runners -- contention,
not architecture.) Their skip reasons now say that instead. They are not
given more headroom because a timing assertion loose enough to survive
arbitrary contention no longer asserts the timing; the honest fix is to
make the assertion event-driven rather than deadline-driven, which is
deferred. -
The coverage ratchet had been failing every Linux CI build, on a
comparison it should never have made. The floor tracked the figure from a
complete run (macOS, nothing skipped) but was enforced against the reduced
Linux run, so CI reported the skipped tests as a coverage "regression" --
86.9% against an 87.0% floor -- with every test passing. The floors are now
enforced only when the whole suite ran (+suite-complete-p+); otherwise
coverage is reported with an explicit note that it is not comparable. -
A guard-clause test that names a program which does not exist proves
nothing.t/run-timeout-test.lisp's "run rejects invalid timeout
controls before spawning" spawned/bin/true, which macOS does not ship
(truelives in/usr/binthere). All ten of its assertions passed on a
process-launch-errorfor the missing file -- identically, and just as
green, whether or not the guard under test existed. That is what hid the
:on-timeoutgap above: the assertion only had to mean something once
the suite was first run on Linux, where/bin/truedoes exist. A
%true-programfixture now resolvestruethrough the ambientPATH,
matching the existing%spawn-sleepingfixture, and is used by the 17
call sites that actually spawn. The remaining literal/bin/true
occurrences are inmake-commanddata tables that never spawn, where the
string is inert. -
Added a regression test asserting that all four entry points reject an
unrecognised outcome policy rather than reading it as:return. -
Coverage ratchet advanced to 87.4% expression / 81.5% branch (from
87.0/79.5), measured on a complete run.
Documentation
-
Documented the timeout and cleanup deadlines.
grace-period,
poll-interval,timeout-signal,kill-signaland
drain-timeout-secondspreviously appeared only insiderun's signature
in the options table -- five knobs with no stated meaning or default,
covering the escalation behaviour that is the library's whole reason to
exist. The execution guide now gives them a table of their own, explains
why signals go to the process group rather than the child, and explains
why draining has a deadline separate from the child's (a descendant that
outlives the leader inherits the same pipe and can hold it open
indefinitely). -
Added a "Running the suite on both platforms" section to the development
guide, covering the container invocation for checking Linux behaviour
locally, and the two habits this release's bugs argue for: never name a
program a guard-clause test does not intend to execute, and prefer fixing
a platform difference to skipping the test that catches it.
Build/environment
flake.nixhardcoded"0.2.0"in four separate places (the docs
derivation, thecl-process-kitpackage, and thecl-process-kit-pty
package), plus two more incl-process-kit.asd(cl-process-kitand
cl-process-kit/test) -- six places a release has to remember to bump in
lockstep, with no build-time check that they agree.nerima-lisp/cl-boundary-kit
v0.6.0 already fixed the identical problem in its ownflake.nix(found
while auditing this project's own environment setup for the same class
of drift); ported the same technique here: aversionletbinding
parses the:versionform out ofcl-process-kit.asdline-by-line
(Nix'sbuiltins.matchis whole-string-anchored and.doesn't span
newlines, so a single multi-line regex doesn't work) and every Nix
package nowinherits it. A release now only ever edits the.asd.cl-process-kit/ptyandcl-process-kit/pty-testwere missing the
:version/:author/:maintainer/:license/:homepage/:bug-tracker/
:source-controlmetadata the other two systems in the same file
already carry -- added it for consistency.
Production readiness
- Added
SECURITY.md,SUPPORT.md, andCONTRIBUTING.md(GitHub's
standard community-health filenames, which its UI surfaces automatically
regardless of README content) -- this repository had none, unlike sibling
nerima-lispprojects (cl-weavehas all three). Scoped and sized for
this project specifically rather than copied verbatim:SECURITY.md
names the concrete classes of report that actually apply here (shell/
argument injection viarun-shell/make-command, process-group
isolation failures,spawn-native's privilege/credential handling), not
a generic template;CONTRIBUTING.mddocuments the coverage ratchet and
nix flake checkas the authoritative (not just convenient) verification
step, matching how this project is actually developed and verified
throughout this changelog.
CPS
src/copier.lisp's%drain-copiershad a two-step "join within the
remaining time; if that times out, force-close the stream and give it
one more brief join" escalation inlined as nestedwhens. Extracted
%join-copier-unless-timed-out (copier timeout on-timeout), which calls
theon-timeoutcontinuation only if the join actually times out --
the same shape ascommunicate.lisp'sescalate-unless-gone, applied
one level down at the copier-thread join instead of the process-group
signal escalation.%drain-copiers's two-step escalation is now two
nested calls to it instead of two nestedwhens, matching this
codebase's establishedcall-with-X/continuation-parameter idiom.
Correctness
src/copier.lisp's%join-copieraccepted an&optional timeoutand, if
omitted, joined a copier thread with no bound at all -- an unbounded wait
on a thread that can be blocked reading a live (possibly-hung) child
process's stdout/stderr. Audited everySB-THREAD:JOIN-THREAD/
PROCESS-WAIT/SB-EXT:PROCESS-WAITcall insrc/for this same class of
risk: the other four unboundedPROCESS-WAITcalls (communicate.lisp,
native-spawn.lisp,process-handle.lisp,spawn.lisp) are each
provably safe by construction -- reached only after the process has
already been SIGKILLed (which POSIX guarantees terminates it) or is
already OS-reported as terminal, so the wait there reaps a dying/dead
process rather than blocking on a live one.%join-copierwas the one
genuine gap. Madetimeouta required parameter (it was never actually
called without one -- the unbounded branch was dead in practice, not just
in principle) and updated its docstring to state the "never unbounded"
contract explicitly.
Data/logic separation
src/native-spawn.lisp's%native-spawn-phasemapped the native spawn
trampoline's byte phase codes to keywords with acaseform embedded
directly in the lookup function. Hoisted the mapping into a top-level
+native-spawn-phases+alist, leaving%native-spawn-phaseas a plain
(or (cdr (assoc number +native-spawn-phases+)) :unknown)traversal --
the phase table can now be read, diffed, or extended on its own, separate
from the lookup logic that applies it, matching the pattern already used
elsewhere insrc/(e.g.+task-terminal-state-rules+).
Readability
src/communicate.lisp:communicate's cancellation watcher body -- a
~20-line anonymous lambda passed straight tosb-thread:make-thread--
is now a named local function,run-cancellation-watcher, alongside its
labelssiblings (finished-p,mark-cancelled, etc.), so the thread
creation call reads as#'run-cancellation-watcherinstead of an
unlabeled inline block. Extracted withparedit-cli.src/spawn.lispandsrc/command.lispeach independently validated a
"KEY=VALUE"environment-string entry (non-empty key before=,
rejecting a duplicate key) with byte-for-byte identical logic under two
different error-message labels. Extracted the shared shape into
%validate-environment-entry-shape(src/command.lisp, next to
%validate-environment-entries, its other caller) via
paredit refactor extract-function;spawn.lisp's
%validate-environmentnow calls it with"ENVIRONMENT", matching the
uppercase label convention%validate-environment-entriesalready uses
for"ENVIRONMENT-POLICY"/"ENVIRONMENT-UPDATE"(previously
spawn.lispalone used mixed-case "Environment" in its error text; no
test asserted on the literal message, so this also fixes a real
inconsistency, not just a duplication). Found viaparedit inspect similarity --threshold 0.85 src(score 44.5, 46 shared AST nodes) rather
than searched for. Net effect: fewer total expression/branch points to
cover (4736/676 -> 4726/670), sosrc/'s coverage percentage rose
slightly (87.5% -> 87.6% expression) as a side effect of having less
duplicated code, not a coverage change pursued for its own sake.
Testing
t/validation-test.lisp'smake-command/spawn-nativeguard-clause
tables now usecl-weave:it-eachinstead of a hand-rolleddolistover a
list of(label . thunk)conses: every one of the 22 + 3 malformed-input
rows is now its own independently-named, independently-reported test case
(e.g. "rejects an empty program", "rejects a dotted (improper)
:environment-update list") instead of one aggregate pass/fail that hid
which row actually failed. Test count 152 -> 174; behavior and coverage
unchanged, since it's the same guard clauses exercised the same way.- Added
t/mutation-test.lisp:cl-weave:run-mutationsmutation-tests
process-success-pandpipeline-success-p(src/command.lisp),
systematically flipping their comparison/boolean logic and asserting the
existing case battery notices every one-operator change
(cl-weave:assert-mutation-scoreat a required 1.0 -- no surviving
mutants). Coverage proves a line executed, not that a wrong result there
would be caught; mutation testing closes that specific gap for these two
pure predicates. Follows the exact patternnerima-lisp/cl-tty-kit
already established for this ecosystem (contrib/weave-mutation-tests.lisp):
theDEFUNbody is read live fromsrc/command.lispon every run (never
copied into the test file), so there is nothing to fall out of sync with
the real implementation.
Dependencies
- Bumped
cl-weavev0.11.0 -> v1.0.0,cl-boundary-kitv0.5.0 -> v0.6.0,
cl-log-kitv1.1.0 -> v1.6.0,cl-tty-kitv0.5.0 -> v0.6.0. Checked each
upstream changelog/diff for breaking changes before bumping: none touch
the surface this library actually calls (cl-boundary-kit's v0.6.0
unbounded-wait-by-default change is scoped to its own
process-boundary-run/process-kit-run-fn, which this library never
calls;cl-log-kit's "logger as explicit first argument" breaking change
landed at v1.0.0 and%log(src/logging.lisp) already passed it
explicitly). Verified with a fullnix flake check(checkout-tests
152/152 at 87.5%/79.9% coverage,pty-tests6/6), not just a local
sbcl --scriptrun. flake.nix's four nerima-lisp inputs are consumed purely as raw ASDF
source trees (buildASDFSystemsrc, orCL_SOURCE_REGISTRYat
runtime) -- none of their own flakepackages/checksoutputs are ever
used.cl-weave,cl-boundary-kit, andcl-log-kitwere still declared
as full flakes (missingflake = false, unlike the already-correct
cl-tty-kit), sonix flake updatepulled in their entire transitive
dev-only input graphs (treefmt-nix,cl-json-kit, andcl-json-kit's
own sub-inputs) into this project'sflake.lockfor no reason. Added
flake = falseto all three, shrinkingflake.lockfrom 34 nodes to 6.cl-tty-kitv0.6.0 vendorsnerima-lisp/cl-prologas a git submodule
(vendor/cl-prolog) and self-registers that path from its own.asd;
a plaingithub:fetch does not follow submodules (regardless of a
?submodules=1query string, which only thegit+https://fetcher
honors), so the Nix sandbox had an emptyvendor/cl-prologand
cl-process-kit/pty-testfailed to load withComponent #:CL-PROLOG not found, required by #<SYSTEM "cl-tty-kit">. This only broke insidenix flake check's sandboxedpty-tests, not localsbcl --scriptruns
against a manually-cloned~/ghqcheckout that already had the
submodule populated -- another case (see the[0.2.0]entry below) of a
local run passing while the authoritative sandboxed check does not.
Fixed by switchingcl-tty-kit's input to
git+https://github.com/nerima-lisp/cl-tty-kit?ref=refs/tags/v0.6.0&submodules=1.
Documentation
- Audited every public function signature documented in
README.mdand
docs/src/guide/*.mdagainst the actual&keylists insrc/run.lisp,
src/spawn.lisp, andsrc/communicate.lisp(plus the condition
hierarchy insrc/conditions.lisp, theprocess-result/pipeline-result
structs insrc/types.lisp, and the PTY/event/task accessors in
src/pty.lisp/src/async-events.lisp/src/async-task.lisp, which were
already accurate). Foundrun'soutputkeyword (%run-base) -- the
stdout counterpart oferror's policy, controlling whether stdout is
:captured,:inherited live, or sent to a stream -- completely
undocumented, even though it has been part of the public API since the
[0.1.0]entry below;error's own accepted values were also
under-described (listed only:capture/:output, omitting:inherit
and stream targets that%valid-run-output-policy-palso accepts). Also
foundREADME.md'srun/communicatesignatures missing
decoding-error-policyanddocs/src/guide/execution.md'srun
signature anddocs/src/guide/async.md'sspawnsignature missing
fd-limit-- each doc had independently drifted to cover only the
option added most recently in its own edit history. Fixed all of these
indocs/src/guide/execution.mdanddocs/src/guide/async.md, which
remain the single authoritative reference (see the next entry for why
README.mdno longer duplicates them). README.mdhad grown a ~450-line "API" section duplicating the option
reference already maintained on the MkDocs site (docs/src/guide/*.md,
docs/src/reference/results-and-conditions.md) -- the exact duplication
responsible for the signatures drifting out of sync in the entry above,
and for the site's own "Native spawn trampoline" page having no README
counterpart at all. SimplifiedREADME.mdto a landing page (intro,
install, one minimal usage example, links to the published docs) and
removed the API section and the Development section's test-suite/
source-layout breakdown entirely, leavingdocs/src/as the only place
a signature or option list is written down.