Repository navigation
Releases: nerima-lisp/cl-process-kit
Release list
v3.4.0
Highlights
- Added opt-in
:use-posix-spawnsupport for process launches while preserving the existing default behavior. - Added directory-aware POSIX spawn launches without changing
:search nilsemantics. - Propagated the option through synchronous, asynchronous, command-spec, and pipeline entry points.
No caller changes are required unless opting into POSIX spawn launches.
v3.3.1
Adds an event-driven async process-completion API and fixes a bug that broke executable builds in v3.3.0, which was not published.
Added
communicate-async/process-task: an event-driven counterpart tocommunicatethat returns aprocess-taskimmediately instead of blocking.await-processblocks for the result (optionally with a timeout, returning(values nil nil)on expiry), andcancel-processcancels a pending task. Options includeevent-callback,event-queue-capacity(default 64),event-overflow-policy(:drop-newestby default, or:block), andevent-history-capacity.
Fixed
- v3.3.0 created its two async task executors in top-level
defvarinitforms, so merely loading the system started twelve threads. SBCL refusessave-lisp-and-die(and soasdf:program-op) andsb-posix:forkwhile other threads are running, which broke executable builds for every consumer -- this is why v3.3.0 was not published. The executors are now created lazily on first use, and the new exportedshutdown-process-kitstops and joins them; it also runs fromsb-ext:*save-hooks*, so a saved image carries no thread objects and recreates the pools on demand after restart.
v3.1.0
Added
select-fds/wait-for-input: aselect(2)-based file-descriptor
readiness primitive for event loops that need to block on several raw
descriptors at once -- several PTY masters, a unix socket, and stdin,
say -- and be told which woke it. This is the one primitive an event
loop could not previously build out of the rest of the library:
process-wait/communicateeach drive a single child.select-fds
returns the three ready read/write/exceptional descriptor sub-lists;
wait-for-inputis the common read-only, single-descriptor-list case
reduced to one argument and one return value. Both retry transparently
onEINTRagainst the original deadline, so a periodic signal (
SIGWINCH,SIGCHLD) never truncates or extends the caller's timeout.
New public conditionsfd-set-overflow(a descriptor at or above
+maximum-fd+) andfd-wait-failed(any other failingerrno) cover
the two failure modes; every other bound documented onselect-fds
cannot be reached fromSB-UNIX:UNIX-FAST-SELECT's own argument checks.
See the File-Descriptor Readiness guide.
v3.0.1
v3.0.1: bump cl-codec-kit to v0.3.1, fixing double-replacement on a U…
v3.0.0
v3.0.0: migrate UTF-8/octet handling onto cl-codec-kit (breaking: EXT…
v2.0.0
Added
-
process-kit/pty:pty-try-wait/pty-alive-p: non-blocking liveness
checks for a PTY session, the analogues ofprocess-try-wait/
process-alive-p. Found the same way aswith-pty-processtwo commits
earlier -- comparing the PTY subsystem's public surface against the
main library's and findingpty-wait(blocking, possibly escalating)
had no non-blocking counterpart.pty-wait's own polling loop now
calls the publicpty-try-waitinstead of the internal%try-wait
directly, matchingprocess-wait's existing call-the-public-function
convention inprocess-handle.lisp. -
process-kit/pty:with-pty-process/call-with-pty-process: the PTY
analogue ofwith-process/call-with-process's continuation-passing
cleanup contract, binding a variable to aspawn-ptyresult for a body
and closing it viapty-closeon the way out, success or error. The
main library had this idiom forprocess-handle; the PTY subsystem
never got the equivalent, so all 6 oft/pty-test.lisp's cases
hand-rolled the identicalunwind-protect/pty-closeshape. Now they
read as(with-pty-process (process (spawn-pty ...)) ...). -
cl-weave:it-fuzzproperty (t/property-test.lisp):run's:replace
UTF-8 decoding path must never signal on arbitrary octet input, across
100 generated trials, generalizing the 3 hand-picked malformed sequences
t/edge-coverage-test.lispalready covered into a genuine property of
the full byte space. Firstit-fuzzuse in this codebase;it-property,
it-each, andrun-mutationswere already adopted in earlier passes. -
checks.maxLineLengthinflake.nix: failsnix flake checkif any
src//t/line exceeds 100 columns (the org'sCODING_STANDARD.md
rule), matching the existingchecks.maxFileLengthgate's shape. Added
after actually closing the gap it enforces, not before: an earlier
pass's "wrap every line at 100 columns" commit was scoped tosrc/
only (its own message says so), andt/was never swept -- 100 lines
across 12 test files exceeded 100 columns until this pass fixed them by
hand, preserving each file's existing hand-crafted formatting rather
than running a canonical auto-formatter (paredit edit format's
canonical style turned out to differ enough from this codebase's actual
conventions -- breaking every keyword-argument pair onto its own line,
removing blank lines betweenitcases -- that using it wholesale
would have been a large, unrelated stylistic churn; used only on the two
files,t/pty-test.lispandt/native-spawn-test.lisp, that had no
existing hand-crafted style to preserve in the first place, both single
giant crammed one-liners).
Changed
-
Extracted
%task-submit-output's and%task-finish's identical "flush
pending-drops into an:overflowevent" block (src/async-events.lisp)
into a shared%flush-pending-drops-event, viaparedit refactor extract-function --at <offset> --infer-params. Raised the coverage
ratchet floors to match (expression 87.4% -> 87.7%, branch 81.3% ->
81.4%): the dedup genuinely closed one branch gap (124 -> 123 uncovered),
not just diluted the denominator. -
Unified three near-duplicate deadline-polling loops onto the single
%poll-untilprimitive thatcommunicate.lisp/pipeline.lispalready
shared -- found by grepping for the "loop until a predicate is true or a
deadline passes, sleeping between checks" shape directly, since it
differs enough call to call (different clock sources, a returned value
vs. a boolean) that the exact-matchparedit inspect duplicatestool
cannot see it as one shape.%poll-untilitself is now defined in
process-handle.lisp(the earliest-loading file among its callers, so
bothprocess-group.lispandcommunicate.lispcan reach it) and
generalized to accept a NILdeadlinefor an unbounded wait:process-handle.lisp'sprocess-waitnow delegates to%poll-until
instead of hand-rolling the same loop; a genuine (not just cosmetic)
fix falls out of this, since%poll-untilcaps its final sleep at
whatever time is actually left before the deadline, where the old loop
always slept the fullpoll-intervalregardless --process-wait
could previously overshoot its deadline by up topoll-interval
before ever re-checking.process-group.lisp's%wait-until-process-group-gonenow delegates
to the same%poll-untilinstead of a hand-rolled copy of exactly the
donepcommunicate.lisp's own%wait-until-group-gonealready used --
the two are the same "process group is gone" check, previously written
down twice.%poll-until's own docstring previously claimed "this is the one
place [every deadline-bounded wait] is written down," which was false
given the two hand-rolled copies above; corrected, and now explains why
process-kit/pty'spty-waitdeliberately keeps its own copy (a
genuinely different, three-way donep with no clock/sleeper injection,
not an oversight).
Verified via two fullnix flake checkruns (the first hit the
documentedSETUP-SERVE-EVENT-PIPEflake onawait-process, an
unrelated async-cursor test that shares no code with any of the three
functions touched here; the second was fully green, 180/180, coverage
88.0%/82.1%, both above the ratchet floor).
-
Extracted
run-pipeline's inline thread-spawningloop/lambdablock
(src/pipeline.lisp) into a namedspawn-stage-threadssibling inside
the samelabelsform as the pre-existingawait-pipeline-stages.
Found via a fresh size-ranked scan of everysrc/*.lispfunction
(paredit inspect outline --output json):run-pipelinewas the
codebase's 2nd-largest function and, unlike its neighbors
communicate-async/communicate, had never had its inline
thread-spawning lambda pulled into a named function. Hand-authored
rather than viaparedit refactor extract-function --infer-params,
since that flag mishandlesloop's clause-keyword syntax (for/in/
collectread as ordinary symbols to infer as parameters).run-pipeline
itself now reads as a linear sequence -- build pipes, spawn stages, close
write ends,(await-pipeline-stages (spawn-stage-threads)), build
result -- with no behavior change.
Added
- Both test entry points (
run-tests.lisp,run-pty-tests.lisp) now bind
cl-weave:*default-timeout-ms*to 30 seconds before running the suite.
The org-wideTEST_STANDARD.mdrequires every repository to set this;
cl-process-kit's suite previously had no PER-TEST ceiling, only a
whole-suite one (nix flake check's owntimeout 180/CI's
timeout-minutes) -- so a single hung test burned the entire budget
before failing, with nothing naming which test hung. 30s leaves generous
headroom above the slowest observed test (~1.1s) while catching a
genuine deadlock with a specificCL-WEAVE:TEST-TIMEOUTdiagnostic.
Reformattedrun-pty-tests.lispfrom a single unreadable one-line form
to matchrun-tests.lisp's style while making this change.
Changed
flake.nix's.asd:versionextraction usescl-nix-forge's dedicated
fromAsdSystemlexer instead of a hand-rolled line-by-line regex.
Strictly stronger, not just shorter:cl-process-kit.asddeclares four
systems sharing one version (cl-process-kit,/test,/pty,
/pty-test), and the old regex read whichever:versionline happened
to come first with no cross-check, whilefromAsdSystemfails the build
loudly if any of them ever disagreed.flake.nix'spackages.*/checks.checkout-tests/checks.pty-tests
now build oncl-nix-forge'slispDerivation/mkScriptCheckprimitives
(new flake input, pinnedv0.4.0) instead ofpkgs.sbcl.buildASDFSystem
plus hand-rolledpkgs.runCommandtest derivations.noForbiddenMarkers,
maxFileLength,formatting, anddocsare untouched -- none of them
involve ASDF, so cl-nix-forge has nothing to offer there. Two things
cl-process-kit's own code needed that neither of the org's two existing
adopters (cl-weave,cl-json-kit) demonstrated:src/pty.lispreads
CL_PROCESS_KIT_PTY_LIBRARYexplicitly at load time (anSB-ALIEN LOAD-SHARED-OBJECTcall needs a real path, not a bare SONAME resolved
offnativeLibraries'LD_LIBRARY_PATH/DYLD_LIBRARY_PATHpropagation
alone), so that env var is still set directly alongsidenativeLibraries;
andnative/spawn.c's trampoline binary is compiled twice on purpose --
once into$out/binfor the distributed package (postInstall), once
into the build sandbox forchecks.checkout-tests(preCheck), since a
check needs the binary to exist before its own build/install phases (what
actually produces the packaged copy) have run.apps/devShellsstay on
the original hand-rolledCL_SOURCE_REGISTRYstring, not
cl-nix-forge:lispScript-- this migration's own research did not verify
that primitive's exact parameter shape the waylispDerivation/
mkScriptCheckwere verified against real builds, and guessing an
unconfirmed API for a part of the flake with no correctness benefit over
the proven string was not worth the risk.
Verified withnix buildon each new package/check individually before
the fullnix flake check(179/179 + 6/6 PTY, coverage unchanged at
87.8%/81.5%, both native artifacts confirmed present and loadable).
Changed (BREAKING)
next-process-eventnow returns a singleprocess-event-stepstruct
(process-event-step-event/-cursor/-status/-gap-count) instead of
four raw values. The org-wideAPI_STANDARD.mdnames this exact function,
by name, as the canonical bad example of returning more than two values
((values a b)at most, or a struct beyond that) -- a caller that only
wantedgap-counthad to destructure all four positionally. Requires a
major version bump at the next...
v1.0.1
A packaging fix. No source change, and the exported API is identical to 1.0.0.
Fixed
-
flake.nixasked forcl-log-kit/v1.6.0, a tag that has never existed —
cl-log-kithas only ever releasedv1.0.0. The v1.0.0 tag of this
repository still carries that reference, so anyone pinning
cl-process-kit/v1.0.0as a flake fails atnix flake lockfrom a cold
cache. Existing lock files hide it by holding an already-resolved revision,
which is why CI stayed green. Three repositories pin this one that
way —cl-boundary-kit,cl-cc-runtimeandcl-cli— and would all have
broken at once on the nextflake-update.ymlrun.Tags are never moved, so v1.0.0 stays as it is and this release exists to
give downstream something that resolves. Sibling versions are now read out of
each pinned source's own.asdrather than repeated here, so the two cannot
drift apart again.
v1.0.0
First stable release. The exported API is unchanged from 0.2.0 and is now
covered by semantic versioning; what this release adds is evidence about how
it behaves on the platforms it claims to support, gathered by actually
running the suite on Linux rather than inferring from macOS. That turned up
two real defects, both invisible on macOS, and one broken CI check.
The first of the two ### Known Limitations recorded under 0.1.0 -- the
drain-timeout bound -- is resolved. The second, a group of timing-sensitive
process-group tests, is not: it is narrowed, re-diagnosed, and still skipped
on Linux.
Correctness
-
drain-timeout-secondsis now honoured on Linux. A copier thread
drained a child's stdout/stderr with a blockingread(2), and
%drain-copiersunstuck one that overran its deadline by force-closing
the stream under it. Closing a descriptor wakes a thread already parked
inread(2)on macOS/BSD, but that is a BSD courtesy rather than
anything POSIX promises, and Linux does not do it: the reader stayed
parked until whoever else held the pipe's write end let go. After
sh -c "sleep 5 & exit 0"that is a backgrounded descendant which
outlives the leader by design, so the wait was effectively unbounded.
The observed symptom on Linux was worse than 0.1.0's note described --
not merelyrunoverrunning its documented bound, butrunsignalling
process-io-error :cleanup("Copier thread did not terminate after its
stream was closed") once the force-close failed to land.The read loop now waits with
poll(2)on a bounded timeout and checks a
stop flag between turns (%await-fd-readable), so the decision to give up
is taken by the reader itself instead of being inflicted on it through the
descriptor. The bound then holds by construction on any POSIX platform
rather than by accident on some. The flag is checked before readiness,
which bounds the loop even against a child that never stops producing:
one poll interval plus one boundedread(2)per turn. Polling costs no
wakeup rate the library was not already paying, sincecommunicate's own
deadline loop already runs at+default-poll-interval+.%drain-copiers' escalation gained the cooperative stop as its first
rung -- ask, then force-close, then signal -- ordered by what each costs
when it fires, in the same continuation-passing shape as
escalate-unless-gone's SIGTERM -> SIGKILL -> give-up. Force-closing is
now a fallback for a copier parked where a flag cannot reach it (a
read-sequenceon a non-fd stream) rather than the primary mechanism.
%communicate-base's cleanup path was reordered to match: it retires the
copiers before their streams, soclose-process-streamsno longer closes
a descriptor under a live reader on every pass -- a descriptor the OS is
free to reissue the moment it is closed. -
:on-timeoutand:on-cancelare validated at the entry points that
own them. Every entry point resolves these by comparing against
:errorand treating anything else as:return, which made an
unrecognised value indistinguishable from a deliberate:returninstead
of an error.run,run-commandandrun-pipelinecompounded it: each
handscommunicatea hardcoded:on-timeout :returnand decides for
itself whether to signal, so%validate-communication-options' guard
never saw what the caller wrote. A misspelt:errrortherefore read as
:returnand silently swallowed the very timeout or cancellation the
caller had asked to have signalled.%validate-outcome-policynow guards
each policy where it is accepted --run(:on-timeout,:on-cancel),
run-command(both),run-pipeline(both), andcommunicate
(:on-cancel, which had no guard either).
Testing
-
The drain-timeout regression test ("run returns boundedly when an exited
leader leaves a pipe-holding descendant") is no longer skipped on Linux --
the fix above makes it platform-independent, and it passes on CI. -
The other seven Linux skips are re-diagnosed rather than removed. They had
been filed under the same "process-group/communicate timing is not
guaranteed identical on Linux" heading as the drain bug, which conflated
two unrelated things. They do not share its cause: each asserts that a
process group is gone within a 0.1s grace period, which a contended
shared CI runner cannot reliably deliver. (They pass on an uncontended
aarch64 Linux container and fail on GitHub's x86_64 runners -- contention,
not architecture.) Their skip reasons now say that instead. They are not
given more headroom because a timing assertion loose enough to survive
arbitrary contention no longer asserts the timing; the honest fix is to
make the assertion event-driven rather than deadline-driven, which is
deferred. -
The coverage ratchet had been failing every Linux CI build, on a
comparison it should never have made. The floor tracked the figure from a
complete run (macOS, nothing skipped) but was enforced against the reduced
Linux run, so CI reported the skipped tests as a coverage "regression" --
86.9% against an 87.0% floor -- with every test passing. The floors are now
enforced only when the whole suite ran (+suite-complete-p+); otherwise
coverage is reported with an explicit note that it is not comparable. -
A guard-clause test that names a program which does not exist proves
nothing.t/run-timeout-test.lisp's "run rejects invalid timeout
controls before spawning" spawned/bin/true, which macOS does not ship
(truelives in/usr/binthere). All ten of its assertions passed on a
process-launch-errorfor the missing file -- identically, and just as
green, whether or not the guard under test existed. That is what hid the
:on-timeoutgap above: the assertion only had to mean something once
the suite was first run on Linux, where/bin/truedoes exist. A
%true-programfixture now resolvestruethrough the ambientPATH,
matching the existing%spawn-sleepingfixture, and is used by the 17
call sites that actually spawn. The remaining literal/bin/true
occurrences are inmake-commanddata tables that never spawn, where the
string is inert. -
Added a regression test asserting that all four entry points reject an
unrecognised outcome policy rather than reading it as:return. -
Coverage ratchet advanced to 87.4% expression / 81.5% branch (from
87.0/79.5), measured on a complete run.
Documentation
-
Documented the timeout and cleanup deadlines.
grace-period,
poll-interval,timeout-signal,kill-signaland
drain-timeout-secondspreviously appeared only insiderun's signature
in the options table -- five knobs with no stated meaning or default,
covering the escalation behaviour that is the library's whole reason to
exist. The execution guide now gives them a table of their own, explains
why signals go to the process group rather than the child, and explains
why draining has a deadline separate from the child's (a descendant that
outlives the leader inherits the same pipe and can hold it open
indefinitely). -
Added a "Running the suite on both platforms" section to the development
guide, covering the container invocation for checking Linux behaviour
locally, and the two habits this release's bugs argue for: never name a
program a guard-clause test does not intend to execute, and prefer fixing
a platform difference to skipping the test that catches it.
Build/environment
flake.nixhardcoded"0.2.0"in four separate places (the docs
derivation, thecl-process-kitpackage, and thecl-process-kit-pty
package), plus two more incl-process-kit.asd(cl-process-kitand
cl-process-kit/test) -- six places a release has to remember to bump in
lockstep, with no build-time check that they agree.nerima-lisp/cl-boundary-kit
v0.6.0 already fixed the identical problem in its ownflake.nix(found
while auditing this project's own environment setup for the same class
of drift); ported the same technique here: aversionletbinding
parses the:versionform out ofcl-process-kit.asdline-by-line
(Nix'sbuiltins.matchis whole-string-anchored and.doesn't span
newlines, so a single multi-line regex doesn't work) and every Nix
package nowinherits it. A release now only ever edits the.asd.cl-process-kit/ptyandcl-process-kit/pty-testwere missing the
:version/:author/:maintainer/:license/:homepage/:bug-tracker/
:source-controlmetadata the other two systems in the same file
already carry -- added it for consistency.
Production readiness
- Added
SECURITY.md,SUPPORT.md, andCONTRIBUTING.md(GitHub's
standard community-health filenames, which its UI surfaces automatically
regardless of README content) -- this repository had none, unlike sibling
nerima-lispprojects (cl-weavehas all three). Scoped and sized for
this project specifically rather than copied verbatim:SECURITY.md
names the concrete classes of report that actually apply here (shell/
argument injection viarun-shell/make-command, process-group
isolation failures,spawn-native's privilege/credential handling), not
a generic template;CONTRIBUTING.mddocuments the coverage ratchet and
nix flake checkas the authoritative (not just convenient) verification
step, matching how this project is actually developed and verified
throughout this changelog.
CPS
src/copier.lisp's%drain-copiershad a two-step "join within the
remaining time; if that times out, force-close the stream and give it
one more brief join" escalation inlined as nestedwhens. Extracted
%join-copier-unless-timed-out (copier timeout on-timeout), which calls
theon-timeoutcontinuation only if the join actually times out --
the same shape ascommunicate.lisp's `escalate-unles...