Skip to content

Releases: nerima-lisp/cl-process-kit

v3.4.0

Choose a tag to compare

@github-actions github-actions released this 28 Sep 05:12
v3.4.0
1276a8b

Highlights

  • Added opt-in :use-posix-spawn support for process launches while preserving the existing default behavior.
  • Added directory-aware POSIX spawn launches without changing :search nil semantics.
  • Propagated the option through synchronous, asynchronous, command-spec, and pipeline entry points.

No caller changes are required unless opting into POSIX spawn launches.

v3.3.1

Choose a tag to compare

@github-actions github-actions released this 27 Sep 07:09
v3.3.1
9800ec1

Adds an event-driven async process-completion API and fixes a bug that broke executable builds in v3.3.0, which was not published.

Added

  • communicate-async/process-task: an event-driven counterpart to communicate that returns a process-task immediately instead of blocking. await-process blocks for the result (optionally with a timeout, returning (values nil nil) on expiry), and cancel-process cancels a pending task. Options include event-callback, event-queue-capacity (default 64), event-overflow-policy (:drop-newest by default, or :block), and event-history-capacity.

Fixed

  • v3.3.0 created its two async task executors in top-level defvar initforms, so merely loading the system started twelve threads. SBCL refuses save-lisp-and-die (and so asdf:program-op) and sb-posix:fork while other threads are running, which broke executable builds for every consumer -- this is why v3.3.0 was not published. The executors are now created lazily on first use, and the new exported shutdown-process-kit stops and joins them; it also runs from sb-ext:*save-hooks*, so a saved image carries no thread objects and recreates the pools on demand after restart.

v3.1.0

Choose a tag to compare

@github-actions github-actions released this 02 Aug 07:51
v3.1.0
7454fc1

Added

  • select-fds/wait-for-input: a select(2)-based file-descriptor
    readiness primitive for event loops that need to block on several raw
    descriptors at once -- several PTY masters, a unix socket, and stdin,
    say -- and be told which woke it. This is the one primitive an event
    loop could not previously build out of the rest of the library:
    process-wait/communicate each drive a single child. select-fds
    returns the three ready read/write/exceptional descriptor sub-lists;
    wait-for-input is the common read-only, single-descriptor-list case
    reduced to one argument and one return value. Both retry transparently
    on EINTR against the original deadline, so a periodic signal (
    SIGWINCH, SIGCHLD) never truncates or extends the caller's timeout.
    New public conditions fd-set-overflow (a descriptor at or above
    +maximum-fd+) and fd-wait-failed (any other failing errno) cover
    the two failure modes; every other bound documented on select-fds
    cannot be reached from SB-UNIX:UNIX-FAST-SELECT's own argument checks.
    See the File-Descriptor Readiness guide.

v3.0.1

Choose a tag to compare

@github-actions github-actions released this 01 Aug 17:01
v3.0.1
a7fe23a
v3.0.1: bump cl-codec-kit to v0.3.1, fixing double-replacement on a U…

v3.0.0

Choose a tag to compare

@github-actions github-actions released this 01 Aug 16:16
v3.0.0
4c65751
v3.0.0: migrate UTF-8/octet handling onto cl-codec-kit (breaking: EXT…

v2.0.0

Choose a tag to compare

@github-actions github-actions released this 30 Jul 16:48
v2.0.0
140e8b7

Added

  • process-kit/pty:pty-try-wait/pty-alive-p: non-blocking liveness
    checks for a PTY session, the analogues of process-try-wait/
    process-alive-p. Found the same way as with-pty-process two commits
    earlier -- comparing the PTY subsystem's public surface against the
    main library's and finding pty-wait (blocking, possibly escalating)
    had no non-blocking counterpart. pty-wait's own polling loop now
    calls the public pty-try-wait instead of the internal %try-wait
    directly, matching process-wait's existing call-the-public-function
    convention in process-handle.lisp.

  • process-kit/pty:with-pty-process/call-with-pty-process: the PTY
    analogue of with-process/call-with-process's continuation-passing
    cleanup contract, binding a variable to a spawn-pty result for a body
    and closing it via pty-close on the way out, success or error. The
    main library had this idiom for process-handle; the PTY subsystem
    never got the equivalent, so all 6 of t/pty-test.lisp's cases
    hand-rolled the identical unwind-protect/pty-close shape. Now they
    read as (with-pty-process (process (spawn-pty ...)) ...).

  • cl-weave:it-fuzz property (t/property-test.lisp): run's :replace
    UTF-8 decoding path must never signal on arbitrary octet input, across
    100 generated trials, generalizing the 3 hand-picked malformed sequences
    t/edge-coverage-test.lisp already covered into a genuine property of
    the full byte space. First it-fuzz use in this codebase; it-property,
    it-each, and run-mutations were already adopted in earlier passes.

  • checks.maxLineLength in flake.nix: fails nix flake check if any
    src//t/ line exceeds 100 columns (the org's CODING_STANDARD.md
    rule), matching the existing checks.maxFileLength gate's shape. Added
    after actually closing the gap it enforces, not before: an earlier
    pass's "wrap every line at 100 columns" commit was scoped to src/
    only (its own message says so), and t/ was never swept -- 100 lines
    across 12 test files exceeded 100 columns until this pass fixed them by
    hand, preserving each file's existing hand-crafted formatting rather
    than running a canonical auto-formatter (paredit edit format's
    canonical style turned out to differ enough from this codebase's actual
    conventions -- breaking every keyword-argument pair onto its own line,
    removing blank lines between it cases -- that using it wholesale
    would have been a large, unrelated stylistic churn; used only on the two
    files, t/pty-test.lisp and t/native-spawn-test.lisp, that had no
    existing hand-crafted style to preserve in the first place, both single
    giant crammed one-liners).

Changed

  • Extracted %task-submit-output's and %task-finish's identical "flush
    pending-drops into an :overflow event" block (src/async-events.lisp)
    into a shared %flush-pending-drops-event, via paredit refactor extract-function --at <offset> --infer-params. Raised the coverage
    ratchet floors to match (expression 87.4% -> 87.7%, branch 81.3% ->
    81.4%): the dedup genuinely closed one branch gap (124 -> 123 uncovered),
    not just diluted the denominator.

  • Unified three near-duplicate deadline-polling loops onto the single
    %poll-until primitive that communicate.lisp/pipeline.lisp already
    shared -- found by grepping for the "loop until a predicate is true or a
    deadline passes, sleeping between checks" shape directly, since it
    differs enough call to call (different clock sources, a returned value
    vs. a boolean) that the exact-match paredit inspect duplicates tool
    cannot see it as one shape. %poll-until itself is now defined in
    process-handle.lisp (the earliest-loading file among its callers, so
    both process-group.lisp and communicate.lisp can reach it) and
    generalized to accept a NIL deadline for an unbounded wait:

    • process-handle.lisp's process-wait now delegates to %poll-until
      instead of hand-rolling the same loop; a genuine (not just cosmetic)
      fix falls out of this, since %poll-until caps its final sleep at
      whatever time is actually left before the deadline, where the old loop
      always slept the full poll-interval regardless -- process-wait
      could previously overshoot its deadline by up to poll-interval
      before ever re-checking.
    • process-group.lisp's %wait-until-process-group-gone now delegates
      to the same %poll-until instead of a hand-rolled copy of exactly the
      donep communicate.lisp's own %wait-until-group-gone already used --
      the two are the same "process group is gone" check, previously written
      down twice.
    • %poll-until's own docstring previously claimed "this is the one
      place [every deadline-bounded wait] is written down," which was false
      given the two hand-rolled copies above; corrected, and now explains why
      process-kit/pty's pty-wait deliberately keeps its own copy (a
      genuinely different, three-way donep with no clock/sleeper injection,
      not an oversight).
      Verified via two full nix flake check runs (the first hit the
      documented SETUP-SERVE-EVENT-PIPE flake on await-process, an
      unrelated async-cursor test that shares no code with any of the three
      functions touched here; the second was fully green, 180/180, coverage
      88.0%/82.1%, both above the ratchet floor).
  • Extracted run-pipeline's inline thread-spawning loop/lambda block
    (src/pipeline.lisp) into a named spawn-stage-threads sibling inside
    the same labels form as the pre-existing await-pipeline-stages.
    Found via a fresh size-ranked scan of every src/*.lisp function
    (paredit inspect outline --output json): run-pipeline was the
    codebase's 2nd-largest function and, unlike its neighbors
    communicate-async/communicate, had never had its inline
    thread-spawning lambda pulled into a named function. Hand-authored
    rather than via paredit refactor extract-function --infer-params,
    since that flag mishandles loop's clause-keyword syntax (for/in/
    collect read as ordinary symbols to infer as parameters). run-pipeline
    itself now reads as a linear sequence -- build pipes, spawn stages, close
    write ends, (await-pipeline-stages (spawn-stage-threads)), build
    result -- with no behavior change.

Added

  • Both test entry points (run-tests.lisp, run-pty-tests.lisp) now bind
    cl-weave:*default-timeout-ms* to 30 seconds before running the suite.
    The org-wide TEST_STANDARD.md requires every repository to set this;
    cl-process-kit's suite previously had no PER-TEST ceiling, only a
    whole-suite one (nix flake check's own timeout 180/CI's
    timeout-minutes) -- so a single hung test burned the entire budget
    before failing, with nothing naming which test hung. 30s leaves generous
    headroom above the slowest observed test (~1.1s) while catching a
    genuine deadlock with a specific CL-WEAVE:TEST-TIMEOUT diagnostic.
    Reformatted run-pty-tests.lisp from a single unreadable one-line form
    to match run-tests.lisp's style while making this change.

Changed

  • flake.nix's .asd :version extraction uses cl-nix-forge's dedicated
    fromAsdSystem lexer instead of a hand-rolled line-by-line regex.
    Strictly stronger, not just shorter: cl-process-kit.asd declares four
    systems sharing one version (cl-process-kit, /test, /pty,
    /pty-test), and the old regex read whichever :version line happened
    to come first with no cross-check, while fromAsdSystem fails the build
    loudly if any of them ever disagreed.
  • flake.nix's packages.*/checks.checkout-tests/checks.pty-tests
    now build on cl-nix-forge's lispDerivation/mkScriptCheck primitives
    (new flake input, pinned v0.4.0) instead of pkgs.sbcl.buildASDFSystem
    plus hand-rolled pkgs.runCommand test derivations. noForbiddenMarkers,
    maxFileLength, formatting, and docs are untouched -- none of them
    involve ASDF, so cl-nix-forge has nothing to offer there. Two things
    cl-process-kit's own code needed that neither of the org's two existing
    adopters (cl-weave, cl-json-kit) demonstrated: src/pty.lisp reads
    CL_PROCESS_KIT_PTY_LIBRARY explicitly at load time (an SB-ALIEN LOAD-SHARED-OBJECT call needs a real path, not a bare SONAME resolved
    off nativeLibraries' LD_LIBRARY_PATH/DYLD_LIBRARY_PATH propagation
    alone), so that env var is still set directly alongside nativeLibraries;
    and native/spawn.c's trampoline binary is compiled twice on purpose --
    once into $out/bin for the distributed package (postInstall), once
    into the build sandbox for checks.checkout-tests (preCheck), since a
    check needs the binary to exist before its own build/install phases (what
    actually produces the packaged copy) have run. apps/devShells stay on
    the original hand-rolled CL_SOURCE_REGISTRY string, not
    cl-nix-forge:lispScript -- this migration's own research did not verify
    that primitive's exact parameter shape the way lispDerivation/
    mkScriptCheck were verified against real builds, and guessing an
    unconfirmed API for a part of the flake with no correctness benefit over
    the proven string was not worth the risk.
    Verified with nix build on each new package/check individually before
    the full nix flake check (179/179 + 6/6 PTY, coverage unchanged at
    87.8%/81.5%, both native artifacts confirmed present and loadable).

Changed (BREAKING)

  • next-process-event now returns a single process-event-step struct
    (process-event-step-event/-cursor/-status/-gap-count) instead of
    four raw values. The org-wide API_STANDARD.md names this exact function,
    by name, as the canonical bad example of returning more than two values
    ((values a b) at most, or a struct beyond that) -- a caller that only
    wanted gap-count had to destructure all four positionally. Requires a
    major version bump at the next...
Read more

v1.0.1

Choose a tag to compare

@github-actions github-actions released this 25 Jul 20:41
v1.0.1
359d127

A packaging fix. No source change, and the exported API is identical to 1.0.0.

Fixed

  • flake.nix asked for cl-log-kit/v1.6.0, a tag that has never existed —
    cl-log-kit has only ever released v1.0.0. The v1.0.0 tag of this
    repository still carries that reference, so anyone pinning
    cl-process-kit/v1.0.0 as a flake fails at nix flake lock from a cold
    cache. Existing lock files hide it by holding an already-resolved revision,
    which is why CI stayed green. Three repositories pin this one that
    way — cl-boundary-kit, cl-cc-runtime and cl-cli — and would all have
    broken at once on the next flake-update.yml run.

    Tags are never moved, so v1.0.0 stays as it is and this release exists to
    give downstream something that resolves. Sibling versions are now read out of
    each pinned source's own .asd rather than repeated here, so the two cannot
    drift apart again.

v1.0.0

Choose a tag to compare

@takeokunn takeokunn released this 25 Jul 16:59
v1.0.0
9f305d6

First stable release. The exported API is unchanged from 0.2.0 and is now
covered by semantic versioning; what this release adds is evidence about how
it behaves on the platforms it claims to support, gathered by actually
running the suite on Linux rather than inferring from macOS. That turned up
two real defects, both invisible on macOS, and one broken CI check.

The first of the two ### Known Limitations recorded under 0.1.0 -- the
drain-timeout bound -- is resolved. The second, a group of timing-sensitive
process-group tests, is not: it is narrowed, re-diagnosed, and still skipped
on Linux.

Correctness

  • drain-timeout-seconds is now honoured on Linux. A copier thread
    drained a child's stdout/stderr with a blocking read(2), and
    %drain-copiers unstuck one that overran its deadline by force-closing
    the stream under it. Closing a descriptor wakes a thread already parked
    in read(2) on macOS/BSD, but that is a BSD courtesy rather than
    anything POSIX promises, and Linux does not do it: the reader stayed
    parked until whoever else held the pipe's write end let go. After
    sh -c "sleep 5 & exit 0" that is a backgrounded descendant which
    outlives the leader by design, so the wait was effectively unbounded.
    The observed symptom on Linux was worse than 0.1.0's note described --
    not merely run overrunning its documented bound, but run signalling
    process-io-error :cleanup ("Copier thread did not terminate after its
    stream was closed") once the force-close failed to land.

    The read loop now waits with poll(2) on a bounded timeout and checks a
    stop flag between turns (%await-fd-readable), so the decision to give up
    is taken by the reader itself instead of being inflicted on it through the
    descriptor. The bound then holds by construction on any POSIX platform
    rather than by accident on some. The flag is checked before readiness,
    which bounds the loop even against a child that never stops producing:
    one poll interval plus one bounded read(2) per turn. Polling costs no
    wakeup rate the library was not already paying, since communicate's own
    deadline loop already runs at +default-poll-interval+.

    %drain-copiers' escalation gained the cooperative stop as its first
    rung -- ask, then force-close, then signal -- ordered by what each costs
    when it fires, in the same continuation-passing shape as
    escalate-unless-gone's SIGTERM -> SIGKILL -> give-up. Force-closing is
    now a fallback for a copier parked where a flag cannot reach it (a
    read-sequence on a non-fd stream) rather than the primary mechanism.
    %communicate-base's cleanup path was reordered to match: it retires the
    copiers before their streams, so close-process-streams no longer closes
    a descriptor under a live reader on every pass -- a descriptor the OS is
    free to reissue the moment it is closed.

  • :on-timeout and :on-cancel are validated at the entry points that
    own them.
    Every entry point resolves these by comparing against
    :error and treating anything else as :return, which made an
    unrecognised value indistinguishable from a deliberate :return instead
    of an error. run, run-command and run-pipeline compounded it: each
    hands communicate a hardcoded :on-timeout :return and decides for
    itself whether to signal, so %validate-communication-options' guard
    never saw what the caller wrote. A misspelt :errror therefore read as
    :return and silently swallowed the very timeout or cancellation the
    caller had asked to have signalled. %validate-outcome-policy now guards
    each policy where it is accepted -- run (:on-timeout, :on-cancel),
    run-command (both), run-pipeline (both), and communicate
    (:on-cancel, which had no guard either).

Testing

  • The drain-timeout regression test ("run returns boundedly when an exited
    leader leaves a pipe-holding descendant") is no longer skipped on Linux --
    the fix above makes it platform-independent, and it passes on CI.

  • The other seven Linux skips are re-diagnosed rather than removed. They had
    been filed under the same "process-group/communicate timing is not
    guaranteed identical on Linux" heading as the drain bug, which conflated
    two unrelated things. They do not share its cause: each asserts that a
    process group is gone within a 0.1s grace period, which a contended
    shared CI runner cannot reliably deliver. (They pass on an uncontended
    aarch64 Linux container and fail on GitHub's x86_64 runners -- contention,
    not architecture.) Their skip reasons now say that instead. They are not
    given more headroom because a timing assertion loose enough to survive
    arbitrary contention no longer asserts the timing; the honest fix is to
    make the assertion event-driven rather than deadline-driven, which is
    deferred.

  • The coverage ratchet had been failing every Linux CI build, on a
    comparison it should never have made. The floor tracked the figure from a
    complete run (macOS, nothing skipped) but was enforced against the reduced
    Linux run, so CI reported the skipped tests as a coverage "regression" --
    86.9% against an 87.0% floor -- with every test passing. The floors are now
    enforced only when the whole suite ran (+suite-complete-p+); otherwise
    coverage is reported with an explicit note that it is not comparable.

  • A guard-clause test that names a program which does not exist proves
    nothing.
    t/run-timeout-test.lisp's "run rejects invalid timeout
    controls before spawning" spawned /bin/true, which macOS does not ship
    (true lives in /usr/bin there). All ten of its assertions passed on a
    process-launch-error for the missing file -- identically, and just as
    green, whether or not the guard under test existed. That is what hid the
    :on-timeout gap above: the assertion only had to mean something once
    the suite was first run on Linux, where /bin/true does exist. A
    %true-program fixture now resolves true through the ambient PATH,
    matching the existing %spawn-sleeping fixture, and is used by the 17
    call sites that actually spawn. The remaining literal /bin/true
    occurrences are in make-command data tables that never spawn, where the
    string is inert.

  • Added a regression test asserting that all four entry points reject an
    unrecognised outcome policy rather than reading it as :return.

  • Coverage ratchet advanced to 87.4% expression / 81.5% branch (from
    87.0/79.5), measured on a complete run.

Documentation

  • Documented the timeout and cleanup deadlines. grace-period,
    poll-interval, timeout-signal, kill-signal and
    drain-timeout-seconds previously appeared only inside run's signature
    in the options table -- five knobs with no stated meaning or default,
    covering the escalation behaviour that is the library's whole reason to
    exist. The execution guide now gives them a table of their own, explains
    why signals go to the process group rather than the child, and explains
    why draining has a deadline separate from the child's (a descendant that
    outlives the leader inherits the same pipe and can hold it open
    indefinitely).

  • Added a "Running the suite on both platforms" section to the development
    guide, covering the container invocation for checking Linux behaviour
    locally, and the two habits this release's bugs argue for: never name a
    program a guard-clause test does not intend to execute, and prefer fixing
    a platform difference to skipping the test that catches it.

Build/environment

  • flake.nix hardcoded "0.2.0" in four separate places (the docs
    derivation, the cl-process-kit package, and the cl-process-kit-pty
    package), plus two more in cl-process-kit.asd (cl-process-kit and
    cl-process-kit/test) -- six places a release has to remember to bump in
    lockstep, with no build-time check that they agree. nerima-lisp/cl-boundary-kit
    v0.6.0 already fixed the identical problem in its own flake.nix (found
    while auditing this project's own environment setup for the same class
    of drift); ported the same technique here: a version let binding
    parses the :version form out of cl-process-kit.asd line-by-line
    (Nix's builtins.match is whole-string-anchored and . doesn't span
    newlines, so a single multi-line regex doesn't work) and every Nix
    package now inherits it. A release now only ever edits the .asd.
  • cl-process-kit/pty and cl-process-kit/pty-test were missing the
    :version/:author/:maintainer/:license/:homepage/:bug-tracker/
    :source-control metadata the other two systems in the same file
    already carry -- added it for consistency.

Production readiness

  • Added SECURITY.md, SUPPORT.md, and CONTRIBUTING.md (GitHub's
    standard community-health filenames, which its UI surfaces automatically
    regardless of README content) -- this repository had none, unlike sibling
    nerima-lisp projects (cl-weave has all three). Scoped and sized for
    this project specifically rather than copied verbatim: SECURITY.md
    names the concrete classes of report that actually apply here (shell/
    argument injection via run-shell/make-command, process-group
    isolation failures, spawn-native's privilege/credential handling), not
    a generic template; CONTRIBUTING.md documents the coverage ratchet and
    nix flake check as the authoritative (not just convenient) verification
    step, matching how this project is actually developed and verified
    throughout this changelog.

CPS

  • src/copier.lisp's %drain-copiers had a two-step "join within the
    remaining time; if that times out, force-close the stream and give it
    one more brief join" escalation inlined as nested whens. Extracted
    %join-copier-unless-timed-out (copier timeout on-timeout), which calls
    the on-timeout continuation only if the join actually times out --
    the same shape as communicate.lisp's `escalate-unles...
Read more