Repository navigation
Version 0.3.0
[0.3.0] - 2026-09-08
Added
- Lineage-backed DAG reconstruction (
--diff-dag) — when a run's project
directory has a Nextflow data-lineage store (.lineage/, produced with
lineage.enabled = trueon Nextflow 25.04+), the process wiring is now read
from the authoritative provenance Nextflow persisted instead of being
inferred from work-dir symlinks. A newLineageStorereads each.data.json
record directly off disk (no compile-time dependency on thenf-lineage
module), indexes the run'sTaskRunrecords by their session id, and
reconstructs producer→consumer edges from each task's recordedinputLID
references (lid://<producerTaskHash>/…). Because it reads what Nextflow
recorded, this needs no work directories and is unaffected by cleanup.
DagComparator.graphOfnow takes the run's project directory and prefers the
lineage store, falling back to the existing best-effort symlink
reconstruction (symlinkGraphOf) when no lineage store recorded the run. Each
run'sRunGraphcarries asource(LINEAGE/SYMLINK/NONE), and the
wiring layer's note now states whether the graph is authoritative (lineage) or
best-effort (symlinks), including the mixed case. The layer remains
informational only — it never affects the "identical" verdict or
--fail-on-change. Only the default<projectDir>/.lineagestore location is
auto-detected; a customlineage.store.locationstill falls back to symlinks.
Changed
-
Unknown plugin command now exits with the usage code (2), not 1. An
unrecognized verb (anything other thandiff) is a usage error, in the same
class as bad arguments, so it now returns2— matching the documented
exit-code table — instead of1(which is reserved for runtime errors). The
message also notes that onlydiffis supported. -
Help summary lists all current diff layers. The one-line description shown
by-h/--helpstill read "metadata, processes, and per-task
resources/scripts" from the 0.1.0 days; it now enumerates the always-on layers
(parameters, configuration, software & versions, failure rollup, performance
regressions, resource-efficiency) and the three opt-in flags. -
--diff-dagno longer requires work directories when lineage is enabled.
Previously the wiring layer always needed the tasks' work directories to still
exist locally; with a lineage store present it is reconstructed from persisted
provenance instead. -
Continuous integration & tag-based releases — a GitHub Actions CI
workflow (.github/workflows/ci.yml) now runs the full verification suite
(make check) on every push and pull request tomain, across JDK 17 and 21,
uploading test reports as build artifacts. A companion release workflow
(.github/workflows/release.yml) publishes to the
Nextflow plugin registry when av*version
tag is pushed: it verifies the tag matchesbuild.gradle'sversion(so a tag
can never publish a mismatched artifact), runsmake check, thenmake release, authenticating with anNPR_API_KEYrepository secret. See the
README's "Continuous integration" and "Releasing" sections for setup and the
tagging flow. -
DAG (process wiring) diff (
--diff-dag) — a new opt-in layer that
reconstructs each run's process;process wiring and diffs the two edge
sets, so nf-diff surfaces topology changes the task-count-per-process view
cannot see — e.g. a pipeline rewired fromA → CtoA → B → C. Nextflow
does not persist DAG edges in its history or cache, so there is no
authoritative edge list to read; what it does leave on disk is every task's
staged inputs, materialised as symbolic links inside the task's work
directory.DagComparatorwalks each task's work dir, resolves every input
symlink, and attributes any target that resolves into another task's work dir
(walking the parent chain so a link into a nested output subdir still
attributes to the producer) as a producer;consumer edge; links that
resolve outside every work dir are external inputs and yield no edge. Because
it walks work directories, this layer needs them to still exist locally (like
--diff-outputs/--diff-logs) and is a best-effort reconstruction: if some
work dirs were cleaned up, the recovered wiring is incomplete, and a note
reports how many task work dirs were missing so a partial diff is not read as
authoritative. Surfaced in all three report formats (HTML "Process wiring
(DAG)" section + nav link, Markdown section, and adagblock with
dagEdgesAdded/dagEdgesRemovedin JSON). Because the reconstruction is
best-effort, this layer is informational only and never affects the
"identical" verdict or--fail-on-change. -
Failure rollup (top-level "what failed and why") — a new always-on layer
that answers, at a glance, which tasks failed and why, instead of leaving that
scattered across per-task detail. Failed tasks are detected from the cached
status/exitfields (an explicitFAILED/ABORTEDstatus, or a non-zero
exit code — theNO_EXITsentinel and blanks are ignored), so the layer reads
straight from the run cache and needs no work directories. Failures are rolled
up by their(process, status, exit)signature and counted per run, sorted by
biggest blast radius first; a signature seen only in Run B is flagged new
(a regression), one present in Run A but gone in Run B is resolved, and one
in both is persistent. The layer also surfaces a run-level error state
(historystatusstartingERRor equal toFAILED/ABORTED/KILLED) even
when no individual task failure was recorded. Surfaced in all three report
formats (HTML "Failure rollup" section + nav link + summary cards, Markdown
section, and afailuresblock withfailedA/failedB/newFailures/
resolvedFailuressummary counts in JSON). Because the meaningful identity
signal — a task whose status or exit changed — is already carried by the task
field diffs, this rollup is informational only and never separately affects
isIdentical()/--fail-on-change. -
Resource-efficiency layer (requested vs. measured-peak provisioning) — a
new always-on layer that, per process, compares what each run requested
(cpus,memory) against what it actually peaked at (%cpu,peak_rss).
Both requested and peak values are read straight from the run cache trace and
taken as the max across a process's tasks (a retried task that used more, or
was bumped a higher request, is the honest worst case), so the layer needs no
work directories and is always computed. The efficiency ratio is
measured-peak / requested; each process is classified per run as over
(below 50% — wasted allocation, e.g. "requested 32 GB, peaked at 4 GB"),
tight (90%+ — risk of OOM kills or CPU throttling), or ok in between.
Surfaced in all three report formats (HTML section + an "Over-provisioned (B)"
summary card, Markdown table, and anefficiencyarray plus
overProvisionedA/overProvisionedBandtightA/tightBsummary counts in
JSON). Because provisioning is a tuning signal rather than a correctness
change, this layer is informational only — it never affectsisIdentical()/
--fail-on-change. -
Cross-project comparison (
--dir-a/--dir-b) — the two runs no longer
have to live in the same project. Previously a single--dirresolved both
runs'.nextflow/history, cache, config and params, so you could not compare
"the same pipeline in two checkouts" (or on two machines).--dir-a=<dir>and
--dir-b=<dir>now set each run's project directory independently; each falls
back to--dirwhen omitted, so existing invocations are unchanged. Run A is
loaded from and resolved againstdir-a, run B againstdir-b: the parameters
layer reads each run's own-params-file, and the configuration layer rebuilds
each run's effectivenextflow.configfrom its own working tree, so a
-profile dockerin project A is diffed against project B's config. The
git-provenance caveat became per-tree:ConfigProvenancenow carries a
crossProjectflag plus each side's directory, current HEAD and dirty state,
and its warning describes the two working trees separately (currentRevisionB,
dirtyB,dirA,dirBare surfaced in the JSON report).--laststill needs
a single history, so it is rejected when combined with differing
--dir-a/--dir-b.
Fixed
- Diff errors with no message printed a blank line.
DiffPlugin.exec()
reported a caught throwable viae.messageonly, so a message-less exception
(notablyNullPointerException) produced a barenf-diff:line with nothing
after it, while the stack trace went only to the debug-gated log. It now falls
back to the exception's simple class name, so both the stderr line and the log
always name the failure. - Cache-lock retry backoff was not interruptible. The backoff between
attempts to open a contended run cache used Groovy'ssleep(), which swallows
InterruptedExceptionand clears the interrupt flag, so aCtrl-Cduring a
contended open was ignored and the loop kept retrying. It now uses
Thread.sleep(), restoring the interrupt flag and aborting the retry on
interruption. GitProvenancesubprocess timeout was ineffective — the git subprocess's
stdout/stderr were read inline withgetText()before the timedwaitFor,
which blocks until the process exits, so a hunggitcould never be timed out.
Both streams are now drained on background threads started beforewaitFor, so
the 5s timeout actually fires and a chatty command cannot deadlock on a full
pipe buffer.