Releases: mribeirodantas/nf-diff
Release list
Version 0.9.1
[0.9.1] - 2026-09-21
Fixed
- Bump Gradle plugin version. This way, the documentation will show up at https://registry.nextflow.io/plugins/nf-diff@0.9.1
Full Changelog: v0.9.0...v0.9.1
Version 0.9.0
[0.9.0] - 2026-09-21
Added
- Central run archive for comparing runs after their project directories are
gone. When a pipeline opts in withdiff.archive.enabled = true, an
end-of-run trace observer snapshots each finished run into a shared archive
directory, anddifftransparently falls back to that archive whenever a
requested run is no longer in the local.nextflow/history. The observer is
off by default and inert unless enabled, so existing users see no behaviour
change.mode = 'lightweight'(default) stores run and per-task metadata;
mode = 'complete'also copies each task's output and.command.*log files
(staged input symlinks are skipped). The archive location resolves from
diff.archive.dir/--archive-dir, thenNXF_DIFF_ARCHIVE_DIR, then the
default$NXF_HOME/nf-diff/archive. Archiving is best-effort and never fails
the pipeline run.
Fixed
- File-content comparison now decides identity on the full SHA-256 digest,
not a truncated 12-character prefix. Two same-size files that differed only
beyond the first 48 bits of their digest could previously collide and be
misreported asUNCHANGED— and such a change would not count toward
--fail-on-change. The equality decision now compares full digests; the
12-character prefix is retained only for compact display in the report, so
HTML/JSON output is unchanged. Affects--diff-outputsand the
--published-a/--published-bcomparison.
Documentation
- Documented the central run archive (
diff.archive.*config, the
--archive-diroption,NXF_DIFF_ARCHIVE_DIR, and the lightweight/complete
modes) in the usage guide and website docs. - Documented that archiving requires the plugin to be active for the run.
The archive section now shows theplugins { id 'nf-diff@<version>' }
declaration (or-plugins nf-diff@<version>) needed so the end-of-run
observer is loaded —diff.archive.enabled = truealone has no effect unless
the pipeline loads the plugin.
Full Changelog: v0.8.0...v0.9.0
Version 0.8.0
[0.8.0] - 2026-09-18
Added
- The GitHub Action now enables
diff-allby default. Action users get
task output, log, and DAG-wiring comparisons without adding the flag;
settingdiff-all: 'false'restores the opt-out behavior. - The GitHub Action's PR comment now falls back to a short summary when the
full report would exceed GitHub's 65,536-character comment limit. The
fallback keeps the summary/runs tables and links to the uploaded HTML
report artifact for the complete diff, instead of failing to post.
Documentation
- GitHub Action examples now link to a working demo pull request in the
demo-nf-pipelinerepository, making the CI integration easier to try and
verify. - GitHub Action examples now enable Nextflow data lineage for compared runs.
The cache-backed examples pass a sharedlineage.configfile to both
pipeline runs and cache the.lineagestore alongside.nextflow. The
dual-run example enables lineage for both directories, and the documentation
explains that lineage must be enabled before execution for authoritative DAG
and runtime metadata comparison. - Documented that Nextflow lineage is optional. When no
.lineagestore is
available,nf-difffalls back to work-directory symlinks where needed; the
documentation also explains when to enable lineage and that cache-backed
workflows should preserve.lineagealongside.nextflow.
Full Changelog: v0.7.0...v0.8.0
Version 0.7.0
[0.7.0] - 2026-09-12
Added
-
A reusable composite GitHub Action (
action.yml) that wraps the
nf-diff:diffverb for CI. It resolves the plugin from
registry.nextflow.io(no build orplugins {}config needed), optionally
sets up a JDK and Nextflow, runs the comparison against a project's existing
.nextflow/history, uploads the report as an artifact, and can post/edit a
Markdown pull-request comment. Typed inputs surface the common flags
(last,run-a/run-b,format,fail-on-change,only/exclude,
perf-threshold, thediff-*work-dir layers,verbose,dir,
dir-a/dir-b, andpublished-a/published-b) with anextra-args
passthrough, and it exposesreport-path,html-report-path,artifact-url,
artifact-id,exit-code, andidenticaloutputs. When commenting on PRs
with Markdown, it automatically generates and uploads the standalone HTML
report dashboard and prepends a direct download link to the PR comment.
The action honours the plugin's launcher contract —
bare id (nf-diff:diff, never pinned) and inline--output=— and maps its
exit codes so1/2fail immediately while3(runs differ +
--fail-on-change) is deferred until after the report is uploaded and
commented. -
A dogfood workflow (
.github/workflows/nf-diff.yml) that generates two
real runs from a trivial demo pipeline and compares them with the local
action (uses: ./) on pull requests, exercising the action wiring
end-to-end against the published plugin. -
An
actionlintjob in CI that lints every workflow (and shellchecks the
embedded run-step scripts) via rhysd's official install script. -
The HTML report's "Resolved configuration" section now offers a git-style
diff view alongside the table. The section is now a tabset with three views
of the same resolved config: Table (the existing key / Run A / Run B
grid, still the default), Full diff (every key askey = valuelines,
unchanged keys shown as context and changed keys shown as a-Run A /+
Run B pair — keys present on only one side render a lone-/+), and
Changes (the same diff with the unchanged context lines dropped). The diff
reuses the report's existing.clline-diff styling, so it inherits the
colorblind-safe palette. The DAG section's tab machinery was generalized from
.dag-tab*/data-dagto shared.tabset/.tab/.tab-panel/
data-tabclasses (with a single generic tab-switching script) so both
sections drive off the same code.HtmlReportRendereronly; the Markdown /
JSON / terminal renderers are unchanged. -
A new published-outputs layer compares two runs' result directories
directly. Pass--published-a=<dir>and--published-b=<dir>to diff each
run'soutdir/publishDirtree instead of — or alongside — the per-task
work-dir outputs (--diff-outputs). Files are keyed by their path relative to
each published root and classified by size, then SHA-256 for same-size files,
then a line-level diff for changed text files, reusing the same
FileContentComparatorengine as the work-dir layer. Symlinks are followed,
so it works whetherpublishDircopied or symlinked. Unlike--diff-outputs,
this reads the durable published results, so it still works after the work
directories are gone (cleaned up, or on remote object storage). A published
file change counts as a difference forisIdentical()/--fail-on-change.
Implemented byPublishedComparator→DiffResult.PublishedDiff, wired
throughCompareOptions,DiffCommand, andRunComparator, and rendered by
all four output formats (HTML, JSON, Markdown, terminal).
Fixed
- CommandParams.parse swallowed negative-number values (CommandParams.groovy)
A flag whose value was a negative number (--min_log2fc -1.5, --seed -42, --scale -1e-3) was misread — the parser saw the leading -, treated the number as another option, and recorded the flag as boolean true. Now isNumericValue() recognizes signed decimals/integers/scientific notation as legitimate values. Impact: high — silently wrong parameter diffs for any pipeline using signed numeric params. - DiffResult.failedExit counted the NO_EXIT sentinel as a failure (DiffResult.groovy)
Nextflow writes Integer.MAX_VALUE to a task's exit field when it never produced one. The old check only excluded 0, so 2147483647 was treated as a non-zero failure. Now it excludes the sentinel explicitly. Impact: high — false-positive failures in the log/failure layers. - Format.pctDelta dropped zero-baseline regressions (Format.groovy)
A 0 → N change (a metric appearing from nothing) returned null and was silently discarded instead of ranking as a regression. Now 0→0 is 0.0 and 0→N is ±Infinity, with signedPct rendering it as ±∞%. Verified the threshold check and worst-first sort in RunComparator handle infinity correctly. Impact: medium. - OutputComparator skipped nested files sharing a control-file name (OutputComparator.groovy)
Control files (.command.sh, etc.) were matched by leaf name anywhere in the tree, so a genuine output like results/.command.sh was dropped. Now only work-dir-root entries (nameCount == 1) are skipped. Impact: low–medium. - ConfigLoader.configFilesFrom could consume a following option as the -c path (ConfigLoader.groovy)
A bare -c followed by another flag mis-consumed it, which could drop a real later -c config. Now it only consumes a non-- token. Impact: low. - Duplicate java.nio.file.Path import (DagComparator.groovy) — cosmetic cleanup.
Each fix has a dedicated regression test (new DiffResultTest.groovy plus cases added to FormatTest, CommandParamsTest, ConfigLoaderTest, OutputComparatorTest), and the full suite passes. Nothing is committed — the changes are staged in your working tree for you to review.
Full Changelog: v0.6.0...v0.7.0
nf-diff GitHub Action 0.7.0
Full Changelog: v0.7.0...ga-0.7.0
Version 0.6.0
[0.6.0] - 2026-09-10
Added
-
The report's vertical section nav now flags which sections hold
differences. Each.sidenaventry whose section carries a change gets a
trailing warning icon (.nav-alert, colored with the shared--changed
token), so the reader can see at a glance where the runs diverge without
opening every page. The flag reuses each section's own "changed" predicate
and respects the verbose (--verbose) view for the metadata / parameters /
configuration layers; purely informational sections (Summary, Efficiency)
never alert.renderNavnow routes every link through anavLinkhelper. -
The report header now shows each run's
main.nfpath and puts the verdict
on the heading row. When a run's.lineage/store recorded the main
script's absolute path (scriptFile), it is threaded through
LineageStore.RunEnv→RunSnapshot.pipelinePathand rendered dimmed
beneath the pipeline name in the run chip (and as a hover tooltip), so the
reader can see exactly which file ran. The "These runs differ" / "identical"
verdict badge moved from a separate band below the heading into a flex row
(.hero-title) beside theRun comparisontitle, so it reads immediately
without spending extra vertical space.
Changed
-
Under
--verbose, a flagged performance regression now counts as a
difference. In the default meaningful-changes view the performance-
regressions layer stays informational (it is derived from always-changing
numeric metrics), but the verbose view promotes always-changing fields to
meaningful — soDiffResult.isIdentical()now also requires
regressionCount() == 0whenshowObviousis set, keeping the layer
consistent with the rest of the verbose verdict. -
The HTML report's cards get a softer, uniform surface and lose their
colored accent bars. The flat 1px-outline treatment is replaced by shared
CSS tokens —--radius(14px), a near-invisible--hairborder, and layered
--elev/--elev-hovershadows (defined for both themes) — applied across
every card surface (.card,.run-chip,.summary-headline,.disp,
.rp-plot,.dag-graph,.task, tables,.warn-note); interactive cards
gain a subtle hover lift. The status/run colored accent borders are removed
in favour of the color cues already present elsewhere: the active section-nav
item is now a solid brand pill (not a soft fill with a left bar); task cards
drop their 4px colored left bar (the header pill already states status);
summary stat cards drop the inset bottom-bar shadow (the status-colored
numeral stays); and the warning note drops its left bar (amber tint kept).
Run cards drop the.run-chip::beforeleft bar and instead render the run
name as an inline pill badge colored by run — green (Run A) / blue (Run B).
Per-line diff gutters and the DAG tab underline are kept as line/tab markers.
HtmlReportRenderer(CSS +runChip) and the regenerated demo report only;
noDiffResultaccessors or comparison logic changed. -
The HTML report is paginated instead of long-scroll. Sections render one
at a time (.sectiondefaults to hidden,.is-activereveals it) inside a new
.layoutwrapper. A bottom pager (renderPager) steps through the visible nav
entries, labelling its Previous/Next buttons from the adjacent sections and
disabling at the ends. The page JS drives selection from the nav, keeps the URL
hash in sync viahistory.replaceState, responds tohashchange, and scrolls
to top on each page change, so deep links to a section anchor still land on the
right page. -
The report navigation is now a vertical sidebar. The sticky horizontal
.tabsstrip is replaced by a.sidenavcolumn (sticky, own scroll, active
item marked with a left border). Below 820px it collapses back to a horizontal
scrolling strip so narrow screens keep their content width. The old scroll-spy
is gone — active state follows the shown page. -
Run status now uses Nextflow / Seqera Platform vocabulary everywhere. The
terseOK/ERRhistory-file tokens were still surfaced raw in the Metadata
table's Status row, even though the header pill already mapped them to
SUCCEEDED/FAILED. The mapping now lives once on
RunSnapshot.statusLabel();RunComparator.compareMetadata()and
HtmlReportRendererboth delegate to it, so the reader is never left
guessing what "OK" means and the two views can't drift apart. Resource-efficiency
pills are relabelled for the same reason:over→over-provisionedand the
neutralok→right-sized(class names and colours unchanged). -
The Performance regressions section now leads with a diverging-bar plot.
A newHtmlReportRenderer.regressionPlot()emits a self-contained inline-SVG
chart — one row per flagged metric (worst first, as the comparator already
orders them), bars growing right for regressions (Run B slower/heavier, red)
and left for improvements (green), scaled to the largest absolute delta, with
a zero axis, per-bar signed-percentage labels, dashed outlines on same-work
rows (identical cache hash), and a legend. Like the DAG diagram it needs no
JavaScript or external assets. The existing table stays beneath as the exact
A/B detail and large-list fallback. -
The process-wiring section gains per-run DAG views. Alongside the union
"Changes" diagram, new "Run A (before)" / "Run B (after)" tabs project the
union DAG down to each run's own edges (runEdges) and render them neutrally
(dagSvggained adiffLegendflag;dagPanelhandles the empty case). A
small self-contained tab script switches panels. -
Run chips now show a facts list. Each run chip gained a Started timestamp
(via a newFormat.datetime()) and, when a.lineage/store recorded it, the
run's Nextflow version. The Nextflow row is omitted rather than padded with a
placeholder when the version is unknown.
Fixed
- Two independent runs of the same pipeline are no longer reported as
differing solely because of their task cache hashes. Nextflow folds the
per-run session UUID into every task's cache hash, so two independent runs
always compute a different hash for every task even when the script, inputs
and container are byte-identical. The taskhashfield was compared but not
marked "obvious", so each matched task was flaggedchangedand the verdict
flipped to "These runs differ" — e.g. comparing two plainnextflow run helloruns.hashis now part ofRunComparator.OBVIOUS_TASK_FIELDS, so a
bare hash change is shown for context (and still surfaced as the recompute
count) but never flips the "identical" verdict or trips--fail-on-change
on its own. Under--verboseit is flagged like any other obvious field.
Full Changelog: v0.5.0...v0.6.0
Version 0.5.0
[0.5.0] - 2026-09-10
Added
-
Run metadata now compares the Nextflow version and runtime environment.
A run's Nextflow version can change between two runs and silently explain a
behaviour difference, but it was never surfaced.RunLoadernow reads the
run's data-lineageWorkflowRunrecord (via a new
LineageStore.environmentForSession()) and populatesRunSnapshotwith the
Nextflowversionandbuild, the container engine, and whether Wave and
Fusion were enabled.RunComparator.compareMetadata()adds these as metadata
rows (version/engine/Wave/Fusion are meaningful changes; the build number is
treated as context viaOBVIOUS_METADATA), so they render in every report
through the existing metadata table with no renderer-specific plumbing. Rows
appear only when a run recorded the value, so runs without a.lineage/store
(lineage.enabled=true, Nextflow 25.04+) are not padded with blanks.
Per-run plugin versions are deliberately not compared — Nextflow does not
persist them in the lineage store, the history file, or the task cache — and
the HTML metadata section now says so explicitly. -
The task layer now compares I/O counters and reports execution hardware.
RunComparator.TASK_FIELDSgained the disk-I/O counters (read_bytes,
write_bytes,syscr,syscw,vol_ctxt,inv_ctxt) and the
execution-environment fields (cpu_model,hostname,native_id). These are
already present in everyTraceRecordNextflow writes to the cache, so no new
data source is needed — the report simply stopped throwing them away. All nine
are added toOBVIOUS_TASK_FIELDS, so they surface for context (notably,
cpu_model/hostnameexplain a performance regression the perf layer already
flags — run B's task landing on a slower CPU) without flipping the "identical"
verdict or tripping--fail-on-changeunless--verboseis set.
Changed
-
The Parameters note now explains where non-launch params go. The HTML
report's params section only ever lists values from the launch command and
-params-file; params left at their defaults or set insidenextflow.config
/ an activated profile were silently absent, which read as "unset". The note
now states those are resolved config, not launch input, and links to the
Configuration layer where they actually appear. -
--helpnow documents the full exit-code contract.DiffPlugin.dispatch()
maps outcomes to four exit codes —0success,1runtime error,2usage
error,3--fail-on-changeon a difference — butusage()only mentioned
3(buried in the--fail-on-changeentry). A CI author reading--helphad
no way to tell2("I typed the command wrong") from1("the diff itself
failed"). A newExit codes:block inusage()spells out all four; doc-only,
no behaviour change. -
The process-wiring section now leads with a node-link diagram of the DAG,
not just a table of changed edges.renderDag()already had the full union
of process→process edges taggedUNCHANGED/ADDED/REMOVED
(RunComparatorbuilds it), but the report threw the unchanged edges away and
listed only added/removed rows — so the reader never saw where in the
topology a change sat. A newdagSvg()helper lays the union graph out with a
lightweight longest-path (Kahn) layering — columns = topological depth — and
emits a self-contained inline SVG: unchanged edges are neutral hairlines,
added edges solid green, removed edges dashed red, and a process appearing in
only one run gets a matching node outline. The layout is computed in Groovy so
the SVG needs no JavaScript or external assets (preserving the report's
no-network-assets guarantee), and the existing per-edge table is kept beneath
it as the precise detail and large-graph fallback. On the rich-report demo the
ALIGN→QCedge is now visibly rerouted through the newly insertedMARKDUP
node. OnlyHtmlReportRendererand its test changed; noDiffResult
accessors or other renderers were touched.
Changed
- HTML report summary reorganised from a flat wall of boxes into a scannable
hierarchy. The summary section previously rendered up to 14 identical,
equal-weightstatCards in a singleauto-fitgrid — the reader had to read
every box one by one to find the answer, with no cue that the first five were
a single task distribution and the rest were per-layer change counts.
renderSummary()now leads with a headline number (total task-level
differences, coloured by the identical/different verdict), collapses the four
mutually-exclusive task buckets (Changed / Only in A / Only in B / Unchanged)
into one stacked proportion bar with a counted legend — withRecomputed
demoted to an annotation since it is a cross-cut ofchanged, not a fifth
bucket — and groups the remaining diff-layer counts under labelled
Failures and Changes by layer bands. NewstatGroup()and
dispositionBar()helpers plus supporting CSS (.summary-headline,
.card-group,.disp*) reuse the existing colour tokens; no counts,
DiffResultaccessors, or other renderers changed.
Fixed
-
Report percentages are now locale-independent.
Format.signedPct()and
HtmlReportRenderer.fmt()formatted floating-point values with
String.format('%.1f', …)/String.format('%.2f', …), which use the JVM's
defaultLocale. Under a comma-decimal locale (e.g.pt_BR,de_DE) the
report emitted values like+12,5%and1,50instead of+12.5%and
1.50, corrupting the rendered percentages and breaking any downstream
numeric parsing that expects.as the decimal separator. Both call sites
now passLocale.ROOTso output is stable regardless of the host locale. -
The lineage-derived DAG now reads the
lineage/v1beta1store Nextflow
actually writes, instead of silently falling back to the symlink heuristic.
LineageStoreparsed a pre-v1beta1flat record shape — discriminator
type, withsessionId/name/inputat the top level. Current Nextflow
(25.04+) instead writes alineage/v1beta1envelope whose discriminator is
kindand whose payload is nested underspec. Every record therefore failed
thetype != 'TaskRun'guard,edgesForSession()returnednull, and
DagComparatorfell back to reconstructing edges from work-dir input symlinks
— so the authoritative-provenance path this class exists to provide was dead
against any real store, with nothing logged above debug level.LineageStore
now reads the discriminator viakindOf()(kind, falling back totype)
and the payload viaspecOf()(thespecmap, falling back to the record
itself), so both the current envelope and legacy flat stores reconstruct. The
existing unit tests were green only because they encoded the same obsolete
flat shape; a new fixture of realv1beta1records captured from a
rich-reportrun (src/test/resources/lineage/rich-report) now pins the
end-to-endINDEX_REF→ALIGN→MARKDUP→QC→MULTIQCreconstruction, alongside
directv1beta1envelope cases. Theexamples/rich-reportreport was
regenerated so its DAG layer reflects the authoritative lineage.
Full Changelog: v0.4.0...v0.5.0
Version 0.4.0
[0.4.0] - 2026-09-09
Fixed
--fail-on-changenow actually exits3when invoked vianextflow plugin. The documented CLI exit-code contract (1 runtime error, 2 usage
error, 3--fail-on-changeon a difference) is the whole reason
--fail-on-changeexists — a CI job keys off it. But thenextflow plugin <id>:<verb>launcher (CmdPlugin, on Nextflow'svoidLauncher.run()
path) invokes the plugin'sexec()and then discards its returned int, so
every invocation exited0no matter what — verified on 26.04.1, where two
identical:falseruns still exited0.DiffPlugin.exec()now forces the
process exit code itself viaSystem.exit(code)for any non-zero result
(flushing stdout/stderr first, since that path skips the trait's session
teardown), so 1/2/3 reach the shell regardless of whether the launcher
propagates the value. The exit-code mapping was split into a package-visible
dispatch()so it stays unit-testable without the terminalSystem.exit(),
ande2e/smoke.shnow asserts the real3end to end instead of noting the
bug as a known limitation.
Added
-q/--quiet/--summary-onlyprints just the summary block and skips
the report body. Every invocation rendered and wrote (or streamed) the full
HTML/JSON/Markdown report, even when a CI gate only cares about the one-line
N changed, …signal — so the report body was pure noise in the job log (and
a wasted file write). The new flag suppresses rendering entirely (contentis
never computed) and prints the summary to stdout regardless of--output; the
summary'sReport:line reads(suppressed by --summary-only). Exit-code
behaviour is unchanged, so--summary-only --fail-on-changeis now the leanest
CI gate. Accepts the launcher-injected--flag true, inline=, and explicit
=falseforms like the other boolean flags.- Apache-2.0 license headers on every source file. The repository ships an
Apache-2.0LICENSE, but none of the 40 Groovy sources (src/main+
src/test) carried the per-file SPDX/copyright header that the license text
itself recommends and that a registry-published plugin wants for clean
provenance. The canonical Nextflow header (Copyright 2026, Seqera Labs) is
now prepended to each file above itspackagedeclaration. - End-to-end smoke test that drives the real
nextflow plugin nf-diff:diff
launcher. The Spock suite exercises every component in isolation, but
nothing resolved the plugin by its bare id and ran the actual CLI verb
against a genuine.nextflow/history+ LevelDB cache — the exact path where
the internal Nextflow APIs the plugin reuses (HistoryFile,CacheDB/
DefaultCacheStore,ConfigBuilder) can drift between Nextflow lines. The
CI matrix already pinnedNXF_VERto both25.04.0and26.04.0for this
reason, yet only compiled and unit-tested against them. A newe2e/smoke.sh
(wired in asmake smoke) now runs a trivial pipeline twice to produce two
real runs, then invokes the plugin verb and asserts exit codes and report
content (JSONschemaVersion/summary, a standalone HTML document, and the
documented exit code3for--fail-on-changewhen the runs differ). CI
installs the matrix Nextflow version viaget.nextflow.io(which honours
NXF_VER) and runs it on every matrix leg, so drift is caught at the launcher
layer where it actually surfaces. --format=jsonoutput now carries a top-levelschemaVersionfield.
The JSON model emittedgeneratedAt,identical,runA/runB,summary
and the layers but no version marker, so a downstreamjqassertion in a CI
pipeline, PR bot or dashboard — the very consumers the README markets JSON to
— had no way to detect a breaking shape change.schemaVersion: "1"is now
emitted as the first key of the document, establishing an explicit contract
that can be bumped when the shape changes incompatibly.
Changed
- The opt-in work-dir layers now share one executor instead of one pool
each.--diff-outputs,--diff-logsand--diff-dageach fan their
independent, read-only work-dir I/O across a bounded thread pool — but
mapMatchedInParallel/runInParallelcreated (andshutdownNow()-tore-down)
a fresh pool per layer, so--diff-allpaid for three create/destroy cycles
in a single comparison.RunComparator.comparenow builds one daemon-threaded
pool (sized toavailableProcessors()) up front — only when at least one of
the three layers is enabled — threads it through the threecompute*methods,
and shuts it down once in afinally. Results still return inresult.tasks
order, and the single-pair sequential fast path (which never touches the pool)
is unchanged, so reports are byte-for-byte identical. - HTML report restyled to match the
nf-docsdesign language. The report
previously leaned on a dark-by-default, gradient-heavy look (radial body/hero
"glows", gradient-filled cards and chips, a gradient logo, 16px radii). The
internalnf-docs-generated pages use a flat, light-by-default documentation
aesthetic — a slate palette (slate-50/100/200surfaces in light,
slate-900/800/700in dark) with the shared Seqera primary green#0DC09Das
the sole brand accent, bordered white cards with only a hairline0 1px 2px
shadow, and tighter geometry. The report's inline CSS now adopts those exact
tokens: both theme palettes were re-mapped to slate +#0DC09D, the radial
gradients and gradient fills were removed in favour of flat bordered surfaces,
card/table radii dropped from 16px to 10px, the body gained
line-height:1.625and aui-sans-serif, system-ui, …stack, the page title
is now primary-green, and table rows gained a:hoverhighlight. Accent
colours (pills, row highlights, verdict, code-diff, warn-note, source badges)
were re-based onto the green/blue/green-500/red-500/yellow-500system,
with darkera16207/dc2626/16a34avariants in light mode for contrast on
white. Only the inlineCSSconstant changed — the report markup, JavaScript,
data-themetoggle mechanism and every existing class name are untouched, and
the report remains self-contained (no web fonts or external assets). - Derived summary counts now live on
DiffResult, not inline in each
renderer. The "software changed", "regressions", "outputs changed" and
"logs changed" stats were each recomputed inline in the HTML, JSON and
Markdown renderers (diff.software.count { it.changed },
diff.regressions.count { it.regression }, etc.).DiffResultalready
exposed peer accessors for the same class of derived count
(failedCountA(),newFailureCount(),dagEdgesAdded(),
overProvisionedA()) — these four just weren't pulled in, so if the
changed/regression/hasChangespredicate ever shifted the three
renderers could silently disagree. NewsoftwareChangedCount(),
regressionCount(),outputsChangedCount()andlogsChangedCount()
accessors sit next to the existing ones as the single source of truth, and
all three renderers now call them. RunComparatornow takes a singleCompareOptionsvalue object instead of
a twelve-argument positional constructor. The old signature interleaved four
booleans (showObvious,diffOutputs,diffLogs,diffDag) and three
numeric limits (outputsMaxBytes,logsMaxLines,outputsMaxLines), so only
argument order told them apart and a transposition compiled silently under
@CompileStatic— the classic long-parameter-list hazard, made worse because
adding a layer meant threading a new positional through every call site. The
newCompareOptionsnames each knob and defaults each to the constructor's old
default, soDiffCommandsets them by name andnew RunComparator()still
reproduces the no-argument behaviour. Future layers become one added field, not
a signature change at every call site.
Added
- Direct unit tests for the
ArgCursorparsing primitive. A new
ArgCursorTestpins the cursor's contract in isolation fromparse(): inline
--key=valuesplitting, the token-consumption semantics that the old manual
if( inlineVal == null ) i++bookkeeping encoded (requireValue/boolValue
consume the space-separated value; a bare or non-boolean-followed flag does
not), the shared numeric parse/floor helpers (intValue/longValue/
doubleValue), andpeek/consumePeekediteration.ArgCursorwas widened
fromprivateto package-visible for this. --diff-allconvenience flag enables the three opt-in work-dir layers
(--diff-outputs,--diff-logs,--diff-dag) at once. They share the same
precondition — the tasks' work directories must still exist — and are commonly
wanted together. The flag only enables, never forces off, so a later explicit
--diff-<layer>=falsestill opts an individual layer back out.- Direct unit tests for
RunLoader's pure helpers.RunLoaderis the
riskiest component (it reuses Nextflow's internalHistoryFile/CacheDB) yet
had no test. A newRunLoaderTestpins the pieces that are pure and
standalone — transient-lock detection (isLockError), the trace-store value
coercions (asLong/asString), and session-id shortening (shortId) — which
required only making those static helpers package-visible. - Fixture-backed test for
RunLoader.lastPair.lastPairreads only
.nextflow/history(no LevelDB cache), soRunLoaderTestnow writes a
hand-built history fixture and exercises the real selection logic end to end:
the default-B single-offset form, the explicitA:Bpair form (and its
equivalence toA:0), and the out-of-range guards (negative B,A <= B, and
too f...
Version 0.3.0
[0.3.0] - 2026-09-08
Added
- Lineage-backed DAG reconstruction (
--diff-dag) — when a run's project
directory has a Nextflow data-lineage store (.lineage/, produced with
lineage.enabled = trueon Nextflow 25.04+), the process wiring is now read
from the authoritative provenance Nextflow persisted instead of being
inferred from work-dir symlinks. A newLineageStorereads each.data.json
record directly off disk (no compile-time dependency on thenf-lineage
module), indexes the run'sTaskRunrecords by their session id, and
reconstructs producer→consumer edges from each task's recordedinputLID
references (lid://<producerTaskHash>/…). Because it reads what Nextflow
recorded, this needs no work directories and is unaffected by cleanup.
DagComparator.graphOfnow takes the run's project directory and prefers the
lineage store, falling back to the existing best-effort symlink
reconstruction (symlinkGraphOf) when no lineage store recorded the run. Each
run'sRunGraphcarries asource(LINEAGE/SYMLINK/NONE), and the
wiring layer's note now states whether the graph is authoritative (lineage) or
best-effort (symlinks), including the mixed case. The layer remains
informational only — it never affects the "identical" verdict or
--fail-on-change. Only the default<projectDir>/.lineagestore location is
auto-detected; a customlineage.store.locationstill falls back to symlinks.
Changed
-
Unknown plugin command now exits with the usage code (2), not 1. An
unrecognized verb (anything other thandiff) is a usage error, in the same
class as bad arguments, so it now returns2— matching the documented
exit-code table — instead of1(which is reserved for runtime errors). The
message also notes that onlydiffis supported. -
Help summary lists all current diff layers. The one-line description shown
by-h/--helpstill read "metadata, processes, and per-task
resources/scripts" from the 0.1.0 days; it now enumerates the always-on layers
(parameters, configuration, software & versions, failure rollup, performance
regressions, resource-efficiency) and the three opt-in flags. -
--diff-dagno longer requires work directories when lineage is enabled.
Previously the wiring layer always needed the tasks' work directories to still
exist locally; with a lineage store present it is reconstructed from persisted
provenance instead. -
Continuous integration & tag-based releases — a GitHub Actions CI
workflow (.github/workflows/ci.yml) now runs the full verification suite
(make check) on every push and pull request tomain, across JDK 17 and 21,
uploading test reports as build artifacts. A companion release workflow
(.github/workflows/release.yml) publishes to the
Nextflow plugin registry when av*version
tag is pushed: it verifies the tag matchesbuild.gradle'sversion(so a tag
can never publish a mismatched artifact), runsmake check, thenmake release, authenticating with anNPR_API_KEYrepository secret. See the
README's "Continuous integration" and "Releasing" sections for setup and the
tagging flow. -
DAG (process wiring) diff (
--diff-dag) — a new opt-in layer that
reconstructs each run's process;process wiring and diffs the two edge
sets, so nf-diff surfaces topology changes the task-count-per-process view
cannot see — e.g. a pipeline rewired fromA → CtoA → B → C. Nextflow
does not persist DAG edges in its history or cache, so there is no
authoritative edge list to read; what it does leave on disk is every task's
staged inputs, materialised as symbolic links inside the task's work
directory.DagComparatorwalks each task's work dir, resolves every input
symlink, and attributes any target that resolves into another task's work dir
(walking the parent chain so a link into a nested output subdir still
attributes to the producer) as a producer;consumer edge; links that
resolve outside every work dir are external inputs and yield no edge. Because
it walks work directories, this layer needs them to still exist locally (like
--diff-outputs/--diff-logs) and is a best-effort reconstruction: if some
work dirs were cleaned up, the recovered wiring is incomplete, and a note
reports how many task work dirs were missing so a partial diff is not read as
authoritative. Surfaced in all three report formats (HTML "Process wiring
(DAG)" section + nav link, Markdown section, and adagblock with
dagEdgesAdded/dagEdgesRemovedin JSON). Because the reconstruction is
best-effort, this layer is informational only and never affects the
"identical" verdict or--fail-on-change. -
Failure rollup (top-level "what failed and why") — a new always-on layer
that answers, at a glance, which tasks failed and why, instead of leaving that
scattered across per-task detail. Failed tasks are detected from the cached
status/exitfields (an explicitFAILED/ABORTEDstatus, or a non-zero
exit code — theNO_EXITsentinel and blanks are ignored), so the layer reads
straight from the run cache and needs no work directories. Failures are rolled
up by their(process, status, exit)signature and counted per run, sorted by
biggest blast radius first; a signature seen only in Run B is flagged new
(a regression), one present in Run A but gone in Run B is resolved, and one
in both is persistent. The layer also surfaces a run-level error state
(historystatusstartingERRor equal toFAILED/ABORTED/KILLED) even
when no individual task failure was recorded. Surfaced in all three report
formats (HTML "Failure rollup" section + nav link + summary cards, Markdown
section, and afailuresblock withfailedA/failedB/newFailures/
resolvedFailuressummary counts in JSON). Because the meaningful identity
signal — a task whose status or exit changed — is already carried by the task
field diffs, this rollup is informational only and never separately affects
isIdentical()/--fail-on-change. -
Resource-efficiency layer (requested vs. measured-peak provisioning) — a
new always-on layer that, per process, compares what each run requested
(cpus,memory) against what it actually peaked at (%cpu,peak_rss).
Both requested and peak values are read straight from the run cache trace and
taken as the max across a process's tasks (a retried task that used more, or
was bumped a higher request, is the honest worst case), so the layer needs no
work directories and is always computed. The efficiency ratio is
measured-peak / requested; each process is classified per run as over
(below 50% — wasted allocation, e.g. "requested 32 GB, peaked at 4 GB"),
tight (90%+ — risk of OOM kills or CPU throttling), or ok in between.
Surfaced in all three report formats (HTML section + an "Over-provisioned (B)"
summary card, Markdown table, and anefficiencyarray plus
overProvisionedA/overProvisionedBandtightA/tightBsummary counts in
JSON). Because provisioning is a tuning signal rather than a correctness
change, this layer is informational only — it never affectsisIdentical()/
--fail-on-change. -
Cross-project comparison (
--dir-a/--dir-b) — the two runs no longer
have to live in the same project. Previously a single--dirresolved both
runs'.nextflow/history, cache, config and params, so you could not compare
"the same pipeline in two checkouts" (or on two machines).--dir-a=<dir>and
--dir-b=<dir>now set each run's project directory independently; each falls
back to--dirwhen omitted, so existing invocations are unchanged. Run A is
loaded from and resolved againstdir-a, run B againstdir-b: the parameters
layer reads each run's own-params-file, and the configuration layer rebuilds
each run's effectivenextflow.configfrom its own working tree, so a
-profile dockerin project A is diffed against project B's config. The
git-provenance caveat became per-tree:ConfigProvenancenow carries a
crossProjectflag plus each side's directory, current HEAD and dirty state,
and its warning describes the two working trees separately (currentRevisionB,
dirtyB,dirA,dirBare surfaced in the JSON report).--laststill needs
a single history, so it is rejected when combined with differing
--dir-a/--dir-b.
Fixed
- Diff errors with no message printed a blank line.
DiffPlugin.exec()
reported a caught throwable viae.messageonly, so a message-less exception
(notablyNullPointerException) produced a barenf-diff:line with nothing
after it, while the stack trace went only to the debug-gated log. It now falls
back to the exception's simple class name, so both the stderr line and the log
always name the failure. - Cache-lock retry backoff was not interruptible. The backoff between
attempts to open a contended run cache used Groovy'ssleep(), which swallows
InterruptedExceptionand clears the interrupt flag, so aCtrl-Cduring a
contended open was ignored and the loop kept retrying. It now uses
Thread.sleep(), restoring the interrupt flag and aborting the retry on
interruption. GitProvenancesubprocess timeout was ineffective — the git subprocess's
stdout/stderr were read inline withgetText()before the timedwaitFor,
which blocks until the process exits, so a hunggitcould never be timed out.
Both streams are now drained on background threads started beforewaitFor, so
the 5s timeout actually fires and a chatty command cannot deadlock on a full
pipe buffer.
nf-diff 0.2.0
Added
- Line-level output diffing — under
--diff-outputs, a file classified as
changed that is text on both sides is now additionally diffed line by line
(reusing the sameLineDiffengine as--diff-logs), so the report answers
what changed rather than merely that it changed — a VCF, CSV, JSON, or
report file shows its added/removed lines inline.OutputComparatorsniffs
the head of each changed file for a NUL byte; binary files fall back to the
existing size/hash verdict and produce no line diff. Reads are bounded: the
first--outputs-max-lineslines (new option, default1000) and a hard byte
cap, so a huge file never blows up memory, with atruncatedmarker when a
cap dropped content. Surfaced in all three report formats (HTML unified-diff
pane with an added/removed line count, Markdown fenceddiffblock, and
diff/linesAdded/linesRemoved/truncatedfields on each output file in
JSON). This enriches the existing layer only — an output change still counts
toward the "identical" verdict and--fail-on-changeexactly as before. - Config-provenance caveat — the configuration layer now inspects the git
state of the working tree it resolves config from and warns when that tree has
drifted from the revision a run was actually launched at. Because
ConfigLoaderrebuilds each run's effective config from the files as they
exist now, a run launched at commit A and re-run at commit B have their
configs both resolved against whatever is checked out now — silently masking
config differences driven by code changes between those revisions. The
metadata layer already records each run'srevisionId; this compares it to
the currentHEAD(viagit rev-parse, prefix-matching abbreviated ids) and
also flags an uncommitted (dirty) working tree. When either run drifted, or
the tree is dirty, a prominent caveat is surfaced in all three report formats
(HTML warning banner, Markdown blockquote, and aconfigProvenanceobject
withdriftedA/driftedB/workingTreeDirty/warningfields in JSON). Git
state is inspected best-effort — a non-git project, missinggit, or a
command timeout degrades to "unknown" without breaking the diff. The caveat is
informational only: it never affects the "identical" verdict or
--fail-on-change. - Software & versions diffing — a new always-on layer that compares, per
process, the distinct container image(s) and Conda package spec(s) that
process's tasks ran with in each run. Both values are read straight from the
run cache's trace records, so the layer needs no work directories and is
always computed. It answers "did a tool version change?" directly — e.g.
biocontainers/fastqc:0.11.9→biocontainers/fastqc:0.12.1— instead of
leaving it buried in the per-task container field. A process present in only
one run is added/removed; a process in both whose container or Conda set
differs is flagged changed. A changed software environment counts toward the
"identical" verdict and--fail-on-change, which also makes a Conda-only
change (previously invisible to the task layer) break identity. Surfaced in
all three report formats (HTML section + summary card, Markdown section +
summary column, and asoftwarearray with asoftwareChangedsummary count
in JSON). - Output-file diffing — a new opt-in
--diff-outputslayer that compares
the files each task matched in both runs wrote to its work directory,
classifying them as added / removed / changed / unchanged. Files are compared
by size first, then by a streamed SHA-256 for same-size files. Staged inputs
(symlinks) and Nextflow control files (.command.*,.exitcode) are skipped;
cache-resumed tasks that share a work directory short-circuit as identical.
Unlike the performance-regressions layer, an output-file change counts toward
the "identical" verdict and--fail-on-change, so this is the layer that
answers "did my pipeline actually produce different results?". --outputs-max-bytes=<n>— caps the size of same-size files that are
hashed under--diff-outputs; larger files are reported content-unverified.
Default0means no limit.- Output diffs are surfaced in all three report formats (HTML, JSON, Markdown),
with anoutputsChangedcount in the JSON/HTML summary. - Failure / log diffing — a new opt-in
--diff-logslayer that compares the
standard log files (.command.out,.command.err,.command.log) each
matched task wrote to its work directory, line by line. It surfaces the
exit-code and status change alongside the log contents, so a task that went
from exit 0 to exit 1 can be inspected side by side — answering not that a
task failed but what it printed before it did. Reads are bounded (tailed to
a line cap and a hard byte cap) so an enormous log never blows up memory, and
tasks sharing a work directory (cache-resumed) short-circuit as identical.
Because task stdout/stderr legitimately varies between runs (timestamps,
paths, ordering), this layer is informational only: it never affects the
"identical" verdict or--fail-on-change— the exit-code change already
captured by the per-task diff does that. --logs-max-lines=<n>— keeps only the last<n>lines of each log file
before diffing under--diff-logs(default200).- Log diffs are surfaced in all three report formats (HTML, JSON, Markdown),
with alogsChangedcount in the JSON/HTML summary.
Fixed
--lastand bare boolean flags dropped by the plugin launcher — Nextflow's
pluginlauncher rewrites forwarded arguments before they reach the verb: a
bare--flagarrives as--flag true, and--opt=valuearrives
space-separated as--opt value. Thediffparser assumed the inline=
form survived, sonextflow plugin nf-diff:diff --lastreached it as
['--last', 'true']and the injectedtrueleaked into the positional list,
tripping the "--last cannot be combined with explicit run identifiers" guard.
The same latent bug affected--last=Nand every bare boolean flag
(--fail-on-change,--verbose,--diff-outputs,--diff-logs). The parser
now tolerates the launcher-normalized forms, consuming an injected/inline
true/false(or an integer for--last N) instead of treating it as a
positional run identifier.