Skip to content

Version 0.4.0

Choose a tag to compare

@mribeirodantas mribeirodantas released this 09 Sep 01:45
· 65 commits to main since this release

[0.4.0] - 2026-09-09

Fixed

  • --fail-on-change now actually exits 3 when invoked via nextflow plugin. The documented CLI exit-code contract (1 runtime error, 2 usage
    error, 3 --fail-on-change on a difference) is the whole reason
    --fail-on-change exists — a CI job keys off it. But the nextflow plugin <id>:<verb> launcher (CmdPlugin, on Nextflow's void Launcher.run()
    path) invokes the plugin's exec() and then discards its returned int, so
    every invocation exited 0 no matter what — verified on 26.04.1, where two
    identical:false runs still exited 0. DiffPlugin.exec() now forces the
    process exit code itself via System.exit(code) for any non-zero result
    (flushing stdout/stderr first, since that path skips the trait's session
    teardown), so 1/2/3 reach the shell regardless of whether the launcher
    propagates the value. The exit-code mapping was split into a package-visible
    dispatch() so it stays unit-testable without the terminal System.exit(),
    and e2e/smoke.sh now asserts the real 3 end to end instead of noting the
    bug as a known limitation.

Added

  • -q / --quiet / --summary-only prints just the summary block and skips
    the report body.
    Every invocation rendered and wrote (or streamed) the full
    HTML/JSON/Markdown report, even when a CI gate only cares about the one-line
    N changed, … signal — so the report body was pure noise in the job log (and
    a wasted file write). The new flag suppresses rendering entirely (content is
    never computed) and prints the summary to stdout regardless of --output; the
    summary's Report: line reads (suppressed by --summary-only). Exit-code
    behaviour is unchanged, so --summary-only --fail-on-change is now the leanest
    CI gate. Accepts the launcher-injected --flag true, inline =, and explicit
    =false forms like the other boolean flags.
  • Apache-2.0 license headers on every source file. The repository ships an
    Apache-2.0 LICENSE, but none of the 40 Groovy sources (src/main +
    src/test) carried the per-file SPDX/copyright header that the license text
    itself recommends and that a registry-published plugin wants for clean
    provenance. The canonical Nextflow header (Copyright 2026, Seqera Labs) is
    now prepended to each file above its package declaration.
  • End-to-end smoke test that drives the real nextflow plugin nf-diff:diff
    launcher.
    The Spock suite exercises every component in isolation, but
    nothing resolved the plugin by its bare id and ran the actual CLI verb
    against a genuine .nextflow/history + LevelDB cache — the exact path where
    the internal Nextflow APIs the plugin reuses (HistoryFile, CacheDB/
    DefaultCacheStore, ConfigBuilder) can drift between Nextflow lines. The
    CI matrix already pinned NXF_VER to both 25.04.0 and 26.04.0 for this
    reason, yet only compiled and unit-tested against them. A new e2e/smoke.sh
    (wired in as make smoke) now runs a trivial pipeline twice to produce two
    real runs, then invokes the plugin verb and asserts exit codes and report
    content (JSON schemaVersion/summary, a standalone HTML document, and the
    documented exit code 3 for --fail-on-change when the runs differ). CI
    installs the matrix Nextflow version via get.nextflow.io (which honours
    NXF_VER) and runs it on every matrix leg, so drift is caught at the launcher
    layer where it actually surfaces.
  • --format=json output now carries a top-level schemaVersion field.
    The JSON model emitted generatedAt, identical, runA/runB, summary
    and the layers but no version marker, so a downstream jq assertion in a CI
    pipeline, PR bot or dashboard — the very consumers the README markets JSON to
    — had no way to detect a breaking shape change. schemaVersion: "1" is now
    emitted as the first key of the document, establishing an explicit contract
    that can be bumped when the shape changes incompatibly.

Changed

  • The opt-in work-dir layers now share one executor instead of one pool
    each.
    --diff-outputs, --diff-logs and --diff-dag each fan their
    independent, read-only work-dir I/O across a bounded thread pool — but
    mapMatchedInParallel/runInParallel created (and shutdownNow()-tore-down)
    a fresh pool per layer, so --diff-all paid for three create/destroy cycles
    in a single comparison. RunComparator.compare now builds one daemon-threaded
    pool (sized to availableProcessors()) up front — only when at least one of
    the three layers is enabled — threads it through the three compute* methods,
    and shuts it down once in a finally. Results still return in result.tasks
    order, and the single-pair sequential fast path (which never touches the pool)
    is unchanged, so reports are byte-for-byte identical.
  • HTML report restyled to match the nf-docs design language. The report
    previously leaned on a dark-by-default, gradient-heavy look (radial body/hero
    "glows", gradient-filled cards and chips, a gradient logo, 16px radii). The
    internal nf-docs-generated pages use a flat, light-by-default documentation
    aesthetic — a slate palette (slate-50/100/200 surfaces in light,
    slate-900/800/700 in dark) with the shared Seqera primary green #0DC09D as
    the sole brand accent, bordered white cards with only a hairline 0 1px 2px
    shadow, and tighter geometry. The report's inline CSS now adopts those exact
    tokens: both theme palettes were re-mapped to slate + #0DC09D, the radial
    gradients and gradient fills were removed in favour of flat bordered surfaces,
    card/table radii dropped from 16px to 10px, the body gained
    line-height:1.625 and a ui-sans-serif, system-ui, … stack, the page title
    is now primary-green, and table rows gained a :hover highlight. Accent
    colours (pills, row highlights, verdict, code-diff, warn-note, source badges)
    were re-based onto the green/blue/green-500/red-500/yellow-500 system,
    with darker a16207/dc2626/16a34a variants in light mode for contrast on
    white. Only the inline CSS constant changed — the report markup, JavaScript,
    data-theme toggle mechanism and every existing class name are untouched, and
    the report remains self-contained (no web fonts or external assets).
  • Derived summary counts now live on DiffResult, not inline in each
    renderer.
    The "software changed", "regressions", "outputs changed" and
    "logs changed" stats were each recomputed inline in the HTML, JSON and
    Markdown renderers (diff.software.count { it.changed },
    diff.regressions.count { it.regression }, etc.). DiffResult already
    exposed peer accessors for the same class of derived count
    (failedCountA(), newFailureCount(), dagEdgesAdded(),
    overProvisionedA()) — these four just weren't pulled in, so if the
    changed/regression/hasChanges predicate ever shifted the three
    renderers could silently disagree. New softwareChangedCount(),
    regressionCount(), outputsChangedCount() and logsChangedCount()
    accessors sit next to the existing ones as the single source of truth, and
    all three renderers now call them.
  • RunComparator now takes a single CompareOptions value object instead of
    a twelve-argument positional constructor.
    The old signature interleaved four
    booleans (showObvious, diffOutputs, diffLogs, diffDag) and three
    numeric limits (outputsMaxBytes, logsMaxLines, outputsMaxLines), so only
    argument order told them apart and a transposition compiled silently under
    @CompileStatic — the classic long-parameter-list hazard, made worse because
    adding a layer meant threading a new positional through every call site. The
    new CompareOptions names each knob and defaults each to the constructor's old
    default, so DiffCommand sets them by name and new RunComparator() still
    reproduces the no-argument behaviour. Future layers become one added field, not
    a signature change at every call site.

Added

  • Direct unit tests for the ArgCursor parsing primitive. A new
    ArgCursorTest pins the cursor's contract in isolation from parse(): inline
    --key=value splitting, the token-consumption semantics that the old manual
    if( inlineVal == null ) i++ bookkeeping encoded (requireValue/boolValue
    consume the space-separated value; a bare or non-boolean-followed flag does
    not), the shared numeric parse/floor helpers (intValue/longValue/
    doubleValue), and peek/consumePeeked iteration. ArgCursor was widened
    from private to package-visible for this.
  • --diff-all convenience flag enables the three opt-in work-dir layers
    (--diff-outputs, --diff-logs, --diff-dag) at once. They share the same
    precondition — the tasks' work directories must still exist — and are commonly
    wanted together. The flag only enables, never forces off, so a later explicit
    --diff-<layer>=false still opts an individual layer back out.
  • Direct unit tests for RunLoader's pure helpers. RunLoader is the
    riskiest component (it reuses Nextflow's internal HistoryFile/CacheDB) yet
    had no test. A new RunLoaderTest pins the pieces that are pure and
    standalone — transient-lock detection (isLockError), the trace-store value
    coercions (asLong/asString), and session-id shortening (shortId) — which
    required only making those static helpers package-visible.
  • Fixture-backed test for RunLoader.lastPair. lastPair reads only
    .nextflow/history (no LevelDB cache), so RunLoaderTest now writes a
    hand-built history fixture and exercises the real selection logic end to end:
    the default-B single-offset form, the explicit A:B pair form (and its
    equivalence to A:0), and the out-of-range guards (negative B, A <= B, and
    too few runs in history).

Changed

  • Argument parsing is centralised behind an ArgCursor. DiffCommand.parse
    previously hand-rolled, for every option, the inline --key=value split, the
    space-separated --key value fallback with its manual if( inlineVal == null ) i++ index bookkeeping, and — for each numeric flag — a duplicated
    parse/NumberFormatException/range-check block. A private ArgCursor now owns
    position tracking and exposes requireValue/boolValue/intValue/longValue/
    doubleValue helpers, so each option case collapses to a single assignment and
    the off-by-one hazard in the repeated i++ dance is gone. Behaviour is
    unchanged (all forms — inline, space-separated, and launcher-injected
    --flag true/--last N/--last A:B — parse exactly as before); only the
    numeric-validation messages are now generated from a shared template.
  • --last gained an explicit A:B pair form and clearer docs. A single
    --last=N still compares the run N positions before the latest against the
    latest — but that silently skips the runs in between, which was easy to
    misread as "the N most recent runs". You can now name an exact pair by their
    offsets back from the latest (0 = latest, requiring A > B >= 0): e.g.
    --last=2:1 compares the run two back against the run one back. The bare
    --last and single-integer forms are unchanged (--last=N ≡ --last=N:0).
    The -h text now spells out the skip behavior, and the info line logged at
    selection names each side's offset explicitly.
  • Per-task raw trace map is no longer deep-copied. RunLoader.toTaskInfo
    built each TaskInfo with raw = new LinkedHashMap<>(store), duplicating the
    entire trace store on top of the already-copied display map — roughly
    doubling per-task memory on large runs. CacheDB.eachRecord deserializes a
    fresh TraceRecord (and store map) per iteration and the record is discarded
    immediately, so nothing can mutate or reuse it; the defensive copy bought no
    isolation. TaskInfo now references the store map directly.
  • Cache-lock detection is no longer coupled to a single literal message.
    RunLoader.isLockError — which decides whether a failed cache open is a
    transient lock contention worth retrying — previously matched exactly one
    string (Unable to acquire lock). A phrasing change in LevelDB or Nextflow
    would silently disable the retry loop. It now matches a small, case-insensitive
    allow-list of known lock signatures (still lock-specific, so unrelated failures
    are never retried pointlessly) while continuing to walk the cause chain.
  • The opt-in work-dir layers now compare tasks in parallel. --diff-outputs
    and --diff-logs previously walked matched task pairs one at a time, so a
    pipeline with many tasks and large outputs paid for single-threaded
    SHA-256/log I/O. Because each matched pair is independent, read-only work-dir
    I/O, RunComparator now fans the comparisons out across a bounded pool (sized
    to the smaller of the work size and the available processors) while still
    returning results in result.tasks order, so the report is byte-for-byte
    unchanged. --diff-dag likewise reconstructs both runs' graphs concurrently.
    Failures propagate unchanged (the underlying exception is unwrapped from the
    executor), and pool threads are daemon so a stuck read never keeps the JVM
    alive.
  • Command-line tokenisation now has a single source of truth. ConfigLoader
    previously carried its own copy of the quote-aware command tokenizer that
    "mirrored" CommandParams'; the two could silently drift. CommandParams.tokenize
    is now shared and ConfigLoader delegates to it, so -c/-config extraction
    and flag parsing always split commands identically.

Full Changelog: v0.3.0...v0.4.0