Skip to content

proef v0.17.0

Choose a tag to compare

@github-actions github-actions released this 06 Sep 14:11
· 28 commits to main since this release
e5f8d89

Added

  • [http] cookie-store = false runs the whole suite cookie-less — hurl
    8.0's --no-cookie-store, surfaced through the table built for exactly this
    class of setting. No Set-Cookie is retained and none is replayed, which is
    how a stateless API is proven stateless: the fixture-backed test is green
    only because its steps assert the 403 a missing session cookie earns.

    This is the one [http] key with no per-entry [Options] spelling at
    all
    (OptionKind has no cookie variant — verified against the enum), so
    run-wide is not a compromise but the only place it can be said. With the
    store off, the engine also skips both halves of the batch-split cookie
    round-trip: hurl reads a cookie_input_file only when enabling the engine,
    so injecting one would be silently ignored — and there is nothing to write.
    hurl's own FIXME (a handle once given cookie storage cannot lose it) never
    reaches proef, because run_entries builds its client per call (TECH-SPEC
    §5) — a handle never transitions on → off.

    Breaking (library): HttpDefaults gains the cookie_store field, so a
    struct-literal construction needs the new line (..Default::default() sites
    are untouched, and an absent [http] cookie-store key changes nothing).

  • --ctrf <path> — the run's verdicts as a CTRF report. CTRF
    (https://ctrf.io) is the emerging JSON successor to JUnit XML for CI
    dashboards, and it models in the schema what JUnit can only smuggle
    through extensions — which is exactly the data proef already tracks: a
    pass-after-retry carries flaky, retries, and retryAttempts listing
    the real failed attempts with their (redacted) messages; every test carries
    its tags and file path. One serializer off the same fold as JUnit, so the
    two files cannot disagree — most visibly for a quarantined failure, which
    both report as skipped with a message (ADR-0019), because a dashboard
    reading "failed" beside exit 0 would contradict itself. A User/System
    fault stays failed even under a quarantine tag: quarantine is for flaky
    tests, not broken input.

    The R12-3 contract applies from day one: a [run] setup abort still writes
    the file, carrying the setup scenario itself — a job gating on the report
    must never see no file at all. The schema's required wall-clock
    start/stop are measured at the CLI edge like every other clock read
    (ADR-0015); the sans-IO core and the JSONL record are untouched — the
    record remains the only record (ADR-0008).

  • The HTML report answers "what is slowest". After "what failed", it is the
    question a test report is most often asked, and the page could not answer it:
    the timeline showed that workers were busy, never which scenarios to
    attack. Every number needed was already in the fold.

    A ranked section, slowest first, each row linking to its own block, with the
    heading reporting the share of run time the listed scenarios account for —
    "3 of 40 · 71% of run time" is a decision, where a column of durations is
    homework. Capped at eight: a ranking long enough to scroll has stopped
    answering the question.

    Cost is the sum of a scenario's step durations, the same definition
    timings.json uses for shard weights — one notion of what a scenario costs
    across the whole tool. Not the wall-clock span, which includes time waiting
    for a worker: a property of how the run was scheduled, and not something the
    reader can go and fix.

    Absent when there is nothing to rank — fewer than two timed scenarios, or a
    record with no injected durations at all.

  • --shard-weights balances a shard matrix by measured duration. --shard
    assigns by a frozen hash, which guarantees that adding one scenario never
    re-buckets the others but cannot balance by time — and a CI matrix finishes
    when its slowest shard does, so a count-split routinely leaves runners idle.
    Every run that reaches its suite now writes a small timings.json into its
    run directory; CI archives that one file and each matrix job points
    --shard-weights at the same copy.

    The obvious design is silently wrong, and the module says so at length.
    proef already retains records carrying every step's duration, so "weight by
    the newest local record" looks free. But matrix jobs run on different
    machines
    , each with its own (usually empty) runs-dir — every job would
    compute a different weight table, therefore a different assignment, and
    scenarios would run twice or not at all while the suite reported green.
    Nothing about that announces itself. One named file shared by every job is
    what makes the split a pure function of (selected scenarios, that file).

    Two rules place scenarios and they partition rather than compete: a
    scenario the file mentions goes through longest-processing-time-first
    placement, and one it does not mention falls back to the frozen hash. So a
    test added after the timings were captured still runs exactly once. That is
    pinned by a test that runs a whole three-way matrix — with a weights file
    covering only five of nine scenarios, so both rules are exercised at once —
    and asserts set equality both ways; mutating the placement by one bucket drops
    two scenarios and the test names them.

    The weight is the sum of a scenario's step durations, not its wall-clock
    span. The span includes time spent waiting for a worker, which is a property
    of the run's scheduling rather than of the scenario, and feeding it back would
    let one crowded run's queueing distort the next split.

    What this gives up is exactly what hash mode was chosen for: a balanced split
    is not stable under insertion. That is what balancing means, which is why the
    flag is opt-in. A missing or malformed weights file is exit 2 — falling back
    silently would hand back the unbalanced split the flag was passed to avoid.

  • The editor tells proef's two variable tiers apart. A pack's hurl: | block
    is the centre of the authoring experience and, to every editor, a plain YAML
    scalar — inside which ${…} (resolved at lower time, by proef, before any
    request exists) and {{…}} (resolved at run time, by hurl) look
    identical. That distinction is ADR-0005's whole model and the thing authors
    most often get wrong, and no generic grammar can see it: a YAML highlighter
    sees a string, and a hurl highlighter never runs because the block is not a
    file. proef is the only party that knows.

    The server now answers textDocument/semanticTokens/full, lighting ${…} as
    macro — a substitution performed before execution, which is what a macro
    is — and {{…}} as variable. Both are coloured differently by every
    mainstream theme, so it works without anyone configuring anything. The $${
    escape stays dark, because telling an author proef will substitute text it
    will in fact leave alone is worse than no highlighting.

    The ${…} scan is proef_core::resolve::reference_spans, walking the same
    first_reference the resolver itself uses — a second implementation of the
    escape rule would drift, and the drift would show as an editor confidently
    colouring literal text. The {{…}} scan lives in proef-lsp rather than
    core, because that spelling is the engine's and ADR-0002's amendment is that
    engine syntax does not accumulate in the core.

    Collapsing the seven-arm request dispatch behind a local macro came with it:
    the chain crossed clippy's line limit the moment an eighth feature landed, and
    the honest fix was to stop repeating an identical frame seven times rather
    than to suppress the lint that noticed.

  • The linear-validation claim is now a test, not a sentence. #138 made pack
    validation linear and recorded the result as a shape: "the curve changed
    shape — 4× per doubling before, ~2× after"
    . That number lived only in the
    changelog, where nothing could re-run it — so a future span locator scanning
    the whole pack file again would have restored the quadratic behaviour
    silently, a regression that costs seconds rather than correctness and which no
    gate measured.

    The guard asserts the ratio between 1000 and 2000 macros, because the
    claim is a ratio. It observes ~2.05× against a bound of 3.0; mutating
    locate::MacroIndex to re-index per lookup — the exact pre-#138 shape —
    measures 4.01×, matching the changelog's own prediction of 4× and turning
    a 0.4-second test into a 73-second one. The failure message names the cause
    rather than reporting a number.

    A ratio rather than a benchmark, for a reason now written into
    TESTING-STRATEGY.md §7: load on a shared runner inflates both measurements
    together and cancels, where an absolute threshold has to be loosened until it
    means nothing. iai-callgrind would be the better CI gate — instruction
    counts ignore runner noise entirely — but it needs valgrind, so it would be a
    gate the maintainer cannot reproduce on macOS; criterion and divan sit in
    the same noise regime as this test while adding a dependency tree to a
    workspace that audits every edge. No new dependency was added.

  • Every diagnostic code is now named by a test, and a guard keeps it that
    way.
    DIAGNOSTICS.md calls codes "a contract: they never change meaning".
    Twenty-three of seventy-five had nothing holding them to it — reachable in
    production, documented, exercised by nothing at all: not a seeded corpus
    directory, not a unit test, not even an assertion on their message text. They
    existed only at their definition site.

    The catalogue itself was found exactly honest — 75 codes defined, 75
    documented, and its corpus column matched disk in both directions with zero
    drift. The gap was never documentation; it was that a documented promise had
    no enforcement.

    Nineteen new tests close it, each reaching its code through a real path rather
    than constructing the diagnostic directly. Two of them exercise guards that
    are unreachable in normal operation and were therefore the most valuable to
    test: lower::kind_unrouted fires only when the engine registry and pack
    validation disagree, so the test makes them disagree on purpose; and
    lower::expansion_too_deep sits behind pack validation's identical depth
    limit, so the test bypasses validation with load_collecting — the only way
    to hand lowering a graph validation would have stopped, and therefore the only
    way to prove the second line of defence is still there.

    Two codes are exempted by name, with reasons recorded in the guard:
    source::unreadable and config::unreadable need a file the process may stat
    but not read, a permissions state CI runners do not reproduce because they run
    as root. The guard also checks its own exemption list, failing if an exempted
    code is deleted or renamed — an exemption that outlives its code silently
    excuses nothing.

    The guard joins the four in source_guards.rs and is mutation-verified:
    rewriting one test to match a code by suffix instead of naming it turns the
    guard red, which is the point — a test that matches the prose pins the
    wording, and only one that names the code pins the contract.

  • [http] now carries the settings that describe an environment: TLS, proxy
    and mTLS.
    The table exposed two of hurl's runner options — timeout-ms and
    follow-location — while the embedded engine has supported the rest all
    along; TECH-SPEC.md:235 even listed insecure among what RunnerOptions
    carries. So a suite that had to run against staging's self-signed certificate,
    or through a corporate proxy, or against an mTLS-protected API, could not say
    so anywhere: the only route was repeating an [Options] block inside every
    macro's raw hurl, which defeats environment profiles exactly where they are
    most useful, since these settings are the difference between environments.

    Eight new keys — insecure, proxy, no-proxy, cacert, client-cert,
    client-key, max-redirs, user-agent — each merging field-wise through the
    existing [http] < [env.<name>.http] chain, so a staging profile turns
    verification off without production inheriting it. No new concept: only more
    of one that already worked.

    Three deliberate edges. insecure = true warns on every run, naming the
    profile that set it — a suite that goes green without verifying a certificate
    has not proved what a green suite normally proves, and since the run record
    carries no config by design, the warning is the entire audit trail. A
    client-key without a client-cert is exit 2
    rather than a pass-through:
    libcurl accepts the pair and then presents nothing, so the failure would
    otherwise surface at the server as an authentication error naming nothing
    about the cause. And credentials are excluded on purpose — there is no
    user or netrc key, because a password belongs in the secret store where it
    is encrypted at rest and masked out of every sink.

    The three path-valued keys resolve against proef.toml, the one-path rule
    every other config path follows; core still reads no filesystem and receives
    them already resolved (ADR-0012). Each option is applied to hurl's builder
    only when actually set, so a project with no [http] table runs
    byte-identically to one built before the keys existed — pinned by a test.
    Per-entry [Options] still override all of them except user-agent, for
    which hurl has no per-entry option at all; that exception is documented rather
    than papered over.

    Breaking (library): proef_core::engine::HttpDefaults gains eight fields and
    loses Copy — it now carries Strings. Default stays hand-written, and
    the reason is now stated in the type: a derive would make timeout_ms zero,
    which libcurl reads as no timeout at all, silently converting ADR-0007's
    budget into an unbounded wait at every existing default() call site.

Changed

  • The toolchain policy is stated in the spec that rust-toolchain.toml
    cites.
    R18-2 corrected the policy to latest stable, adopted at its
    x.y.1 point release
    , and the correction reached RELEASING.md and
    CLAUDE.md while TECH-SPEC §15 — named by rust-toolchain.toml as its
    authority — still said "always latest stable". R18-2's own conclusion was
    that an unwritten policy contradicting the written one is a docs defect;
    fixing it in two files and leaving the source of truth contradicting itself
    reproduced the defect one level down. Now consistent across all four.

  • An artifact is named by its feature's path, not its stem — two scenarios
    can no longer claim one file.
    Slugs were {stem}--{scenario}, dropping the
    directory, so features/x.feature and features/sub/x.feature each with a
    same name scenario both produced x--same-name: the second artifact
    silently overwrote the first while the CLI reported writing two. Silent
    loss of the hand-off ADR-0010 calls a contract — and the project already
    treats same-named scenarios across files as real, which is what
    --scenario-file exists for. The same slug drives the HTML report's anchors
    and artifact links, reproduce: lines, and harness trial names, so all of
    them move together off the one helper.

    Names are now features-sub-x--same-name. Derived from the path rather than
    disambiguated on collision, deliberately: a counter or hash appended only
    when two names clash would make one scenario's artifact name depend on
    whether some other file exists, so adding a feature would rename an
    unrelated artifact — the instability --shard's frozen hash exists to
    avoid. The path fed in is the portable suite-relative name the record
    carries, never a path off the running machine.

    Breaking, and quietly so for library callers: emit::artifact_slug keeps
    its (&str, &str) -> String signature while its first argument changes
    meaning from stem to feature path, so the API gate cannot see it —
    passing a stem still compiles and now yields a different name.
    emit::emit's second parameter changes the same way, and emit::feature_stem
    is removed (it had no remaining consumer). Artifact filenames and report
    anchors change for every suite; the snapshot corpus was regenerated under the
    new names and reviewed.

    The unification that made that a one-line change came first: the stem
    expression (file_stem, falling back to "feature") had existed four times
    across both crates — the emitter's caller, the dispatcher's spec naming, the
    HTML report's anchors, the editor's analysis — and the stem--scenario
    composition twice, with the report's links to artifact files resolving only
    because both sides happened to derive the same name. Worklist item Q6 called
    the four sites a future-drift risk; collapsing them to one helper is what
    let the collision above be fixed in a single place instead of four. In the
    same pass, Q2 (the editor's per-request walks) was found already closed by
    the #146 analysis cache, and its entry now says so with the evidence.

  • "What a scenario costs" is defined once, as ScenarioOutcome::cost. The
    sum of a scenario's step durations was computed in three places on the same
    type — JUnit's per-suite time, JUnit's per-case time, and the new
    timings.json weights — plus a fourth over the record-fold shape in the HTML
    report. Four surfaces free to drift apart about a number they are supposed to
    agree on, and the argument for summing steps rather than taking a wall-clock
    span was written out twice.

    Now a method on the type that owns the steps, with the rationale stated there
    and referenced from the rest. The one behaviour change is a fidelity gain: the
    weights file used to truncate each step to whole milliseconds before summing
    and now truncates the sum, so its numbers agree with the times JUnit has
    always reported. Additive to the library surface.

  • A run whose setup aborted no longer leaves shard weights behind.
    timings.json was written from inside the CI-report block, which a setup
    abort also reaches — with the setup phase's summary. The file that came out
    named setup scenarios, and a weights file naming them is worse than no file:
    those identities never appear in a suite run, so they absorb bucket load on
    behalf of scenarios that never run and skew the very split --shard-weights
    exists to balance, silently. The write moved to the one site where the summary
    is the suite's, pinned by a test that reproduces the old file.

  • lower.rs stops threading the same three values through twelve
    functions.
    out, refs and sinks travelled as separate parameters
    everywhere, and five functions — expand_macro, expand_step,
    expand_ref_step, expand_payload_step, finish_step — carried 8 to 11
    parameters each behind individual arity suppressions. Adding one piece of
    lowering state meant editing five signatures and five call sites, which is the
    shape of change that drops a parameter at one site.

    Two bundles, both of them types that were already implied by the code:
    Emit { out, refs, sinks } (the mutable outputs, always passed together and
    never independently), StepScope { step_ref, ctx, at } (what stays fixed for
    one authored step however deep expansion recurses), and a small Finished
    for the four values that describe a step being completed.

    What was not done matters as much. The obvious refactor — hoist the
    state into a self and make the five methods — would have broken the reason
    they are parameters at all: resolve_in and friends take them explicitly so
    they remain callable while other state is mutably borrowed, and a method on
    &mut self cannot be called while self is borrowed elsewhere. The threading
    discipline is load-bearing, so it stays; only the arity changes.

    Arity suppressions across the workspace: 13 → 6, with lower.rs at zero.
    No behaviour change, and the 241 core tests say so.

  • ADR-0002 now names the core's hurl entry grammar, and a guard keeps it
    closed.
    "Adding an engine leaves proef-core diff-empty" was true of
    engine-types and never of engine-syntax: the core does text surgery on
    entries — splicing [Options] in, merging an expect: block's asserts into
    the previous entry — so it has to find an entry boundary in text hurl will
    later parse. The worklist carried the gap for two rounds as "~290 lines of
    hurl grammar in core", a figure that counted #[cfg(test)] fixtures.

    Measured: twelve literals across four files. Seven the core writes, four it
    recognises to find a boundary, and one it quotes — a hurl snippet inside
    a did-you-mean help string in bind.rs, which generates nothing and parses
    nothing but drifts like any other copy. The four boundary recognisers are
    already one shared pub(crate) set. proef's own pack keys are shaped like
    option lines and are excluded by name rather than listed as sanctioned rows.

    The guard lexes whole files. The first version scanned line by line and so
    could not see a literal that spans lines — which is where a larger piece of
    engine syntax would naturally be written, and where the one entry above that
    nobody had counted was in fact sitting.

    The amendment sanctions that set and closes it. Deferred with a named
    trigger — a second engine being scheduled — is moving the written half behind
    the seam, where the reading half already lives:
    StepKindSpec::options exists precisely so an engine's option spellings stay
    out of the core, and it covers recognising them only, so retry:,
    retry-interval:, delay: and variable: are still core literals. Until a
    second engine exists that migration relocates seven literals that exactly one
    implementation will ever supply, at the cost of a public-API break.

    crates/proef-cli/tests/source_guards.rs (renamed from stderr_hygiene.rs,
    which had not been only about stderr for two rules now) pins the set: a new
    token, or an existing one spreading to another core module, fails the test
    and names both remedies. A claim of this shape decays the moment it is only
    prose — this one already had, by an order of magnitude, in the direction that
    made it look worse than it is.

Fixed

  • A file,…; body in a ref: fragment resolves where its author put it.
    hurl resolves a file body against the directory of the file that wrote the
    reference
    — its --file-root default, and the same rule Karate, pytest and
    Jest use for fixtures. proef resolved every asset against the feature,
    and a fragment lives in another tree entirely ([run] fragments), so the
    same bytes passed under stock hurl and failed under proef, as exit 2,
    blaming the author for a path that was correct. Nothing worked around it:
    moving the file beside the feature breaks the standalone run ADR-0018
    exists to guarantee, a reaching ../ path is refused by hurl's own sandbox,
    and the advice that refusal prints — "check --file-root option" — names a
    flag proef does not expose.

    Each asset is now staged from beside the source that referenced it, feature
    or fragment, into that scenario's own asset root, which is what the engine
    gets as its context dir. Staging rather than two roots because hurl offers
    one context dir per run of entries and no per-entry override, while a
    single batch may mix both body forms — measured, not assumed: an inline step
    and a ref: step in one macro lower to one batch. Copying fixtures into the
    build output is the standard answer to exactly this, and it adds no copy
    operation: the record already copied these files once per scenario, just
    into a shared directory instead of the right one. What it does change is the
    footprint — an asset N scenarios read is now N files in the run record
    rather than one, which is the same fact as the collision below, seen from
    the disk's side rather than the reader's.

  • Two scenarios' assets no longer overwrite each other. Staging was flat
    and keyed by the asset's bare name, so two features that each keep a
    data.json beside them staged to one file — last writer wins, with "0
    warning(s)" — and the loser's artifact replayed against the other's bytes.
    artifact_slug already refuses that trade for the .hurl text, deriving
    from the feature's whole path so two same-named scenarios cannot collide;
    the files it reads now get the same treatment. An artifact that reads a
    file says so in its replay line (--file-root assets/<slug>); one that does
    not is byte-identical to before. Two sources claiming one name inside a
    single scenario — the case a per-scenario root cannot separate — is refused
    rather than narrowed.

    A missing asset is also an error now instead of a silent skip. It had to
    become one: the staged root is what the engine reads, so a file that quietly
    failed to arrive is no longer an incomplete record but a request reading
    nothing.

    [Options] output: resolves through the same root, so the root is created
    for every scenario rather than by the staging loop — which never runs for a
    scenario that reads no file body. A response written that way now lands
    inside the run record, where a run's outputs belong, instead of in the
    feature's own directory.

    Breaking (library): emit::file_references is replaced by
    Artifact::assets, a Vec<AssetRef> carrying each reference with the
    source that wrote it
    — the provenance a whole-artifact text scan destroys,
    and the whole reason the bug was expressible. emit::asset_root names the
    staging directory for the three call sites that must agree on it, and
    pack::split_qualified is now the one reader of the file.hurl#name form
    Fragment::qualified writes — there were two, resolving a ref: and a
    use:, and staging assets was about to make a third in another crate. New
    diagnostic: proef::run::asset_unstageable.

  • A --run-id record is findable again (ADR-0021). --run-id pr-1234
    writes a perfectly good record, and every command that resolves the latest
    run
    — explain, diff, flaky, report, --rerun — enumerated by the
    uuid shape, so that record was invisible to all of them. --rerun was the
    sharp edge: it silently continued some older run instead of the one just
    produced.

    One predicate had been answering two questions whose risks point in opposite
    directions — may I delete this? is unsafe when broad, is this a run I can
    show you?
    is unsafe when narrow — so the deletion-safety choice had
    silently become a visibility choice. They are now separate: a directory is a
    record because it holds an events.jsonl, while rotation still deletes
    only uuid-named directories, so a custom-id run is discoverable and still
    never reclaimed by [run] keep-runs. Ordering stopped riding on the name
    too — uuid-v7 sorted chronologically until a custom-id directory joined the
    set and sorted by its first letter — and now takes the timestamp a uuid-v7
    name carries (48 bits of unix milliseconds, which is why the lexical sort
    worked), falling back to directory mtime for a name that carries none.

  • A run with a failed optional: step no longer prints exactly like a
    spotless one.
    ConsoleMode::Failed's own doc comment states the
    requirement and the Warned arm implementing it was unreachable: a
    scenario's aggregate status was only ever Failed | Skipped | Passed, so a
    real optional failure was invisible under --console failed, showed a .
    rather than the documented w under --console dotted, and left the HTML
    report's warned count and its filter-bar warned button permanently empty —
    four consumers and three docs describing something that could not occur.
    Steps carried Warned; scenarios never did. The aggregate now promotes,
    which changes what a run says and never whether it gates: Warned counts
    as passing in the exit code, the totals and JUnit.

  • explain prints a step's authored name:, like its five siblings.
    step_label's own doc enumerates the six surfaces that must render it —
    console, HTML, JUnit, TAP, the job summary, explain — and explain was
    the one that never called it, so the post-mortem tool showed one sentence
    repeated where the live console had told the steps apart.

  • The HTML "Slowest" section no longer counts [run] setup/teardown into
    "% of run time".
    Every other aggregate on the page excludes phases
    (ADR-0014), including the tag table directly above it, so the share meant
    something different in that one section. A slow phase stays visible in the
    timeline and in its own block.

  • .cargo/audit.toml no longer suppresses advisories deny.toml
    deliberately un-suppressed.
    It carried the quick-xml pair
    (RUSTSEC-2026-0194/0195) with a comment claiming it mirrored deny.toml —
    which had removed them, precisely because the reason had expired
    (quick-junit 0.7 moved to the patched quick-xml 0.41, which the lockfile
    is on). So the nightly cargo audit job was suppressing for no reason, and
    would also have silenced any new advisory filed against that line.

  • A merged report covers the whole suite again, and its headline agrees
    with its page.
    Two independent failures in --rerun composition, against
    docs/CI.md's promise that "one report stands for the composed result". The
    overlay followed only the immediate rerun_of, so the ordinary
    fix → rerun → fix → rerun loop — the workflow the feature exists for —
    silently dropped everything from before the last link, with no banner saying
    so; the page just got smaller. It now walks the chain, newest verdict
    winning, with a cycle guard because rerun_of is a string read out of a
    record and records travel. And the headline took its numbers from the tail
    totals, which belong to the re-run, so one page read 2 passed · 0 failed
    above a tag table summing to eight and a sibling JUnit saying tests="8".
    The composed stream now declines those totals rather than inventing new
    ones, so the headline counts the scenarios actually rendered.

  • Run metadata reaches the two ADR-0020 §5 consumers that never received
    it.
    The GitHub job summary — named in the ADR, and the page a CI reader
    actually opens from the job — carried none, so the commit under test was in
    the record and the HTML report but not there. And diff --format json
    carried env but not metadata while diff's human output printed
    metadata differences, leaving the machine surface a CI gate reads missing
    exactly the context the ADR was written to provide. Still handed over, never
    harvested.

  • The ADR-0002 grammar guard can now see the shapes the ADR names. The
    amendment claims the core's hurl vocabulary is closed and pinned; the guard
    classified four shapes, and method lines — one of the four boundary
    recognisers the amendment's own Measurement section names — was not among
    them. Teaching it surfaced one unenumerated token immediately:
    GET ${url:base}/PATH, sitting in bind.rs in the same literal as the
    already-pinned HTTP 200. The ADR's table and the pinned set both now carry
    it, and the vocabulary is thirteen literals rather than twelve. The scan also
    stopped truncating at a file's first #[cfg(test)] mod and now excises every
    test module: production code placed after one was silently unscanned, and two
    core files already carry a second test module.

  • --rerun on a truncated record no longer reports success over a suite
    that never ran.
    Record::scenarios is built from scenario_finished
    events alone, so a run killed mid-flight — SIGKILL, OOM, a full disk, a
    container eviction — leaves its unreached scenarios absent rather than
    recorded. The candidate list built from such a record named nothing, the
    "no failures" branch fired, and --rerun exited 0 having executed no
    scenario at all. explain saw the truncation the whole time; --rerun did
    not, and CI is exactly where truncation happens. The same class as the
    cancelled-run bug fixed in 0.14.0, which this code's own comment describes.

    A truncated base inverts the question: not "what did the record say to
    re-run" but "what can the record prove finished" — everything else in the
    selected front runs, announced with a warning naming the truncation. That
    distinction now lives in a RerunFilter predicate rather than a list,
    because only the record reader knows which of the two questions applies.

  • diff no longer reports a scenario skipped in both runs as "now skipped
    … (was passing)".
    Both halves were false — it did not become skipped, and
    it was not passing — and it fired for every @skip scenario on every diff,
    including two runs of an unchanged suite, handing --format json consumers
    the same wrong pair. The bucket exists for transitions (ADR-0019 §7); the
    guard makes that true of the code and not only of its name.

  • A tab in a bound value is refused where every other control character
    already was.
    The lower-time guard exempted \t, which hurl's variable:
    grammar rejects like any other control character, so exactly one character
    kept taking the late path the guard exists to close — dying as
    emit::invalid_artifact against generated text the author never wrote,
    rather than as a refusal naming their own bind:.

  • --shard-weights no longer piles every zero-cost scenario into shard 0.
    Costs are whole milliseconds, so anything sub-millisecond stores as 0 —
    routine for a fast suite — and adding 0 never moved a shard's load, so
    shard 0 stayed the minimum forever. An all-zero weights file put the entire
    suite in one shard and left the others selecting nothing: the flag doing the
    exact opposite of its purpose, silently, with the partition still exact so
    nothing complained. Assignments are now a tie-break alongside load, which
    also gives the right answer when weights genuinely cannot separate
    scenarios: equal cost, equal share.

  • A disk filling mid-run now reaches the exit code. A stdout that was
    already broken at start has failed loudly since the correctness series — but
    the human report's own writes go through the console reporter, which
    swallows write errors (a reporter cannot report its own channel dying), so a
    disk filling during the run truncated the report while the run still
    exited by its verdict. The Tee under the reporter is the last place the
    failure is visible; it now latches the same stdout-failure flag outln!
    uses, and the exit funnel turns lost output into exit 3. Same closed-pipe
    exemption as ever — proef … | head is the reader ending the pipeline, not
    a failure — and a stderr console (machine mode) does not claim stdout
    failed. Pinned by a three-case test, mutation-checked.

  • The complexity guard added moments earlier was itself flaky, and now runs
    alone.
    It shipped in the ordinary suite on the reasoning that a ratio
    cancels out runner load. Measurement disagreed on its second full-suite run:
    2.05× isolated, 3.09× under nextest's full parallelism, against a bound of
    3.0. The larger input has the larger working set, so memory-bandwidth
    contention penalises it more than the smaller one — the ratio drifts rather
    than cancelling, and interleaving the samples cannot fix a systematic effect.

    nextest's test-groups bound concurrency within a group and do not isolate
    one from the rest of the suite, so the only mechanism that actually delivers
    isolation is #[ignore] plus a dedicated invocation: a CI step of its own and
    just perf. The samples are interleaved as well, which removes the one skew
    that ordering alone creates.

    TESTING-STRATEGY.md §7 previously asserted the opposite in as many words —
    that a ratio "survives a shared runner" — and is corrected with the numbers.
    The claim was reasoning, not measurement, which is the failure this whole
    section of the changelog exists to record.

Internal

  • A fixture that spells the record by hand can no longer drift off the
    schema.
    explain's truncated-record test wrote its stream as three JSON
    string literals, and all three had drifted: a scenarios count on the head,
    a schema on the body events, a line on the close. Event carries none of
    them. Nothing failed and nothing could — the reader has no
    deny_unknown_fields, so a stale key parses cleanly and is dropped, and a
    fixture built to assert "a record holding one passed scenario" was
    three-quarters describing a format proef has never written. It is typed now,
    through the helpers its two neighbours already use.

    The class is closed by a sixth source_guards.rs rule: every string literal
    in the workspace that parses as a JSON object tagged event must
    deserialize as an Event, and every key in it must matter — a key is
    phantom when deleting it yields the same Event.
    Inertness rather than an
    inventory, so it stays correct through renames, #[serde(default)] and
    skip_serializing_if, none of which a key-set comparison survives.
    Substring assertions against records proef actually emitted
    ("event":"run_finished","passed":1) are skipped by construction — they are
    not objects, and they check the opposite direction.

  • Each doc check now lives in the half of the gate that its own rule
    names.
    tests/docs.rs holds the checks that need a built binary (they ask
    clap, rather than parsing help text into a model that could drift);
    xtask docs-check holds the ones that read files. The changelog-heading
    check added moments earlier read one file and parsed headings, so it sat in
    the wrong half — and the cost was concrete rather than tidy: it never ran in
    the fast doc-only CI step, only under a full nextest that had to build a
    binary it did not use.

  • A PR that changes source now has to record itself. RELEASING.md has
    always said that every landed change adds an [Unreleased] line in the
    commit series that lands it, and nothing checked it — this very entry is the
    one that was missed. Measured before being written: across the previous 21
    source-touching merges the rule would have fired exactly once, on exactly the
    commit that broke it, so the check earns its place by count rather than by
    argument.