Skip to content

Releases: emrecdr/proef

proef v0.19.0

Choose a tag to compare

@github-actions github-actions released this 11 Sep 08:29
21b0e3a

Fixed

  • The Homebrew formula installs the man page and the shell completions. Its
    def install was bin.install "proef" and nothing else, so from 0.16.0 —
    the release that started shipping proef.1 and five completions/ files in
    every archive — until 0.18.0, brew install proef gave no man proef and no
    tab completion, while binstall and a direct download gave both. The formula
    is a heredoc inside release.yml and the archive is staged in a different
    job, so nothing tied the two together and no gate could see the gap; the
    render step now fails if the archive lacks a file the formula installs, and
    the formula's own test do asserts the man page and completion landed.
    Takes effect on the next tag: a tag runs the workflow from its own commit.

  • --dry-run refuses a file,…; asset that is not there. A suite whose
    asset had been deleted reported dry-run OK, and the failure arrived later
    from a different command, against a live backend — from the one gate CI runs
    before standing an environment up. Whether an asset resolves is statically
    knowable, so it is answered there now. The checker is staging's own
    (assets::resolve_assets, split out of stage_assets) rather than a second
    walk over the same artifacts, so validation and the run cannot disagree; the
    message and the diagnostic code are the ones a run already gave.

  • file, inside a JSON or assertion body is no longer mistaken for a file
    asset.
    The emitter found the files an artifact reads by scanning its text
    for the literal file, and a closing ;, so a request body containing that
    substring — {"note": "see file,notes.txt; for details"} — produced a
    phantom asset, and staging then failed the run over a file the request never
    reads. The claiming engine now reads its own AST, where a body reference and
    six characters of prose are different things.

Breaking

  • proef_core::emit::emit takes the registered step kinds, and
    StepKindSpec gains an assets hook.
    Asset recognition was hurl's body
    grammar living in proef-core: emit::file_refs_in scanned for the literal
    "file,", which ADR-0002's amendment forbids and — worse — which the guard
    pinning that amendment could not see. engine_grammar_kind classifies
    fences, HTTP, [Section] headers, method lines and key: value options; a
    body constructor is none of those, so the literal was never sanctioned and
    never reported missing. The ADR's own measurement said thirteen literals; it
    was fourteen.

    The scan moves behind the seam as StepKindSpec::assets, the fourth
    engine-contributed hook beside validate, fragments and options, and the
    guard gains a body arm so the shape is classifiable whether or not anything
    currently uses it. emit() takes &[StepKindSpec] to reach it; FrontEnd
    carries kinds beside the kind_to_engine table it is built with, which
    registry already documents as a pair that must not be re-derived
    separately. emit::file_refs_in is gone.

Internal

  • A fragment's assets stage from where its file was read, not from where its
    name points.
    AssetRoots::source_dir rebuilt a fragment's directory by
    splitting file.hurl#name and joining the file half onto the project root —
    the naming boundary run backwards, without the canonicalize fallback that
    boundary carries precisely because a lexical-only version already shipped a
    bug (a suite reached through a symlink silently failed to match, R11-9). The
    two agreed only because both were seeded from config.root() and discovery
    walked from that same root, so only the lexical case was ever exercised, and
    nothing made them stay inverses. The corpus reader now records the directory
    it read each file from (front::CorpusDirs, carried on FrontEnd beside
    kinds), and staging looks it up — the fragment-side twin of what
    LoadedFeature::read_from already does for features, so both halves of the
    naming boundary are one-way in the same way. AssetRoots loses its project
    field and build_specs its project_root argument: with nothing to
    recompute, the project root is no longer staging's business.

  • --rerun reads its base record once. It called record::read_events for
    the JUnit overlay and then record::rerun_candidates, which read and
    deserialized the same events.jsonl a second time — two full passes bounded
    only by the 256 MiB record ceiling, over a file another process may still be
    writing. rerun_candidates now takes the &[Event] its caller already
    holds, which is the rule read_record's own documentation had already
    stated for exactly this case. The read error is handled once as well: the
    first call swallowed it with .ok() and the second rediscovered it a line
    later.

  • The one doc check that reads only files now runs in the half that reads
    files.
    no_current_behaviour_doc_spells_a_format_as_an_output_path lived
    in tests/docs.rs, whose stated charter is the checks needing a built
    binary
    to ask clap — this one only scans markdown, so it never ran in the
    fast doc-only CI step. It is now xtask docs-check's
    check_output_path_spelling, reusing living_docs() instead of carrying a
    second directory walk. Its allowlist-shrink guard got stricter on the way:
    it counted ADRs into the same total, so a renamed entry could be masked by
    docs/adr being larger than the shortfall — which is the one failure that
    guard exists to catch. All three paths were checked by mutation: a stale
    spelling planted in an allowlisted doc, one planted in an ADR, and an
    allowlisted doc renamed away.

Documentation

  • The worklist stops contradicting what shipped. Three entries in
    OPEN-FINDINGS still called CTRF declined or its trigger unfired — the
    2026-08-31 external re-test, the RF audit's deferred list, and R3-5 under
    "deferred, with the trigger named" — for the eight days after --ctrf
    actually shipped (#160). R3-9, four bullets below R3-5 in that same list, was
    annotated the moment it shipped — the convention the three missed. Two more claims
    had outlived their facts: the shipped-changelog duplicate headers (no release
    carries one now, and check_changelog_kinds fails if one returns) and the
    machine-side note about Homebrew's Rust shadowing rustup. Filed at the same
    time: a_second_interrupt_hard_exits_with_130 failed once on Linux CI and
    passed on a re-run of the same commit, so the evidence, the mechanism and the
    fix shape are written down instead of left to the next re-run. And the stance
    that a scenario-level @retry is deliberately absent — retry-until-green
    hides a one-in-four defect 99.6% of the time — is stated in
    TESTING-STRATEGY §5, which the worklist asked for and nobody had written.

  • The runbook records that the registry skips three versions. 0.15.0–0.17.0
    were tagged and GitHub-released but never published, so crates.io moves
    0.14.0 → 0.18.0. Noted in RELEASING.md so the gap does not read as a failed
    upload. The long-standing homepage question in OPEN-FINDINGS is also
    resolved: the field reached the registry with 0.18.0, exactly as that entry
    predicted; documentation remains unset and still open.

  • The release history records every release again. RELEASING.md's History
    section carried no entry for v0.16.0 or v0.17.0 and filed v0.15.0
    between v0.13.0 and v0.14.0; the order is repaired and all three versions
    are present, v0.18.0 included. The corpus also stops calling the 0.18 series
    unreleased, and an IMPROVEMENT-PLAN pointer into CHANGELOG [Unreleased]
    now names the releases that actually carried the work — [Unreleased] has
    been cut several times since that sentence was written.

proef v0.18.0

Choose a tag to compare

@github-actions github-actions released this 09 Sep 10:38
e8a120c

Added

  • SIGTERM and SIGHUP now take the graceful path (ctrlc's termination
    feature): a CI job timeout or docker stop cancels the run — in-flight
    batches finish, the rest record as skipped, teardown runs, the reports are
    written, and the record closes with a cancelled run_finished — where it
    used to kill the process mid-write and leave a truncated record with no
    tail. A second signal still hard-exits 130 (the handler carries no signal
    identity, so the code is 130 for every second signal). Pinned by
    sigterm_cancels_gracefully_and_the_record_completes and — for the first
    time anywhere — an exit-130 assertion,
    a_second_interrupt_hard_exits_with_130.

  • test --format json and explain --format json now report warned and
    cancelled.
    A warned scenario (an optional: step failed, or a saveAs: global promotion was refused) folded into passed, and cancelled — in
    the record's run_finished — was surfaced by neither, so a script could
    not tell a spotless run from one with warnings, nor a complete run from a
    cancelled one, and the two JSON surfaces disagreed on how to say "did not
    finish" (0.18 survey). Both keys are additive and always present.
    warned also becomes visible in JUnit (a <system-out> note, the status
    stays success since JUnit has no warned) and CTRF (an extra.warned
    flag) — it was previously visible only in the HTML report.

  • A tag that looks like a reserved one but is not exactly it now warns
    (proef::tags::reserved_tag_typo). @quarantined, @skipped, @Skip
    matched no reserved tag and silently did nothing — a scenario the author
    believed was quarantined gated the build. The warning names the spelling
    it likely meant, tuned to catch the real typos without firing on
    legitimate short tags (ship, slip, step).

  • proef flaky gains the 2026-field statistical guards (0.18 survey §6),
    each a pure fold over the JSONL history already retained — no new state, no
    gating mode (advisory stays the design):

    • A minimum-sample floor (--min-samples / [flaky] min-samples,
      default 10): below it a scenario is insufficient-data rather than
      classified, because a verdict on thin data is worse than none.
    • Hysteresis (--recovery-runs / [flaky] recovery-runs, default 5): a
      flapping or latent scenario holds its flag until it earns a trailing clean
      run, so it cannot oscillate flaky↔healthy between adjacent runs.
    • An environment-outage guard (--outage-rate / [flaky] outage-rate,
      default 0.8): a run where over this share of suite scenarios failed is an
      environment incident, not evidence about any one scenario, and is excluded
      — so a single fixture or staging outage cannot mark the whole suite broken.
    • An input-fingerprint equivalence class — the default key. Each run
      writes an inputs.json sidecar carrying a hash of what it executes
      (feature sources + loaded macros/fragments + the resolved
      ${url:…}/${vars:…} scope), so a pack, feature, or proef.toml edit
      correctly ends the comparison window instead of silently mixing runs of
      different inputs. It is a proef-computed fact about proef's own inputs,
      not harvested from the environment (ADR-0020 unchanged — git-commit
      grouping stays handed-over via --meta commit=… and proef flaky --by commit). broken≠flaky, transition-counting, and the quarantine
      lifecycle were already present and are unchanged.

Fixed

  • A run-record write that fails now reaches the exit code. The JSONL
    reporter deliberately swallows write results (a reporter cannot report its
    own channel dying), and events.jsonl was handed a bare File — so a disk
    filling mid-run truncated the record while the run still exited by its
    verdict, the exact class the v0.6–v0.8 series closed for the console. The
    record's writer now latches its first failure (one stderr line, run
    continues) and the exit funnel turns it into a system error, the same shape
    as the stdout latch and the JUnit-write fold — unified in one pinned
    function, escalate_environment_failures. run.log's mirror keeps its
    own contract (creation is warn-and-continue, so a mid-run failure warns
    once and leaves the verdict alone — previously it was silent).

  • The GitHub step summary can fail again. It was the only CI sink that
    couldn't: a failed open or write vanished while JUnit and CTRF failures
    re-classify the exit — so the page a reviewer actually reads could be
    missing on a green exit. write_github_summary now returns the error and
    the caller folds it into the same reports_failed path as its siblings.

  • A custom --run-id no longer collapses the JUnit report identity onto
    the nil uuid.
    ADR-0021 made non-uuid run ids first-class, but the report
    uuid was parse_str(...).unwrap_or(nil) — every --run-id ci run emitted
    00000000-…, colliding in any consumer keyed on it. A non-uuid id now
    derives a stable UUIDv5 from its bytes (a uuid id passes through verbatim).

  • The interrupt window and the interrupt's own words. The handler is
    installed at the top of execute — before the front end, the run dir and
    the record exist — so no startup window takes the process default any
    more. Its installation failure is a printed warning (it was silently
    ignored, unlike --watch's handler). The second-signal path no longer
    prints before exiting: the print took stderr's lock, which a worker
    blocked on a full pipe can hold, wedging the escape hatch behind the very
    stall it exists to escape. And the teardown notice said "Ctrl-C again to
    skip" when a second interrupt actually hard-exits dropping every report —
    it now says what happens.

  • Asset staging no longer depends on the working directory. A feature's
    file,…; assets were resolved by joining its portable name against the
    cwd — but a name's anchor (the project root, or the caller's own typed
    spelling) is not in the string, so a typed-absolute or config-written
    suite path run from any subdirectory failed staging with exit 2, blaming
    the author for a correct file (the feature-side twin of OPEN-FINDINGS
    H5). The resolved discovery path now travels beside the name
    (LoadedFeature::read_from) and staging resolves beside the file the
    parser actually read — the H5 prescription, applied to the feature side.
    Reproduced before the fix and re-verified after, from a subdirectory,
    against the reference corpus; a new integration test pins a project under
    a path with spaces and non-ASCII segments, which nothing in the suite had
    ever exercised.

  • --sarif line numbers survive a cd, and byte-match the parser. The
    SARIF writer re-read each source from disk by its portable name to count
    lines — from any subdirectory every read failed and startLine silently
    vanished, annotating nothing; the re-read could also disagree with the
    span by exactly the parser's normalization. Lines now come from the
    diagnostic's own carried source text — the same normalized bytes the span
    indexes. (On Windows, an absolute out-of-project uri also spells its
    separators as a URI requires.)

  • Staging's two symlink edges. An existing symlink at a staging
    destination was written through — fs::copy follows links, so the
    bytes landed wherever it pointed, outside the root built to contain
    them; it is now replaced. A source symlink stays followed, deliberately:
    stock hurl follows it too, and refusing would break the dual-runner
    rule (the module doc now says so).

  • Asset names that are one file to the filesystem are refused. The
    duplicate-name guard keyed on the raw reference string, so Data.json
    and data.json — one file on macOS and Windows — silently last-writer-won,
    the very overwrite the per-scenario root was built to end. The check now
    runs on the canonical path the copy actually landed on, which is exact on
    every platform: a case-sensitive volume keeps both files legitimately, and
    nothing fires.

  • Artifact slugs cap at 120 bytes. The slug flattens the feature's whole
    directory path into one filename component, and assets/<slug>/ repeats
    it as a directory — so path depth became filename length, and a deep tree
    or a long scenario name (multi-byte scripts at a quarter of the visible
    characters) sailed past NAME_MAX and failed the write. Over the cap, the
    tail is a hash of the whole uncapped slug, so two names differing only
    past the cut still name two artifacts; every slug the existing corpus has
    is under the cap and unchanged byte-for-byte.

  • The ADR-0007 budget family is closed over its inputs, and bounded as a
    product.
    [Options] max-time: was read by the budget calculator (as
    the entry's timeout) while invisible to the lint — max-time: 100000h
    was lint-clean and produced a multi-year watchdog budget; it now carries
    the duration cap, and a test pins the rule the hole broke (every option
    the budget reads must be one the lint can see). retry-interval: — the
    one uncapped multiplicand — carries the cap too. And because individually
    capped values still compose into an unbounded product (retry: 10_000 ×
    a 30 s timeout is ~83 lint-clean hours, saturating to Duration::MAX,
    whose deadline addition panicked as a phantom "scenario thread panicked"
    fault), the computed batch budget now clamps to an absolute four-hour
    ceiling and the dispatcher's deadline arithmetic can no longer overflow.
    ADR-0007 carries the amendment.

  • [http] timeout-ms = 0 is refused. libcurl reads zero as no
    timeout, so the value opted a suite into exactly the unbounded hang the
    default exists to defend against — while reading like "immediately".
    Exit 2, in whichever table it appears.

  • Every sink that renders run values now routes identities through the
    secret masker.
    The event stream masks scenario, file, ...

Read more

proef v0.15.0

Choose a tag to compare

@emrecdr emrecdr released this 09 Sep 09:09
54e85ab

Published retroactively on 2026-09-09. This version was released on
2026-08-25 — the release commit is on main and the changelog section below
was cut then — but its v0.15.0 tag was never pushed, so the release
pipeline never ran for it. This release therefore carries no binaries,
checksum sidecars or build attestations
; every other release has all
three. Install 0.15.0 from source, or use a later version — 0.16.0
superseded this one six days later.

Breaking

  • A quarantined test-failure reaches JUnit as <skipped> with a message,
    not <failure>
    — Jenkins marked builds UNSTABLE while proef exited 0;
    every dashboard now says what the exit code says (ADR-0019). Library:
    ScenarioSpec gains skip, ScenarioOutcome/ScenarioRun gain
    reason, Event::ScenarioFinished gains additive reason,
    write_junit takes the non-gating list.
  • --shard assignments re-deal: the hash gained a mixing finalizer. Raw
    FNV-1a's low bit is the XOR-parity of the input bytes, so a scenario named
    after its feature file — the commonest Gherkin convention — collapsed to
    one shard at N=2 and left odd shards empty at N=4, silently (the empty
    shard exits 0). shard_bucket now finalizes with Murmur3's fmix64; every
    scenario re-buckets, so all jobs of one matrix must run the same proef
    version (already true in practice). Round-18 finding, reproduced and
    mechanism-verified before fixing; the balance test gained the
    name-mirrors-file corpus it was structurally blind to.
  • Tag atoms glob. * and ? in a --tags / [run] exclusive-tags atom
    are now anchored wildcards (@FRD-* selects the family; ? is one
    character) — previously they were literal characters that silently matched
    nothing, the trap this closes. Metacharacter-free atoms are bit-identical
    to before, property-pinned. Case stays sensitive.
  • JUnit test identity is classname + name. classname carries the
    feature file, name the scenario alone; the old single name embedded
    file:line, so an edit above a scenario re-identified every test below it
    in Jenkins history and GitLab's MR diff. Anything keyed on the old
    file:line name strings must re-key. The suite skipped count is now
    spelled skipped (was disabled, which no consumer reads).

Added

  • [tag-links] turns tag cells into tracker links (RF's
    --tagstatlink, reduced to one mechanism): tag glob → URL template with
    {tag} substituted, honored by the HTML report's by-tag table and the
    GitHub summary; the pattern language is the same anchored glob --tags
    uses. Library (Breaking): render_html takes the link map;
    tags::atom_matches_public exposes the one matcher.

  • --console dotted|quiet (RF wave 3): one glyph per scenario (.
    pass, F fail, s skip, w warn — lowercase is non-gating, the
    pytest/RF convention, flushed per glyph, wrapped at 80) or just the frame.
    Purely presentation: the record, every report, the post-pool failure
    details and the exit code are identical in every mode; run.log mirrors
    the console verbatim, dots included — events.jsonl is the full truth.
    Library (Breaking): ConsoleReporter::new takes a ConsoleMode.

  • A --rerun now produces the one JUnit and the one report that cover
    the whole suite
    (E2's rerun half; Robot Framework's rebot --merge
    shape, done as composition): the run head records rerun_of, the JUnit
    carries the base's not-re-run scenarios as ordinary testcases, and
    proef report overlays the base into a merged page (banner named, base
    timestamps stripped so timelines never mix, rotated-away base degrades
    loudly). Exit code and totals stay the rerun's own.

  • --meta key=value and [meta]/[env.<name>.meta] record explicit run
    metadata
    (ADR-0020, RF wave 2): commit, build URL, team — recorded in
    the run head, shown by the HTML report, GitHub summary, explain,
    diff (which now also warns on cross-env comparisons) and the
    --output json body (additive keys). The active --env profile name and
    the --shuffle marker ride the same head. proef never harvests: no git,
    no hostname, no CI env sniffing — the shell harvests, proef records.
    Everything passes the sink-boundary mask, keys and values both. Library
    (Breaking): RunRecord::open and exec::execute take the head inputs.

  • Per-tag verdicts in the HTML report and the GitHub summary (RF wave
    2): tags now reach the record — additive tags on scenario_finished
    (finished-only: the cancel-skip path emits no start), additive
    exclusive on scenario_started (closes R11-6, the scheduler's own
    bool) — and both reports roll them up per tag (suite-only, Warned counts
    with passed). Requirement-tagged suites (@FRD-3.1) get their
    traceability matrix for free. Tags are deduped at the one accumulation
    point (first occurrence wins); the quarantine list is now derived from
    the outcomes' own tags — one owner, same behavior, pinned by the exit
    suite. Library (Breaking): ScenarioSpec/ScenarioOutcome gain tags.

  • @skip and @skip:<reason> park a scenario visibly (ADR-0019, RF
    wave 2): counted in every total, reasoned in the console, JUnit, TAP, the
    record, the HTML report, explain and flows --output json; the harness
    maps it to libtest's ignored flag. All-selected-skipped exits 0; the
    empty-selection refusal stays exit 2. --tags "not @skip*" unselects both
    spellings; an authored skip is never re-queued by --rerun, and diff
    gives skip transitions their own bucket instead of reading them as fixed.

  • flows shows the feature description. The prose block under
    Feature: was parsed and then dropped — the one paragraph written for
    exactly the reader flows serves never reached them. Human output prints
    it under the feature header; --output json rows gain
    featureDescription: string|null (additive). Library: FeatureFile gains
    description.

  • --shuffle re-deals the execution order, seeded by the run id — one
    determinism knob for order and fakes alike, so --shuffle --run-id <id>
    reproduces an order-dependent failure exactly (Robot Framework's
    --randomize, minus the parallel seed it threads separately). Applied
    after --shard, so membership never moves; under --watch every unpinned
    rerun re-deals, deliberately. The permutation is version-stable and
    pinned. Recording a shuffled marker in the run head is deferred to the
    planned RunStarted additions (env/metadata), one wire change instead of
    two.

  • The failing step's reproduce: curl … reaches the record. The engine
    always computed the redacted curl and the live console always printed it —
    and the record dropped it, so explain and the HTML report knew less than
    the console did. StepFinished gains additive reproduce_hint (absent on
    passing steps and every pre-field stream); explain and the report print
    it; the sink-boundary mask covers it like detail.

  • README documents every flag the binary exposes, enforced. v0.14.0
    shipped --shard and --max-fail with no README mention; the docs gate
    gains the flags direction (same vacuity guard as the command half), and the
    measured gap — those two plus schema --add-to — is closed.

  • JUnit carries what GitLab and Jenkins actually read (R3-6, specced from
    GitLab's parser docs and Jenkins' SuiteResult.java): file on each
    testcase (GitLab source linking), time on suite and root. timestamp and
    hostname stay absent deliberately — ignored or substituted by both
    consumers, and a hostname would undo R12-1's provenance fix.

  • The docs corpus is a website: https://emrecdr.github.io/proef/. mdBook
    renders docs/ on every push to main that touches it; the nav is
    docs/SUMMARY.md, which the existing docs gates link-check like any other
    doc, and the pages workflow refuses a corpus doc that is not on the site.
    The crate homepage points there from the next release.

Fixed

  • A failure detail is bounded before it reaches any sink. hurl's rendered
    assert error quotes the actual response, so a failed assert on a large body
    rode full-size into the record, JUnit, the HTML report and the GitHub
    summary at once. The engine now middle-cuts past 40 lines / 8 KiB with a
    marker naming the elision; the artifact pointer survives outside the cut,
    and the full output is one re-run away (Robot Framework's 40-line rule,
    adopted at the boundary where all sinks are covered at once).

  • The machine-body contract closes its last two paths: an empty selection
    (--scenario/--tags matching nothing — loud exit 2 by design) and a
    corrupt global-state file both emitted zero stdout bytes under
    --output json.

  • Identical errors collapse like identical warnings — a broken macro
    usually fails to lower everywhere, so the error wall was the more common
    fifty-block wall; distinct errors still render separately, and SARIF keeps
    every site.

  • Injected [Options] lines respect every section-ending shape. The
    section-end move covered one shape of five: an unfenced JSON/XML body after
    an author [Options] swallowed the injected lines into invalid hurl (exit
    2 on input that worked before), and an entry with an author section but no
    response line leaked its pending lines into the next entry, where hurl
    parsed retry: as an HTTP header and the artifact validated green. The
    section now ends at the first line that could not sit inside it.

  • A # inside a bind: value no longer hides the reads after it. The
    template probe parsed the value in an unquoted position where # opens a
    comment; it now probes the quoted variable: position bake actually
    injects into, so "{{a}} # {{b}}" reports both.

  • A setup that fails to load still emits the machine body — the last
    terminating path returning zero stdout bytes under --output json.

  • **SARIF keeps ...

Read more

proef v0.17.0

Choose a tag to compare

@github-actions github-actions released this 06 Sep 14:11
e5f8d89

Added

  • [http] cookie-store = false runs the whole suite cookie-less — hurl
    8.0's --no-cookie-store, surfaced through the table built for exactly this
    class of setting. No Set-Cookie is retained and none is replayed, which is
    how a stateless API is proven stateless: the fixture-backed test is green
    only because its steps assert the 403 a missing session cookie earns.

    This is the one [http] key with no per-entry [Options] spelling at
    all
    (OptionKind has no cookie variant — verified against the enum), so
    run-wide is not a compromise but the only place it can be said. With the
    store off, the engine also skips both halves of the batch-split cookie
    round-trip: hurl reads a cookie_input_file only when enabling the engine,
    so injecting one would be silently ignored — and there is nothing to write.
    hurl's own FIXME (a handle once given cookie storage cannot lose it) never
    reaches proef, because run_entries builds its client per call (TECH-SPEC
    §5) — a handle never transitions on → off.

    Breaking (library): HttpDefaults gains the cookie_store field, so a
    struct-literal construction needs the new line (..Default::default() sites
    are untouched, and an absent [http] cookie-store key changes nothing).

  • --ctrf <path> — the run's verdicts as a CTRF report. CTRF
    (https://ctrf.io) is the emerging JSON successor to JUnit XML for CI
    dashboards, and it models in the schema what JUnit can only smuggle
    through extensions — which is exactly the data proef already tracks: a
    pass-after-retry carries flaky, retries, and retryAttempts listing
    the real failed attempts with their (redacted) messages; every test carries
    its tags and file path. One serializer off the same fold as JUnit, so the
    two files cannot disagree — most visibly for a quarantined failure, which
    both report as skipped with a message (ADR-0019), because a dashboard
    reading "failed" beside exit 0 would contradict itself. A User/System
    fault stays failed even under a quarantine tag: quarantine is for flaky
    tests, not broken input.

    The R12-3 contract applies from day one: a [run] setup abort still writes
    the file, carrying the setup scenario itself — a job gating on the report
    must never see no file at all. The schema's required wall-clock
    start/stop are measured at the CLI edge like every other clock read
    (ADR-0015); the sans-IO core and the JSONL record are untouched — the
    record remains the only record (ADR-0008).

  • The HTML report answers "what is slowest". After "what failed", it is the
    question a test report is most often asked, and the page could not answer it:
    the timeline showed that workers were busy, never which scenarios to
    attack. Every number needed was already in the fold.

    A ranked section, slowest first, each row linking to its own block, with the
    heading reporting the share of run time the listed scenarios account for —
    "3 of 40 · 71% of run time" is a decision, where a column of durations is
    homework. Capped at eight: a ranking long enough to scroll has stopped
    answering the question.

    Cost is the sum of a scenario's step durations, the same definition
    timings.json uses for shard weights — one notion of what a scenario costs
    across the whole tool. Not the wall-clock span, which includes time waiting
    for a worker: a property of how the run was scheduled, and not something the
    reader can go and fix.

    Absent when there is nothing to rank — fewer than two timed scenarios, or a
    record with no injected durations at all.

  • --shard-weights balances a shard matrix by measured duration. --shard
    assigns by a frozen hash, which guarantees that adding one scenario never
    re-buckets the others but cannot balance by time — and a CI matrix finishes
    when its slowest shard does, so a count-split routinely leaves runners idle.
    Every run that reaches its suite now writes a small timings.json into its
    run directory; CI archives that one file and each matrix job points
    --shard-weights at the same copy.

    The obvious design is silently wrong, and the module says so at length.
    proef already retains records carrying every step's duration, so "weight by
    the newest local record" looks free. But matrix jobs run on different
    machines
    , each with its own (usually empty) runs-dir — every job would
    compute a different weight table, therefore a different assignment, and
    scenarios would run twice or not at all while the suite reported green.
    Nothing about that announces itself. One named file shared by every job is
    what makes the split a pure function of (selected scenarios, that file).

    Two rules place scenarios and they partition rather than compete: a
    scenario the file mentions goes through longest-processing-time-first
    placement, and one it does not mention falls back to the frozen hash. So a
    test added after the timings were captured still runs exactly once. That is
    pinned by a test that runs a whole three-way matrix — with a weights file
    covering only five of nine scenarios, so both rules are exercised at once —
    and asserts set equality both ways; mutating the placement by one bucket drops
    two scenarios and the test names them.

    The weight is the sum of a scenario's step durations, not its wall-clock
    span. The span includes time spent waiting for a worker, which is a property
    of the run's scheduling rather than of the scenario, and feeding it back would
    let one crowded run's queueing distort the next split.

    What this gives up is exactly what hash mode was chosen for: a balanced split
    is not stable under insertion. That is what balancing means, which is why the
    flag is opt-in. A missing or malformed weights file is exit 2 — falling back
    silently would hand back the unbalanced split the flag was passed to avoid.

  • The editor tells proef's two variable tiers apart. A pack's hurl: | block
    is the centre of the authoring experience and, to every editor, a plain YAML
    scalar — inside which ${…} (resolved at lower time, by proef, before any
    request exists) and {{…}} (resolved at run time, by hurl) look
    identical. That distinction is ADR-0005's whole model and the thing authors
    most often get wrong, and no generic grammar can see it: a YAML highlighter
    sees a string, and a hurl highlighter never runs because the block is not a
    file. proef is the only party that knows.

    The server now answers textDocument/semanticTokens/full, lighting ${…} as
    macro — a substitution performed before execution, which is what a macro
    is — and {{…}} as variable. Both are coloured differently by every
    mainstream theme, so it works without anyone configuring anything. The $${
    escape stays dark, because telling an author proef will substitute text it
    will in fact leave alone is worse than no highlighting.

    The ${…} scan is proef_core::resolve::reference_spans, walking the same
    first_reference the resolver itself uses — a second implementation of the
    escape rule would drift, and the drift would show as an editor confidently
    colouring literal text. The {{…}} scan lives in proef-lsp rather than
    core, because that spelling is the engine's and ADR-0002's amendment is that
    engine syntax does not accumulate in the core.

    Collapsing the seven-arm request dispatch behind a local macro came with it:
    the chain crossed clippy's line limit the moment an eighth feature landed, and
    the honest fix was to stop repeating an identical frame seven times rather
    than to suppress the lint that noticed.

  • The linear-validation claim is now a test, not a sentence. #138 made pack
    validation linear and recorded the result as a shape: "the curve changed
    shape — 4× per doubling before, ~2× after"
    . That number lived only in the
    changelog, where nothing could re-run it — so a future span locator scanning
    the whole pack file again would have restored the quadratic behaviour
    silently, a regression that costs seconds rather than correctness and which no
    gate measured.

    The guard asserts the ratio between 1000 and 2000 macros, because the
    claim is a ratio. It observes ~2.05× against a bound of 3.0; mutating
    locate::MacroIndex to re-index per lookup — the exact pre-#138 shape —
    measures 4.01×, matching the changelog's own prediction of 4× and turning
    a 0.4-second test into a 73-second one. The failure message names the cause
    rather than reporting a number.

    A ratio rather than a benchmark, for a reason now written into
    TESTING-STRATEGY.md §7: load on a shared runner inflates both measurements
    together and cancels, where an absolute threshold has to be loosened until it
    means nothing. iai-callgrind would be the better CI gate — instruction
    counts ignore runner noise entirely — but it needs valgrind, so it would be a
    gate the maintainer cannot reproduce on macOS; criterion and divan sit in
    the same noise regime as this test while adding a dependency tree to a
    workspace that audits every edge. No new dependency was added.

  • Every diagnostic code is now named by a test, and a guard keeps it that
    way.
    DIAGNOSTICS.md calls codes "a contract: they never change meaning".
    Twenty-three of seventy-five had nothing holding them to it — reachable in
    production, documented, exercised by nothing at all: not a seeded corpus
    directory, not a unit test, not even an assertion on their message text. They
    existed only at their definition site.

    The catalogue itself was found exactly honest — 75 codes defined, 75
    documented, and its corpus column matched disk in both directions with zero
    drift. The gap was never documentation; it was that a documented promise had
    no enforcement.

    Nineteen new tests close it, each reaching its code through a real path rather
    than constructing the diagnostic directly. Two ...

Read more

proef v0.16.0

Choose a tag to compare

@github-actions github-actions released this 31 Aug 12:48
3feb209

Supersedes 0.15.0, which was cut (release: v0.15.0, 2026-08-25) but never
tagged or published — its changes are all here, and crates.io goes 0.14.0 →
0.16.0 with nothing skipped.

Fixed

  • The record-size ceiling reached two of its four readers. 0.13.0 bounded
    the run-record read at 256 MiB because records travel — diff reads a
    downloaded baseline, flaky reads every retained run — and the read, the
    line split and the parsed Vec<Event> are resident at once, so a corrupt or
    hostile file was an OOM rather than an error. The bound lives in
    record::read_events, and explain and report each opened
    events.jsonl with a bare read_to_string instead, so neither had it.
    report even used the guarded reader for the base record two dozen lines
    below the raw read of the primary one.

    Both now go through read_events, which returns the parsed events — exactly
    the read-once/parse-once its own comment asked for. A source-scanning test
    makes the next reader go through the same door, the shape this project
    already uses for the raw-print and malformed-plural rules: a guard added in
    one place and left for the next call site to rediscover is how it went
    missing the first time.

Added

  • explain, diff and doctor speak --format json. They were the three
    commands with no machine output, and the three a consumer reaches for
    after a run. A run directory is artifacts/ + events.jsonl + run.log and
    carries no structured summary, so anything analysing a run it did not launch
    — a CI job reading another job's artifact, a script, an agent — had to fold
    events.jsonl itself. That is the fold proef's own two internal copies
    disagreed on three ways before report::suite_totals unified them; handing
    the canonical answer over is cheaper than inviting everyone to re-derive the
    one proef got wrong.

    Each object mirrors its prose field for field rather than modelling a richer
    view — the prose is the contract a reader already knows, and a machine
    surface that says something different is a second answer to one question.
    diff's flaky/slower stay the rendered sentences for the same reason.
    The flag is the existing single-variant json enum the listing commands
    already use, renamed from ListFormat to JsonFormat now that it serves
    non-listing commands too. Machine mode owns stdout: notes whose content the
    object already carries are suppressed rather than repeated on stderr.

    doctor needed a real change to get there — it printed each check as it ran,
    so the verdict was the only thing a caller could see. Checks are collected
    before rendering now, which makes the JSON a second rendering rather than a
    second walk: the failure mode where one surface gains a check the other never
    learns about.

  • --console failed — the full BDD tree, but only for scenarios that
    failed or warned. A clean run prints the run line and the summary; a dirty
    one prints exactly what full would. The gap it fills is the CI one:
    full is a wall of green on a large suite, dotted drops the detail you
    need when something breaks, and quiet drops everything.

    Warned scenarios are shown, which the name does not say and the code
    explains: a warned scenario is one whose optional: step failed,
    RunSummary::passed counts it with the passes, and the summary line has no
    warned column — so a mode that showed only Failed would let a run in which
    something did fail print exactly what a spotless one prints. A fourth
    value on the existing flag rather than a new one.

Fixed

  • cargo deny failed on a yanked transitive crate. rand 0.10.2 resolved
    chacha20 0.10.1, which was yanked from crates.io; the lock now takes
    0.10.2. Not the secret store's copy — chacha20poly1305 pins 0.9.1,
    which is unaffected — so nothing about encryption changed. Found by the
    gate, which is what it is for.

  • proef report -o wrote the author's home directory into the file built to
    be shared.
    With the report inside the run dir the artifact links are a
    bare artifacts/…; with -o pointing anywhere else they were made
    absolute, which resolves only on the machine that produced them — and -o
    exists to put the report somewhere it will be published, which is exactly
    where that path is dead. 0.13.0 scrubbed machine identity out of the run
    record (R12-1); this put it back, twelve times over, in the HTML uploaded
    beside it. The href is now relative to the report, which resolves everywhere
    the absolute one did plus wherever report and artifacts travel together,
    and in the CI shape (-o public/report.html) names nothing outside the
    workspace. The href is built from path components joined with /, not
    from Path::display — Windows renders \, which is not a separator in a
    URL, so a Windows-generated report's links would have been dead either way
    (the absolute path it replaces had the same flaw). A report written somewhere
    sharing no ancestor with the run dir
    still names the directories between them — that is what a correct relative
    path from there is, and it is no worse than what it replaces.

  • The report's --skip colour failed WCAG AA, and every status pill failed
    it in dark mode.
    --skip was the one palette token the dark block did not
    redefine: a grey chosen against #0d1117 (5.48:1 there) left carrying white
    text on white at 3.45:1, against a 4.5:1 threshold — on the status a
    reader scans for after an interrupted run. It is now #59636e (6.11:1).

    Writing the guard rather than the fix found a second defect nobody had
    measured: .pill painted color:#fff on the status colour, and the dark
    palette's colours are tuned as text on a dark ground, so all four dark
    pills sat between 2.52:1 and 3.45:1. The pill foreground is now a palette
    token — white on light, the page ground on dark — putting all four between
    5.48:1 and 7.5:1. A test asserts the ratio rather than the hex, so a future
    palette change is free to move a colour and not free to move it below AA,
    and a second test pins that both palettes define the same token set (the
    absence that caused this).

  • The HTML report had one heading and no outline. The timeline carried an
    <h2>; the tag table and the scenario list — the body of the page — had
    none, so there was nothing to navigate by and no anchor to link a section
    with. Both gained one, sharing the class the timeline already used (renamed
    from .timeline-h to .section-h, since it now serves three). Pinned
    structurally, so a section added without a heading fails the test.

  • A step's name: label reached the artifact and nothing else. A macro
    with more than one step turns one feature sentence into several engine
    steps, and they share a StepRef exactly — same file, same line, same
    text. The emitter has always written the authored name: into the
    artifact's entry comment, which is why the .hurl could tell them apart;
    StepRef never carried it, so the console, the HTML report, JUnit, TAP,
    the job summary and explain all printed the same sentence once per step,
    with nothing but the status glyph to distinguish a warning from the failure
    beside it. The reference corpus demonstrated it: three step_finished
    events for the cookie session is exercised, byte-identical in the pinned
    snapshot, are now obtain the session cookie, optional probe (forces a split) and cookie survives the split.

    StepOutcome and step_finished now carry label, exactly as they carry
    fragment — the two answer neighbouring questions (which file did this
    request come from
    / which step of the sentence is this) and travel the
    same channels. One proef_core::report::step_label renders it for every
    sink, so the six cannot drift. Additive on the wire: absent when a step has
    no name:, so every pre-existing record still parses and re-renders
    unchanged, and the event schema stays 1.

    This retires two claims that were not true when written:
    AUTHORING.md's "they anchor artifacts, events, and failure output" and
    LoweredStep::label's own "(events/console)". Same class as
    reproduce_hint in the R18 wave — computed all along, printed all along,
    dropped by the record.

  • A fragment's text ran on into the comments introducing the entry below it.
    hurl attaches the blank and comment lines above a request to that request,
    which is exactly what makes the # @proef binding reliable — but it also
    means an entry has two different starts: where its lines begin and where its
    request begins. The scanner used one value for both, ending each fragment at
    the next entry's request line, so every comment a corpus author wrote to
    introduce the next request was copied into the previous fragment and from
    there into the emitted .hurl. An artifact could carry
    # Destructive. Operators only. while containing no destructive request at
    all, and trim_end could not help — a comment is not whitespace. The same
    applied at the end of a file, where a trailing note became part of the last
    fragment. A fragment now runs from its annotation to the end of its own
    request and response; the gap between two entries documents the one below it
    and belongs to neither. Nothing executed differently, because hurl permits
    only comments and blanks between entries — which is why it survived: the only
    damage was to what the durable record says a request is.

    The property covering this asserted one request line per fragment, which is
    blind to comments; it now also asserts that no fragment holds any of the
    generator's inter-entry filler.

  • explain and the HTML report disagreed about a truncated run's totals.
    A record with no tail run_finished — a run killed mid-flight — is
    reconstructed by counting, and each surface carried its own version of that
    fallback. On the sa...

Read more

proef v0.14.0

Choose a tag to compare

@github-actions github-actions released this 18 Aug 15:10
5c82f33

Fixed

  • --rerun after a cancelled run continues it, instead of a false green.
    --max-fail (and Ctrl-C) stop a run early with the never-reached scenarios
    honestly recorded as skipped — but --rerun filtered to failures alone, so
    stop → fix → rerun ran only the old failures and reported exit 0 with most
    of the suite never executed in either run. Reproduced live before fixing
    (found by round-15 external review): stop at 2 of 6, fix, rerun →
    2 passed · 0 failed, green, four scenarios untested. On a cancelled
    base record --rerun now runs failures plus the cancellation-skipped
    tail, and says so (note: the last run was cancelled before N scenario(s) ran…); scenario-level skips only exist under cancellation, so a completed
    base keeps the old semantics exactly. This also changes --rerun after
    Ctrl-C — continuing the unfinished work is what stop → fix → continue always
    meant. Mutation-tested: reverting the union fails the continuation test.

Added

  • proef test --shard I/N — stable hash-mode sharding (R3-3). A CI matrix
    runs --shard 1/N … N/N on separate machines; scenarios are assigned by a
    frozen FNV-1a hash of the run-wide (file, scenario) identity, so adding
    a scenario never re-buckets the others
    — the measured stability argument
    that rejected index-slicing at triage (inserting one scenario re-bucketed
    the whole shifted tail under slicing, nothing under hashing; the shard tests
    pin both directions, and the assignment itself is frozen by literals — the
    hash is a published contract, and changing it would be breaking). Sharding
    applies after every other selector (the pinned filter→shard order), so
    each matrix job partitions one agreed-on set. An empty shard of a non-empty
    selection is a note and exit 0 — a small suite over a big matrix is a fact,
    not a mistake — while an empty selection keeps the loud typo'd-filter
    refusal, sharded or not.

  • proef flaky — flakiness verdicts over the retained run history (R3-2).
    The 2026 discipline is detect → quarantine → resolve, and proef already
    owned the middle step: @quarantine runs a scenario without gating the
    exit code. This is the missing detect, a fold over the records runs-dir
    already retains — the window is [run] keep-runs, and no new state is
    written. Three signals from fields the record already carries (ADR-0008):
    flapping (verdict changed between consecutive observed runs more than
    once — transition-counting, not fail-rate, which is what separates flaky
    from broken: a scenario failing every run is consistently broken, a
    different problem), passes only on retry (green, but some step needed
    more than one attempt — the latent flake pass/fail-history tools
    structurally miss; the record keeps per-step attempts), and always
    failing
    . A cancellation-skipped row is not evidence and does not count
    toward a scenario's history; phases are excluded (ADR-0014). --output json emits one object per scenario with the counts behind each verdict.
    Fewer than two runs is refused (exit 2), the same answer diff gives.

  • proef test --max-fail N stops the run after N suite-scenario failures
    (1 = fail fast) — the convention Playwright (--max-failures), pytest
    (--maxfail) and cargo-nextest (--max-fail) share, with the shared honest
    semantics: in-flight scenarios finish, the never-run rest record as
    skipped (not absent, never passed), and teardown still runs on its own
    token. The stop rides the graceful-cancel path Ctrl-C already exercises, so
    the record is a complete cancelled run — which diff --fail-on-regression already refuses to certify, exactly right for a
    deliberately-partial one. [run] setup/teardown failures never count
    toward the threshold (a broken fixture is not a failing test, ADR-0014).

Documentation

  • The R3 enhancement registry is triaged (OPEN-FINDINGS): --max-fail
    built; a flakiness verdict over the run history and hash-mode sharding
    validated as build-next (the 2026 flaky pipeline is detect → quarantine →
    resolve, and the @quarantine tag already owns the middle step); CTRF,
    pack doc and the pre-M6 seam refactors deferred with named triggers; OTel
    and Cucumber-Messages exporters declined under ADR-0008's one-record rule;
    items defined only in the absent v1 research document held for a spec.

proef v0.13.0

Choose a tag to compare

@github-actions github-actions released this 17 Aug 07:48
503a38f

Security

  • An encoded reflection of a secret is redacted (S1). Redaction was
    exact-match on the raw secret bytes, and a server that reflects a bearer
    token encoded — an OAuth introspection endpoint, a debug echo, a JWT claim
    — defeated it: a failing assert quoted the base64 form in its detail, and a
    string trivially base64 -d-able back to the live credential reached the
    console and events.jsonl, the retained record CI uploads. Demonstrated
    live against 0.12.0 by an external research pass and reproduced here before
    fixing. Redactions::new now derives each secret's common encoded forms as
    additional needles — base64 (standard and URL-safe alphabets, with and
    without padding), hex (both cases), RFC 3986 percent-encoding, and the
    JSON-string escape — so every construction site (the CLI sink, the engine's
    internal renderer, TAP) is covered by construction. This is the remedy
    GitHub's own log-masking documents for the same limitation: register each
    transformed value too. The needle set covers the reversible transforms that
    occur at HTTP boundaries and does not claim completeness — a secret
    reflected hashed or re-encrypted matches no needle list. Over-redaction is
    the accepted failure direction. Property-tested over every derived form,
    pinned end-to-end by a fixture route that echoes the bearer base64-encoded,
    and recorded as an ADR-0005 amendment.

  • The fragment corpus read is bounded. [run] fragments names a directory
    proef did not write and does not control, and it was read with no per-file or
    total cap: a 279 MB file cost 601 MB of resident memory on proef flows —
    a command that never looks at a fragment — because the text is read whole and
    then copied into an Arc<str>. A file over 8 MiB is now skipped
    (proef::pack::oversized_fragment_file) and the reader stops past 64 MiB
    total (proef::pack::fragment_corpus_too_large). The size comes from the
    directory entry, so an oversized file is never allocated at all; the same
    bound applies in proef lsp, where the corpus is held between requests rather
    than for one command. Skipped, never fatal — a corpus is foreign by design, so
    one bad file must not sink the ones beside it. An unreferenced corpus still
    costs nothing: the scan stays lazy, so nothing is reported unless a pack
    actually names a fragment. Filed as R9-3.

Added

  • proef diff takes a path. Each side is now a run id, a record
    directory, or an events .jsonl file under any name — the stream is the
    record (ADR-0008), so all three must mean the same thing. The file form is
    the CI baseline flow an adopting suite asked for: download the base branch's
    events.jsonl artifact and proef diff baseline.jsonl <new> --fail-on-regression gates the PR, with no shared record store. Previously
    every argument was joined onto runs-dir, so a path produced
    .proef-runs/<your path>/events.jsonl: No such file — the argument mangled
    into the complaint. A path that does not exist now names itself; a --baseline
    flag was considered and declined as a second spelling of the same positional.

Internal

  • A hung test is now a five-minute failure, not a five-day zombie. The
    nextest config had slow-timeout with no terminate-after, which only
    labels a test SLOW and never kills it — an lsp_stdio test wedged on an
    unbounded child.wait() ran for five days with its proef lsp child alive.
    Both layers fixed: the two bare child.wait() sites got the file's own
    bounded-watchdog pattern (a server that fails to exit now fails the test in
    10s, naming what did not exit), and the runner gained terminate-after = 2
    (120s), sized from a cold-cache census of the whole suite (slowest ordinary
    test: 5.1s). The harness_ trio — which shells cargo test inside the test
    and measured 216s on a fully cold cache — gets a per-test override to 600s,
    the nextest docs' own tight-global-plus-overrides pattern. The
    process-group kill (a spawned server dies with its test) was verified
    empirically with a deliberately hung test holding a live child.

  • Cleanup pass over this cycle's four PRs (reuse/simplification/efficiency/
    altitude review). The corpus-bound decision moved into core as
    pack::CorpusBudget — it was abstracted in the CLI and hand-copied in the
    LSP, agreeing by copy rather than by construction; both readers now share it
    and only measurement stays reader-local. Redactions stopped allocating on
    the miss path (nearly every call: per string field per event under the
    reporter-stack mutex, with the needle list ~9× larger since the encoded
    forms) — clean fields now hand back their original Arc. A relative source
    path is left exactly as it arrived, per its documented contract — it had
    been falling through to a per-file canonicalize that could rewrite a
    ../-typed spelling. The LSP's percent-encoder folded onto core's
    (byte-identical copies, one character set to drift). The fixture's
    hand-rolled base64 became the crate call — its dependency-surface rationale
    died when this same cycle made base64 a workspace-wide compile. diff's
    path-or-id resolution moved beside its sibling in record. A deny.toml
    home for the curl floor was tried and reverted by mutation test:
    cargo-deny 0.19.8 mismatches build-metadata versions (curl-sys@<0.4.90
    banned the good 0.4.90+curl-8.21.0); the floor stays a unit test, now
    scanning every lockfile entry rather than the first.

  • The bundled libcurl cannot silently regress under the June-2026 CVE
    batch.
    curl-sys 0.4.90+curl-8.21.0 in the lockfile is past the batch —
    but only as a transitive accident of resolution, and the usual gates are
    structurally blind here: RUSTSEC carries no advisories for CVEs in a
    *-sys-bundled C library, so cargo audit/deny stay green however stale
    the bundled curl is. A test now asserts the lockfile floor, and each release
    build prints the libcurl actually linked into that artifact (proef doctor
    already reported it; the release log now carries it per target). The hurl-8.1
    watch items — variables-file:'s missing sandbox first among them — are
    recorded as a pin-bump checklist in the thin-fork runbook.

  • Fuzzing reaches the fragment rules. fuzz_pack_load ran against an empty
    corpus, so ref: resolution, bind: keys nothing reads, a bind: colliding
    with a variable the fragment supplies itself, and unbound placeholders were
    covered on paper and unreachable in fact. The new fuzz_fragment_binding
    target is structure-aware: it builds a well-formed pack and corpus and
    spends its budget on the name space where those rules live. That shape was
    chosen from measurement, not taste — a byte-oriented version never once
    resolved a ref: in 1.45 million runs, because reaching the rules meant
    discovering valid YAML and a matching corpus at the same time. The corpus is
    read by a synthetic scanner rather than hurl's, which is what keeps the fuzz
    workspace free of native libraries: cargo dependencies are package-level, so
    one engine-dependent target would compile hurl for all of them.

  • Hurl's own annotation scanner is property-tested, in proef-engine-hurl
    where the native libraries already are. The properties pin what the
    entry-boundary arithmetic is for: every reported line lies inside the file,
    every entry is accounted for exactly once, the starts are ordered and
    distinct, and — the one that matters — no fragment's text runs into the
    entry after it
    . That last assertion exists because a first draft without it
    passed while the boundary was deliberately broken.

  • The fuzz target list comes from cargo fuzz list. It had been spelled out
    in ci.yml and nightly.yml, so a new target ran nowhere until both were
    edited, and nothing failed to say so.

Fixed

  • A run record no longer names the machine that produced it. [run] suite
    resolves against the config directory (0.12.0), so a path-less proef test
    handed the front end an absolute path — and every emitter printed it: the
    .hurl # source: header, .map.json's feature.file, every
    step_finished event, the console, and pack diagnostics. Two checkouts of one
    suite stopped producing equal artifacts, which is the property ADR-0010 exists
    to guarantee; an adopting suite hit it as /Users/… in 133 artifact lines and
    64% of its event stream by bytes.

    The resolution rule was right and stands. What was missing is its naming
    dual
    : resolve against the project, then name against the project again.
    front::SourceNaming is now the one boundary that answers "how is this path
    spelled", for features, packs and fragments alike — replacing the fragment
    corpus's separate cwd-relative strip, which was a second anchor for the same
    question. The four ways to name one suite — derived from [run] suite, typed,
    typed absolutely, or reached from a subdirectory — now emit one artifact, byte
    for byte.

    A path that arrives relative is recorded exactly as it arrived; a suite or
    corpus genuinely outside the project keeps its absolute name, there being no
    project-relative spelling of it. Filed as R12-1, and it closes R9-6, which had
    described the same defect as safe from the project root — it no longer was.

    Breaking, by the rule in docs/RELEASING.md: it changes emitted artifact
    bytes, which is inherently breaking and takes a MINOR bump. Migration: nothing
    to do for a suite invoked with a typed relative path — those bytes are
    unchanged. A tool reading step.file or feature.file out of a record
    produced by a path-less run now sees a project-relative path where it saw an
    absolute one; join it onto the directory holding proef.toml. Records written
    by earlier versions are not rewritten.

Added

  • [run] keep-runs bounds how many past run records runs-dir retains. Th...
Read more

proef v0.12.0

Choose a tag to compare

@github-actions github-actions released this 14 Aug 08:46
c70602a

Fixed

  • A runs-dir edited mid---watch no longer feeds the loop its own output.
    Reruns re-read the config (the fix below), so records went to the new
    directory while the watcher's exclusion still named the one it had frozen at
    startup — and every rerun's artifacts/*.hurl, now under an unexcluded
    directory, requeued the next run. One edit produced 39 runs in 12 seconds,
    firing real traffic. This was the third outing for the watch-feedback class,
    so the fix removes the second answer rather than resynchronising it: each
    rerun registers where it is about to write, and the exclusion is derived from
    the same config the run is. A directory a previous run wrote stays excluded
    too, since its events can still be in flight. Filed as R11-8.

  • --watch --config <relative path> retriggers on config edits. The watcher
    compared the config by exact path while notify reports events under the
    spelling the OS resolved them to, so --config proef.toml never matched and
    config edits produced nothing — silently, because feature edits kept working
    and the loop looked alive. Symlinked and /tmp-style aliased paths failed the
    same way and are also fixed: the flag is made absolute when it is stored, and
    identity is settled by comparing canonical paths, which is a stricter question
    than being absolute. The same relative-path flaw silently cost proef lsp --config <relative> go-to-definition across the whole fragment corpus, since
    documents::name_to_url refuses a relative name. Filed as R11-9.

  • doctor reports a proef.toml that will not parse. The discovery arm had
    become a silent unwrap_or_default, so a malformed config left doctor
    reporting on invented defaults and printing "all checks passed", exit 0 — with
    the parse error, which the previous code printed, discarded. It is a project:
    row now, so it reaches worst and the exit code a CI script actually reads.
    Being absent is still not a finding: doctor must run outside a project.
    Filed as R11-10.

  • proef fragments exits non-zero when a [run] setup/teardown phase fails
    to load.
    It printed error: setup feature failed to validate: and exited 0,
    because the phase half flattened its failure to "not measured" while the suite
    half kept its code. Withholding the counts was right; reporting success while
    printing errors was not.

  • proef.toml has one path rule. A path written in the config now resolves
    against the directory holding the config; a path typed on the command line
    still resolves against the working directory. [run] fragments already worked
    this way and everything else did not, so two keys in one table meant two
    different roots: from a subdirectory fragments = "hurl" resolved while
    suite = "features" reported "neither a feature file nor a directory". With
    --config the split was worse than inconsistent — pointing at a config in
    another tree ran dry-run OK over whatever suite happened to sit beside the
    shell, and never looked at the configured one.

    The rule now covers suite, setup, teardown, runs-dir and the tests/
    convention probe, plus two files nothing had inventoried: .proef-state.json
    (the persistent World) and .proef-secrets.json (the secret store), which
    were anchored on the working directory — so two shells in one project were two
    Worlds and two secret stores. It is the convention Cargo, tsconfig.json and
    pytest's rootdir all follow. Absolute values are taken as written, and with no
    proef.toml in scope written paths stay relative to the working directory, so
    the config-independent reference corpus is unaffected. Filed as R11-1.

  • --watch rereads the config it retriggers on. Editing proef.toml
    retriggered a run that still used the snapshot loaded at startup: changing
    [url] base produced a rerun that dutifully called the old host, and the same
    went stale for jobs, [env.*] and exclusive-tags. Watching a file whose
    contents you then ignore is worse than not watching it, because the rerun
    reports that the edit was taken. Each rerun now re-reads the file and
    re-resolves the suite from it; a config that no longer parses fails that rerun
    and leaves the loop watching, since half-typed TOML is the normal state of a
    file being edited. Which directories the loop watches is still fixed at
    startup, so changing [run] fragments or [run] suite needs a restart to be
    watched. Filed as R11-2.

  • --config is honoured or refused by every subcommand. doctor printed the
    error for a missing named file and then reported on defaults, exit 0 — the
    "fall back to defaults" CONFIG.md forbids — while fmt, init, schema and
    secret accepted a nonexistent path silently, against the "global to every
    subcommand" claim in CONFIG.md, README.md and this file. A named file that
    is not there is now exit 2 everywhere, including where nothing reads it;
    doctor stays lenient about discovery, which is a different claim. secret
    additionally uses the flag, since the store is the project's. Filed as R11-3.

Breaking: the secret store, the persistent World and the run records
move with the config rather than with the shell. What decides whether this
reaches you is where you invoked proef, not where proef.toml sits: runs
started from the project root are unchanged, but a run started from a
subdirectory used to write .proef-state.json, .proef-secrets.json and
.proef-runs/ beside the shell, and now writes all three beside the config.

Nothing is migrated, and none of it announces itself. A World written from a
subdirectory reads as empty, so saveAs: global values start over on the
first run after upgrading; stored secrets read as absent; and the old run
records are simply invisible to explain, report and diff, which say "no
run records" rather than erroring. To carry them over, move
.proef-state.json, .proef-secrets.json and .proef-runs/ from the
directory you used to run from into the one holding proef.toml. Otherwise
re-run proef secret set and take a fresh baseline.

Breaking (library): proef_cli is not a published library surface, but for
the record front::run takes the state-file path, ProjectConfig::runs_dir
returns a PathBuf, setup/teardown return Option<PathBuf>, suite is
gone (fold into default_suite_path), and the secretstore entry points take
the store path. proef_core gains one item:
pack::FragmentCorpus::unreadable_file.

  • [run] exclusive-tags validates itself. --dry-run did not parse the
    expression at all, so a malformed one exited 2 from proef test and passed
    dry-run OK … 0 warning(s) from the gate CI runs. And a well-formed
    expression matching no scenario was silent: @soloz against a @solo suite
    put every scenario back in the shared pool, exit 0, nothing said — the exact
    silent degradation the key was designed as a config expression to prevent, and
    one that reads as flakiness rather than as a typo. Both paths now parse it, and
    a zero-match expression warns, naming it and pointing at proef flows. Judged
    over every scenario the suite loaded rather than the ones selected, so a
    --tags filter that removes the matches from one run is not reported as a
    broken setting. Filed as R11-4 and R11-5.

Changed

  • proef fragments says which half it could not measure. --check reported
    "needs a suite that binds" when the suite had bound perfectly well and a
    [run] setup/teardown feature was the thing that failed to load, sending the
    reader to inspect the half that was fine. The degraded listing also now carries
    the note macros prints, so withheld counts read as "not measured" rather than
    as a corpus nothing uses.

  • proef fragments --check refuses to pass with no corpus configured. With
    [run] fragments unset it printed 0 entries and exited 0, indistinguishable
    from a fully-used corpus — so a CI gate disarmed silently the day the key left
    the config. The listing still works; only the gate is now a user error.

  • proef fragments --output json carries annotated on both row shapes. The
    annotated and unannotated rows differ in eight fields, and consumers had to
    probe for the absence of one to tell them apart.

Documentation

  • CONFIG.md's "everything else keeps running at jobs width" was false:
    queueing is strict FIFO, so nothing new starts while an exclusive scenario
    waits at the head. The cost is bounded, not absent, and is now described.
  • The one caveat [run] exclusive-tags carries is written down in CONFIG.md and
    ADR-0007: exclusivity is enforced against the dispatcher's active set, which a
    watchdog-abandoned scenario leaves while its detached thread is still issuing
    requests (hurl cannot be cancelled mid-entry).
  • TECH-SPEC §10 gained proef fragments and the global --config; §11's
    [run] inventory listed three of seven keys.
  • DIAGNOSTICS.md carried a pack::load row nothing emits — a reader who
    grepped it found a plausible cause that could never be one — and filed
    lower::multiline_bind under proef::pack::*. Both fixed, and the two-way
    agreement between the file and the emitted codes is now a test, since this
    drifted twice.
  • OPEN-FINDINGS R9-2 still said fuzz_tag_expr "sits in neither fuzz loop"
    three sections after recording that it is in both.

proef v0.11.1

Choose a tag to compare

@github-actions github-actions released this 12 Aug 21:36
68c16b9

Fixed

  • An output path creates the directories it names. --junit, --sarif and
    report -o failed when the parent directory did not exist, while
    artifacts -o and the run directory created theirs — no rule, four sites
    deciding separately, with the two used most in CI on the failing side. Every
    adopter paid the same mkdir -p. pytest --junitxml, jest-junit,
    cargo-nextest's JUnit store and the hurl proef embeds all create them.
    This does not weaken the "side effects should be explicit" principle: that is
    about writing files the user did not name, and here they named exactly this
    path.

  • proef fragments counts [run] setup/teardown usage. A fragment only a
    phase feature reached was reported UNREACHABLE — no macro refs it, which was
    false, and failed --check — a false CI failure in the workflow --check
    exists for. The verdict also depended on where the phase file sat: inside the
    suite directory it was discovered as an ordinary feature and counted. The
    listing's universe now matches the runner's, and a phase that fails to load
    withholds every count rather than guessing. Filed as R10-2.

  • One predicate answers "is this a fragment file?" (FragmentSupport::claims).
    Three answered it before — CLI discovery via Path::extension, the core scan
    via rsplit('.'), and the LSP's corpus invalidation case-insensitively — so
    they disagreed about api.HURL (the editor rebuilt its corpus for a file
    nothing would scan) and about a dotfile named .hurl. Filed as R10-3.

  • --config reaches proef lsp and --watch. The flag bypasses the
    upward search so a proef.toml beside the suite becomes usable — but the
    editor re-discovered its own config and --watch watched whatever a fresh
    search found. So in exactly the layout the flag exists for,
    proef test --config … ran green while the editor reported every ref: as
    unknown, and editing the config driving the run never retriggered it.
    ProjectConfig now keeps the file it was read from (with root derived from
    it rather than stored beside it), and both consumers use the config actually
    in force. For proef lsp the flag also outranks the client-announced
    workspace root. Filed as R10-1.

proef v0.11.0

Choose a tag to compare

@github-actions github-actions released this 12 Aug 12:52
3dd967e

Added

  • [run] exclusive-tags — a tag expression selecting scenarios that run
    with the pool to themselves. Real suites contain scenarios that cannot run
    beside anything: one asserting absolute positions (items[0]) needs a store
    no concurrent scenario writes to, and the only workaround was several CLI
    invocations driven by tag discipline in a Makefile, each producing its own run
    record, JUnit file and exit code to aggregate in shell.

    A matching scenario waits for the pool to drain, runs alone, and the pool
    refills after it; everything else keeps running at jobs width, and discovery
    order is unchanged so an exclusive scenario never loses its place. A config
    expression rather than a reserved tag name, because with a bare convention a
    scenario added months later lands untagged in the parallel pool and breaks
    isolation intermittently — which reads as flakiness rather than as a missing
    declaration. A malformed expression is a user error, never a silently-ignored
    key.

    This is exclusion, not ordering: a scenario that must run before the
    rest belongs in [run] setup, which already runs once before the pool exists.
    Deliberately one axis of the two cargo-nextest settled on — per-group
    concurrency limits (rate-limiting a shared dependency) are a real future need
    that nobody has asked for, and a group table can be added later without
    breaking this key.

  • proef fragments — the corpus listing, symmetric with macros. Until now
    no proef output stated how many fragments there were, so neither way a
    fragment can die had a denominator to be noticed against: one no macro
    references was unobservable, and one reached only through a macro no scenario
    binds looked covered because the macro was flagged. Both are now named
    apart, unannotated entries are listed by line (they have no name to list by),
    and --check exits 1 when something never runs. --require-annotated extends
    that to unannotated entries and is deliberately opt-in: an unannotated entry is
    inert by design (ADR-0018), so "not done yet" is a porting team's meaning, not
    every adopter's. Reachability is read off the lowered scenarios, so a fragment
    reached through a chain of use: counts as reached.

  • --config <path>, global to every subcommand, naming the proef.toml to
    read instead of searching up from the working directory. Discovery only goes
    up, so a config beside the suite is unreachable from the repository root — a
    layout an adopting team planned and abandoned after it failed. A named file
    that does not exist is a user error rather than a fall back to defaults:
    discovery finding nothing means "no project here", but a named path that is
    not there is a typo, and a silently unconfigured run is what that used to buy.

  • proef doctor sees the fragment corpus — a row reporting how many
    fragments loaded from [run] fragments, warning when the configured root is
    not a directory. A misconfigured path used to surface much later as
    pack::unknown_ref: an error about a name when the cause is a path.

  • proef init scaffolds both body forms — a one-entry .hurl file with a
    # @proef annotation, [run] fragments, and a pack macro of each kind. The
    newcomer with most to gain from ref: is the one who already owns a hurl
    corpus, and a scaffold teaching only hurl: | reads as "proef wants your
    files transcribed into YAML".

Fixed

  • A bind: key nothing reads is refused (proef::pack::unread_bind_key),
    with did-you-mean over the names actually in scope. bind_without_ref only
    caught a table with no ref: at all, so bind: { token: …, toekn: … }
    validated clean — the one authoring mistake in the fragment path that produced
    no signal whatsoever. Checked as a union over the scope, never against one
    fragment: a pack-scope table is the plumbing every macro in the file needs, so
    a key serving one macro and not its siblings stays correct.

  • duplicate_fragment no longer says "in both x and x" for two entries
    in one file, and stops offering file.hurl#name as the remedy there — that
    qualifies by file and cannot separate two entries inside one. Annotating a
    corpus adds many names to few files, which makes same-file the likely
    collision.

  • unbound_placeholder names all three supply routes. The omitted one was
    the fragment's own [Options] variable: — the route that makes a corpus file
    runnable standalone, which is the property ADR-0018 exists to preserve.

  • A fragment's [Options] escaped the ADR-0007 value caps. retry: -1,
    repeat: -1 and an unbounded delay: were rejected in an inline hurl:
    block and accepted in a ref: fragment — byte-identical text, exit 2 one
    way and "dry-run OK, 0 warning(s)" the other, then written verbatim into the
    executed input. The scan lived inside the inline-only linter; only the
    twinned-option half of pass 6 had crossed to fragments. It reads the text
    alone, so it now runs against a fragment's too, anchored on the ref: line
    and naming the fragment file and line. This is the case the caps exist for:
    hurl has no cancellation, so an infinite retry makes the batch budget
    unestimatable and leaves the watchdog abandoning a thread it cannot stop.

  • A step declaring both ref: and a payload was told, falsely, that its pack
    had no ref: at all.
    The conflicted step is reported and dropped, so the
    loaded bodies stop showing every ref: the author wrote — and the pack-scope
    bind_without_ref check then drew a conclusion from the gap. It now infers
    nothing from a pack whose steps did not all normalize.

  • A pack-scope bind: with no ref: anywhere was silently dropped.
    AUTHORING.md said bind_without_ref applies "at every scope" while only the
    macro and step scopes were checked — and a setting ignored in silence is the
    bug those two exist to refuse. The check was the better half of the
    disagreement, so the pack scope now has it too.

  • A multi-line bind: value blamed the artifact. A hurl
    [Options] variable: value is a single-line scalar, so a newline could never
    reach the entry — but it surfaced one stage later as emit::invalid_artifact,
    pointing at generated text the author never wrote. Refused by name at lower
    time as lower::multiline_bind, naming the inline hurl: | form that is
    what splices a multi-line body (ADR-0018's splicing-versus-binding boundary,
    enforced where it can be explained).

Changed

  • Breaking (library): AnalyzeCtx takes the fragment corpus instead of
    building one. Building it internally meant a fresh scan memo per call, so the
    LSP re-read and re-hurl-parsed the whole corpus on every request — each
    completion popup, each go-to-definition, each debounce tick. The server now
    holds one and rebuilds it only when a fragment file changes; editing a pack or
    a feature, which is nearly every keystroke, leaves it alone. It is also what
    core purity already required: the caller does the IO.

  • Breaking (library): StepKindSpec gained options, an engine-contributed
    recogniser mapping a raw option key to what ADR-0007's budget rules should make
    of it. The fragment half of that rule already crossed the seam while the inline
    half matched "retry-interval:" as a literal inside proef-core — one rule at
    two altitudes, and a second engine would have had its fragments linted and its
    inline blocks not. A kind contributing no recogniser is not linted, since the
    core has no way to know what its option keys mean.

  • Breaking (library): proef_core::engine::FragmentScanner returns
    ScannedFile { fragments, unannotated } rather than Vec<ScannedFragment>.
    An engine's scanner now also reports the 1-based lines of entries carrying no
    annotation — lines only, never built-then-discarded fragments, so a foreign
    corpus still costs a push per unannotated entry. Without it "which entries did
    I forget to annotate?" is unanswerable: a missing annotation produces a green
    run and a silently absent test, and the entry that would prove it was never
    built. FragmentCorpus gains fragments(), unannotated() and
    diagnostics(), because the scan is gated on some pack naming a fragment —
    so PackSet::fragments is empty for exactly the suite a listing has most to
    say about.

Documentation

  • Config discovery is a requirement, not a convention. proef.toml is found
    by searching up from the working directory, so a config beside the suite
    (tests/proef/proef.toml) is never found from the repository root — an
    adopting team planned that layout and discovered it by failure. CONFIG.md now
    says so, and notes that keeping the file at the root collapses the one place
    [run] fragments (config-relative) and suite/setup/teardown/runs-dir
    (cwd-relative) differ.

  • The release runbook could not work as written. main is a protected
    branch, and step 4's git push origin main --follow-tags fails in the
    dangerous direction: --follow-tags is not atomic, so the branch is rejected
    while the tag still lands — and the tag is what release.yml triggers on,
    starting a release build from a commit that is not on main. It happened
    cutting 0.10.0. The runbook now routes the release commit through a PR and tags
    the merged commit, and the cargo publish section carries the dry-run,
    tag-check and --locked sequence plus why only four crates go
    ([workspace.package] publish = false is the default). Also drops step 1's
    reference to changelog "bottom links", which do not exist.