Repository navigation
Releases: emrecdr/proef
Release list
proef v0.19.0
Fixed
-
The Homebrew formula installs the man page and the shell completions. Its
def installwasbin.install "proef"and nothing else, so from 0.16.0 —
the release that started shippingproef.1and fivecompletions/files in
every archive — until 0.18.0,brew install proefgave noman proefand no
tab completion, while binstall and a direct download gave both. The formula
is a heredoc insiderelease.ymland the archive is staged in a different
job, so nothing tied the two together and no gate could see the gap; the
render step now fails if the archive lacks a file the formula installs, and
the formula's owntest doasserts the man page and completion landed.
Takes effect on the next tag: a tag runs the workflow from its own commit. -
--dry-runrefuses afile,…;asset that is not there. A suite whose
asset had been deleted reporteddry-run OK, and the failure arrived later
from a different command, against a live backend — from the one gate CI runs
before standing an environment up. Whether an asset resolves is statically
knowable, so it is answered there now. The checker is staging's own
(assets::resolve_assets, split out ofstage_assets) rather than a second
walk over the same artifacts, so validation and the run cannot disagree; the
message and the diagnostic code are the ones a run already gave. -
file,inside a JSON or assertion body is no longer mistaken for a file
asset. The emitter found the files an artifact reads by scanning its text
for the literalfile,and a closing;, so a request body containing that
substring —{"note": "see file,notes.txt; for details"}— produced a
phantom asset, and staging then failed the run over a file the request never
reads. The claiming engine now reads its own AST, where a body reference and
six characters of prose are different things.
Breaking
-
proef_core::emit::emittakes the registered step kinds, and
StepKindSpecgains anassetshook. Asset recognition was hurl's body
grammar living inproef-core:emit::file_refs_inscanned for the literal
"file,", which ADR-0002's amendment forbids and — worse — which the guard
pinning that amendment could not see.engine_grammar_kindclassifies
fences,HTTP,[Section]headers, method lines andkey: valueoptions; a
body constructor is none of those, so the literal was never sanctioned and
never reported missing. The ADR's own measurement said thirteen literals; it
was fourteen.The scan moves behind the seam as
StepKindSpec::assets, the fourth
engine-contributed hook besidevalidate,fragmentsandoptions, and the
guard gains abodyarm so the shape is classifiable whether or not anything
currently uses it.emit()takes&[StepKindSpec]to reach it;FrontEnd
carrieskindsbeside thekind_to_enginetable it is built with, which
registryalready documents as a pair that must not be re-derived
separately.emit::file_refs_inis gone.
Internal
-
A fragment's assets stage from where its file was read, not from where its
name points.AssetRoots::source_dirrebuilt a fragment's directory by
splittingfile.hurl#nameand joining the file half onto the project root —
the naming boundary run backwards, without the canonicalize fallback that
boundary carries precisely because a lexical-only version already shipped a
bug (a suite reached through a symlink silently failed to match, R11-9). The
two agreed only because both were seeded fromconfig.root()and discovery
walked from that same root, so only the lexical case was ever exercised, and
nothing made them stay inverses. The corpus reader now records the directory
it read each file from (front::CorpusDirs, carried onFrontEndbeside
kinds), and staging looks it up — the fragment-side twin of what
LoadedFeature::read_fromalready does for features, so both halves of the
naming boundary are one-way in the same way.AssetRootsloses itsproject
field andbuild_specsitsproject_rootargument: with nothing to
recompute, the project root is no longer staging's business. -
--rerunreads its base record once. It calledrecord::read_eventsfor
the JUnit overlay and thenrecord::rerun_candidates, which read and
deserialized the sameevents.jsonla second time — two full passes bounded
only by the 256 MiB record ceiling, over a file another process may still be
writing.rerun_candidatesnow takes the&[Event]its caller already
holds, which is the ruleread_record's own documentation had already
stated for exactly this case. The read error is handled once as well: the
first call swallowed it with.ok()and the second rediscovered it a line
later. -
The one doc check that reads only files now runs in the half that reads
files.no_current_behaviour_doc_spells_a_format_as_an_output_pathlived
intests/docs.rs, whose stated charter is the checks needing a built
binary to ask clap — this one only scans markdown, so it never ran in the
fast doc-only CI step. It is nowxtask docs-check's
check_output_path_spelling, reusingliving_docs()instead of carrying a
second directory walk. Its allowlist-shrink guard got stricter on the way:
it counted ADRs into the same total, so a renamed entry could be masked by
docs/adrbeing larger than the shortfall — which is the one failure that
guard exists to catch. All three paths were checked by mutation: a stale
spelling planted in an allowlisted doc, one planted in an ADR, and an
allowlisted doc renamed away.
Documentation
-
The worklist stops contradicting what shipped. Three entries in
OPEN-FINDINGSstill called CTRF declined or its trigger unfired — the
2026-08-31 external re-test, the RF audit's deferred list, and R3-5 under
"deferred, with the trigger named" — for the eight days after--ctrf
actually shipped (#160). R3-9, four bullets below R3-5 in that same list, was
annotated the moment it shipped — the convention the three missed. Two more claims
had outlived their facts: the shipped-changelog duplicate headers (no release
carries one now, andcheck_changelog_kindsfails if one returns) and the
machine-side note about Homebrew's Rust shadowing rustup. Filed at the same
time:a_second_interrupt_hard_exits_with_130failed once on Linux CI and
passed on a re-run of the same commit, so the evidence, the mechanism and the
fix shape are written down instead of left to the next re-run. And the stance
that a scenario-level@retryis deliberately absent — retry-until-green
hides a one-in-four defect 99.6% of the time — is stated in
TESTING-STRATEGY§5, which the worklist asked for and nobody had written. -
The runbook records that the registry skips three versions. 0.15.0–0.17.0
were tagged and GitHub-released but never published, so crates.io moves
0.14.0 → 0.18.0. Noted inRELEASING.mdso the gap does not read as a failed
upload. The long-standinghomepagequestion inOPEN-FINDINGSis also
resolved: the field reached the registry with 0.18.0, exactly as that entry
predicted;documentationremains unset and still open. -
The release history records every release again.
RELEASING.md's History
section carried no entry forv0.16.0orv0.17.0and filedv0.15.0
betweenv0.13.0andv0.14.0; the order is repaired and all three versions
are present,v0.18.0included. The corpus also stops calling the 0.18 series
unreleased, and anIMPROVEMENT-PLANpointer into CHANGELOG[Unreleased]
now names the releases that actually carried the work —[Unreleased]has
been cut several times since that sentence was written.
proef v0.18.0
Added
-
SIGTERM and SIGHUP now take the graceful path (ctrlc's
termination
feature): a CI job timeout ordocker stopcancels the run — in-flight
batches finish, the rest record as skipped, teardown runs, the reports are
written, and the record closes with acancelledrun_finished— where it
used to kill the process mid-write and leave a truncated record with no
tail. A second signal still hard-exits 130 (the handler carries no signal
identity, so the code is 130 for every second signal). Pinned by
sigterm_cancels_gracefully_and_the_record_completesand — for the first
time anywhere — an exit-130 assertion,
a_second_interrupt_hard_exits_with_130. -
test --format jsonandexplain --format jsonnow reportwarnedand
cancelled. A warned scenario (anoptional:step failed, or asaveAs: globalpromotion was refused) folded intopassed, andcancelled— in
the record'srun_finished— was surfaced by neither, so a script could
not tell a spotless run from one with warnings, nor a complete run from a
cancelled one, and the two JSON surfaces disagreed on how to say "did not
finish" (0.18 survey). Both keys are additive and always present.
warnedalso becomes visible in JUnit (a<system-out>note, the status
stayssuccesssince JUnit has no warned) and CTRF (anextra.warned
flag) — it was previously visible only in the HTML report. -
A tag that looks like a reserved one but is not exactly it now warns
(proef::tags::reserved_tag_typo).@quarantined,@skipped,@Skip
matched no reserved tag and silently did nothing — a scenario the author
believed was quarantined gated the build. The warning names the spelling
it likely meant, tuned to catch the real typos without firing on
legitimate short tags (ship,slip,step). -
proef flakygains the 2026-field statistical guards (0.18 survey §6),
each a pure fold over the JSONL history already retained — no new state, no
gating mode (advisory stays the design):- A minimum-sample floor (
--min-samples/[flaky] min-samples,
default 10): below it a scenario isinsufficient-datarather than
classified, because a verdict on thin data is worse than none. - Hysteresis (
--recovery-runs/[flaky] recovery-runs, default 5): a
flapping or latent scenario holds its flag until it earns a trailing clean
run, so it cannot oscillateflaky↔healthybetween adjacent runs. - An environment-outage guard (
--outage-rate/[flaky] outage-rate,
default 0.8): a run where over this share of suite scenarios failed is an
environment incident, not evidence about any one scenario, and is excluded
— so a single fixture or staging outage cannot mark the whole suite broken. - An input-fingerprint equivalence class — the default key. Each run
writes aninputs.jsonsidecar carrying a hash of what it executes
(feature sources + loaded macros/fragments + the resolved
${url:…}/${vars:…}scope), so a pack, feature, orproef.tomledit
correctly ends the comparison window instead of silently mixing runs of
different inputs. It is a proef-computed fact about proef's own inputs,
not harvested from the environment (ADR-0020 unchanged — git-commit
grouping stays handed-over via--meta commit=…andproef flaky --by commit).broken≠flaky, transition-counting, and the quarantine
lifecycle were already present and are unchanged.
- A minimum-sample floor (
Fixed
-
A run-record write that fails now reaches the exit code. The JSONL
reporter deliberately swallows write results (a reporter cannot report its
own channel dying), andevents.jsonlwas handed a bareFile— so a disk
filling mid-run truncated the record while the run still exited by its
verdict, the exact class the v0.6–v0.8 series closed for the console. The
record's writer now latches its first failure (one stderr line, run
continues) and the exit funnel turns it into a system error, the same shape
as the stdout latch and the JUnit-write fold — unified in one pinned
function,escalate_environment_failures.run.log's mirror keeps its
own contract (creation is warn-and-continue, so a mid-run failure warns
once and leaves the verdict alone — previously it was silent). -
The GitHub step summary can fail again. It was the only CI sink that
couldn't: a failed open or write vanished while JUnit and CTRF failures
re-classify the exit — so the page a reviewer actually reads could be
missing on a green exit.write_github_summarynow returns the error and
the caller folds it into the samereports_failedpath as its siblings. -
A custom
--run-idno longer collapses the JUnit report identity onto
the nil uuid. ADR-0021 made non-uuid run ids first-class, but the report
uuid wasparse_str(...).unwrap_or(nil)— every--run-id cirun emitted
00000000-…, colliding in any consumer keyed on it. A non-uuid id now
derives a stable UUIDv5 from its bytes (a uuid id passes through verbatim). -
The interrupt window and the interrupt's own words. The handler is
installed at the top ofexecute— before the front end, the run dir and
the record exist — so no startup window takes the process default any
more. Its installation failure is a printed warning (it was silently
ignored, unlike--watch's handler). The second-signal path no longer
prints before exiting: the print took stderr's lock, which a worker
blocked on a full pipe can hold, wedging the escape hatch behind the very
stall it exists to escape. And the teardown notice said "Ctrl-C again to
skip" when a second interrupt actually hard-exits dropping every report —
it now says what happens. -
Asset staging no longer depends on the working directory. A feature's
file,…;assets were resolved by joining its portable name against the
cwd — but a name's anchor (the project root, or the caller's own typed
spelling) is not in the string, so a typed-absolute or config-written
suite path run from any subdirectory failed staging with exit 2, blaming
the author for a correct file (the feature-side twin of OPEN-FINDINGS
H5). The resolved discovery path now travels beside the name
(LoadedFeature::read_from) and staging resolves beside the file the
parser actually read — the H5 prescription, applied to the feature side.
Reproduced before the fix and re-verified after, from a subdirectory,
against the reference corpus; a new integration test pins a project under
a path with spaces and non-ASCII segments, which nothing in the suite had
ever exercised. -
--sarifline numbers survive acd, and byte-match the parser. The
SARIF writer re-read each source from disk by its portable name to count
lines — from any subdirectory every read failed andstartLinesilently
vanished, annotating nothing; the re-read could also disagree with the
span by exactly the parser's normalization. Lines now come from the
diagnostic's own carried source text — the same normalized bytes the span
indexes. (On Windows, an absolute out-of-projecturialso spells its
separators as a URI requires.) -
Staging's two symlink edges. An existing symlink at a staging
destination was written through —fs::copyfollows links, so the
bytes landed wherever it pointed, outside the root built to contain
them; it is now replaced. A source symlink stays followed, deliberately:
stockhurlfollows it too, and refusing would break the dual-runner
rule (the module doc now says so). -
Asset names that are one file to the filesystem are refused. The
duplicate-name guard keyed on the raw reference string, soData.json
anddata.json— one file on macOS and Windows — silently last-writer-won,
the very overwrite the per-scenario root was built to end. The check now
runs on the canonical path the copy actually landed on, which is exact on
every platform: a case-sensitive volume keeps both files legitimately, and
nothing fires. -
Artifact slugs cap at 120 bytes. The slug flattens the feature's whole
directory path into one filename component, andassets/<slug>/repeats
it as a directory — so path depth became filename length, and a deep tree
or a long scenario name (multi-byte scripts at a quarter of the visible
characters) sailed past NAME_MAX and failed the write. Over the cap, the
tail is a hash of the whole uncapped slug, so two names differing only
past the cut still name two artifacts; every slug the existing corpus has
is under the cap and unchanged byte-for-byte. -
The ADR-0007 budget family is closed over its inputs, and bounded as a
product.[Options] max-time:was read by the budget calculator (as
the entry's timeout) while invisible to the lint —max-time: 100000h
was lint-clean and produced a multi-year watchdog budget; it now carries
the duration cap, and a test pins the rule the hole broke (every option
the budget reads must be one the lint can see).retry-interval:— the
one uncapped multiplicand — carries the cap too. And because individually
capped values still compose into an unbounded product (retry: 10_000×
a 30 s timeout is ~83 lint-clean hours, saturating toDuration::MAX,
whose deadline addition panicked as a phantom "scenario thread panicked"
fault), the computed batch budget now clamps to an absolute four-hour
ceiling and the dispatcher's deadline arithmetic can no longer overflow.
ADR-0007 carries the amendment. -
[http] timeout-ms = 0is refused. libcurl reads zero as no
timeout, so the value opted a suite into exactly the unbounded hang the
default exists to defend against — while reading like "immediately".
Exit 2, in whichever table it appears. -
Every sink that renders run values now routes identities through the
secret masker. The event stream masksscenario,file, ...
proef v0.15.0
Published retroactively on 2026-09-09. This version was released on
2026-08-25 — the release commit is onmainand the changelog section below
was cut then — but itsv0.15.0tag was never pushed, so the release
pipeline never ran for it. This release therefore carries no binaries,
checksum sidecars or build attestations; every other release has all
three. Install 0.15.0 from source, or use a later version — 0.16.0
superseded this one six days later.
Breaking
- A quarantined test-failure reaches JUnit as
<skipped>with a message,
not<failure>— Jenkins marked builds UNSTABLE while proef exited 0;
every dashboard now says what the exit code says (ADR-0019). Library:
ScenarioSpecgainsskip,ScenarioOutcome/ScenarioRungain
reason,Event::ScenarioFinishedgains additivereason,
write_junittakes the non-gating list. --shardassignments re-deal: the hash gained a mixing finalizer. Raw
FNV-1a's low bit is the XOR-parity of the input bytes, so a scenario named
after its feature file — the commonest Gherkin convention — collapsed to
one shard atN=2and left odd shards empty atN=4, silently (the empty
shard exits 0).shard_bucketnow finalizes with Murmur3'sfmix64; every
scenario re-buckets, so all jobs of one matrix must run the same proef
version (already true in practice). Round-18 finding, reproduced and
mechanism-verified before fixing; the balance test gained the
name-mirrors-file corpus it was structurally blind to.- Tag atoms glob.
*and?in a--tags/[run] exclusive-tagsatom
are now anchored wildcards (@FRD-*selects the family;?is one
character) — previously they were literal characters that silently matched
nothing, the trap this closes. Metacharacter-free atoms are bit-identical
to before, property-pinned. Case stays sensitive. - JUnit test identity is
classname+name.classnamecarries the
feature file,namethe scenario alone; the old singlenameembedded
file:line, so an edit above a scenario re-identified every test below it
in Jenkins history and GitLab's MR diff. Anything keyed on the old
file:line namestrings must re-key. The suiteskippedcount is now
spelledskipped(wasdisabled, which no consumer reads).
Added
-
[tag-links]turns tag cells into tracker links (RF's
--tagstatlink, reduced to one mechanism): tag glob → URL template with
{tag}substituted, honored by the HTML report's by-tag table and the
GitHub summary; the pattern language is the same anchored glob--tags
uses. Library (Breaking):render_htmltakes the link map;
tags::atom_matches_publicexposes the one matcher. -
--console dotted|quiet(RF wave 3): one glyph per scenario (.
pass,Ffail,sskip,wwarn — lowercase is non-gating, the
pytest/RF convention, flushed per glyph, wrapped at 80) or just the frame.
Purely presentation: the record, every report, the post-pool failure
details and the exit code are identical in every mode;run.logmirrors
the console verbatim, dots included —events.jsonlis the full truth.
Library (Breaking):ConsoleReporter::newtakes aConsoleMode. -
A
--rerunnow produces the one JUnit and the one report that cover
the whole suite (E2's rerun half; Robot Framework'srebot --merge
shape, done as composition): the run head recordsrerun_of, the JUnit
carries the base's not-re-run scenarios as ordinary testcases, and
proef reportoverlays the base into a merged page (banner named, base
timestamps stripped so timelines never mix, rotated-away base degrades
loudly). Exit code and totals stay the rerun's own. -
--meta key=valueand[meta]/[env.<name>.meta]record explicit run
metadata (ADR-0020, RF wave 2): commit, build URL, team — recorded in
the run head, shown by the HTML report, GitHub summary,explain,
diff(which now also warns on cross-env comparisons) and the
--output jsonbody (additive keys). The active--envprofile name and
the--shufflemarker ride the same head. proef never harvests: no git,
no hostname, no CI env sniffing — the shell harvests, proef records.
Everything passes the sink-boundary mask, keys and values both. Library
(Breaking):RunRecord::openandexec::executetake the head inputs. -
Per-tag verdicts in the HTML report and the GitHub summary (RF wave
2): tags now reach the record — additivetagsonscenario_finished
(finished-only: the cancel-skip path emits no start), additive
exclusiveonscenario_started(closes R11-6, the scheduler's own
bool) — and both reports roll them up per tag (suite-only, Warned counts
with passed). Requirement-tagged suites (@FRD-3.1) get their
traceability matrix for free. Tags are deduped at the one accumulation
point (first occurrence wins); the quarantine list is now derived from
the outcomes' own tags — one owner, same behavior, pinned by the exit
suite. Library (Breaking):ScenarioSpec/ScenarioOutcomegaintags. -
@skipand@skip:<reason>park a scenario visibly (ADR-0019, RF
wave 2): counted in every total, reasoned in the console, JUnit, TAP, the
record, the HTML report,explainandflows --output json; the harness
maps it to libtest's ignored flag. All-selected-skipped exits 0; the
empty-selection refusal stays exit 2.--tags "not @skip*"unselects both
spellings; an authored skip is never re-queued by--rerun, anddiff
gives skip transitions their own bucket instead of reading them as fixed. -
flowsshows the feature description. The prose block under
Feature:was parsed and then dropped — the one paragraph written for
exactly the readerflowsserves never reached them. Human output prints
it under the feature header;--output jsonrows gain
featureDescription: string|null(additive). Library:FeatureFilegains
description. -
--shufflere-deals the execution order, seeded by the run id — one
determinism knob for order and fakes alike, so--shuffle --run-id <id>
reproduces an order-dependent failure exactly (Robot Framework's
--randomize, minus the parallel seed it threads separately). Applied
after--shard, so membership never moves; under--watchevery unpinned
rerun re-deals, deliberately. The permutation is version-stable and
pinned. Recording ashuffledmarker in the run head is deferred to the
plannedRunStartedadditions (env/metadata), one wire change instead of
two. -
The failing step's
reproduce: curl …reaches the record. The engine
always computed the redacted curl and the live console always printed it —
and the record dropped it, soexplainand the HTML report knew less than
the console did.StepFinishedgains additivereproduce_hint(absent on
passing steps and every pre-field stream);explainand the report print
it; the sink-boundary mask covers it likedetail. -
README documents every flag the binary exposes, enforced. v0.14.0
shipped--shardand--max-failwith no README mention; the docs gate
gains the flags direction (same vacuity guard as the command half), and the
measured gap — those two plusschema --add-to— is closed. -
JUnit carries what GitLab and Jenkins actually read (R3-6, specced from
GitLab's parser docs and Jenkins'SuiteResult.java):fileon each
testcase (GitLab source linking),timeon suite and root.timestampand
hostnamestay absent deliberately — ignored or substituted by both
consumers, and a hostname would undo R12-1's provenance fix. -
The docs corpus is a website: https://emrecdr.github.io/proef/. mdBook
rendersdocs/on every push tomainthat touches it; the nav is
docs/SUMMARY.md, which the existing docs gates link-check like any other
doc, and the pages workflow refuses a corpus doc that is not on the site.
The cratehomepagepoints there from the next release.
Fixed
-
A failure detail is bounded before it reaches any sink. hurl's rendered
assert error quotes the actual response, so a failed assert on a large body
rode full-size into the record, JUnit, the HTML report and the GitHub
summary at once. The engine now middle-cuts past 40 lines / 8 KiB with a
marker naming the elision; the artifact pointer survives outside the cut,
and the full output is one re-run away (Robot Framework's 40-line rule,
adopted at the boundary where all sinks are covered at once). -
The machine-body contract closes its last two paths: an empty selection
(--scenario/--tagsmatching nothing — loud exit 2 by design) and a
corrupt global-state file both emitted zero stdout bytes under
--output json. -
Identical errors collapse like identical warnings — a broken macro
usually fails to lower everywhere, so the error wall was the more common
fifty-block wall; distinct errors still render separately, and SARIF keeps
every site. -
Injected
[Options]lines respect every section-ending shape. The
section-end move covered one shape of five: an unfenced JSON/XML body after
an author[Options]swallowed the injected lines into invalid hurl (exit
2 on input that worked before), and an entry with an author section but no
response line leaked its pending lines into the next entry, where hurl
parsedretry:as an HTTP header and the artifact validated green. The
section now ends at the first line that could not sit inside it. -
A
#inside abind:value no longer hides the reads after it. The
template probe parsed the value in an unquoted position where#opens a
comment; it now probes the quotedvariable:position bake actually
injects into, so"{{a}} # {{b}}"reports both. -
A setup that fails to load still emits the machine body — the last
terminating path returning zero stdout bytes under--output json. -
**SARIF keeps ...
proef v0.17.0
Added
-
[http] cookie-store = falseruns the whole suite cookie-less — hurl
8.0's--no-cookie-store, surfaced through the table built for exactly this
class of setting. NoSet-Cookieis retained and none is replayed, which is
how a stateless API is proven stateless: the fixture-backed test is green
only because its steps assert the 403 a missing session cookie earns.This is the one
[http]key with no per-entry[Options]spelling at
all (OptionKindhas no cookie variant — verified against the enum), so
run-wide is not a compromise but the only place it can be said. With the
store off, the engine also skips both halves of the batch-split cookie
round-trip: hurl reads acookie_input_fileonly when enabling the engine,
so injecting one would be silently ignored — and there is nothing to write.
hurl's own FIXME (a handle once given cookie storage cannot lose it) never
reaches proef, becauserun_entriesbuilds its client per call (TECH-SPEC
§5) — a handle never transitions on → off.Breaking (library):
HttpDefaultsgains thecookie_storefield, so a
struct-literal construction needs the new line (..Default::default()sites
are untouched, and an absent[http] cookie-storekey changes nothing). -
--ctrf <path>— the run's verdicts as a CTRF report. CTRF
(https://ctrf.io) is the emerging JSON successor toJUnitXML for CI
dashboards, and it models in the schema whatJUnitcan only smuggle
through extensions — which is exactly the data proef already tracks: a
pass-after-retry carriesflaky,retries, andretryAttemptslisting
the real failed attempts with their (redacted) messages; every test carries
its tags and file path. One serializer off the same fold asJUnit, so the
two files cannot disagree — most visibly for a quarantined failure, which
both report as skipped with a message (ADR-0019), because a dashboard
reading "failed" beside exit 0 would contradict itself. AUser/System
fault staysfailedeven under a quarantine tag: quarantine is for flaky
tests, not broken input.The R12-3 contract applies from day one: a
[run] setupabort still writes
the file, carrying the setup scenario itself — a job gating on the report
must never see no file at all. The schema's required wall-clock
start/stopare measured at the CLI edge like every other clock read
(ADR-0015); the sans-IO core and the JSONL record are untouched — the
record remains the only record (ADR-0008). -
The HTML report answers "what is slowest". After "what failed", it is the
question a test report is most often asked, and the page could not answer it:
the timeline showed that workers were busy, never which scenarios to
attack. Every number needed was already in the fold.A ranked section, slowest first, each row linking to its own block, with the
heading reporting the share of run time the listed scenarios account for —
"3 of 40 · 71% of run time" is a decision, where a column of durations is
homework. Capped at eight: a ranking long enough to scroll has stopped
answering the question.Cost is the sum of a scenario's step durations, the same definition
timings.jsonuses for shard weights — one notion of what a scenario costs
across the whole tool. Not the wall-clock span, which includes time waiting
for a worker: a property of how the run was scheduled, and not something the
reader can go and fix.Absent when there is nothing to rank — fewer than two timed scenarios, or a
record with no injected durations at all. -
--shard-weightsbalances a shard matrix by measured duration.--shard
assigns by a frozen hash, which guarantees that adding one scenario never
re-buckets the others but cannot balance by time — and a CI matrix finishes
when its slowest shard does, so a count-split routinely leaves runners idle.
Every run that reaches its suite now writes a smalltimings.jsoninto its
run directory; CI archives that one file and each matrix job points
--shard-weightsat the same copy.The obvious design is silently wrong, and the module says so at length.
proef already retains records carrying every step's duration, so "weight by
the newest local record" looks free. But matrix jobs run on different
machines, each with its own (usually empty)runs-dir— every job would
compute a different weight table, therefore a different assignment, and
scenarios would run twice or not at all while the suite reported green.
Nothing about that announces itself. One named file shared by every job is
what makes the split a pure function of (selected scenarios, that file).Two rules place scenarios and they partition rather than compete: a
scenario the file mentions goes through longest-processing-time-first
placement, and one it does not mention falls back to the frozen hash. So a
test added after the timings were captured still runs exactly once. That is
pinned by a test that runs a whole three-way matrix — with a weights file
covering only five of nine scenarios, so both rules are exercised at once —
and asserts set equality both ways; mutating the placement by one bucket drops
two scenarios and the test names them.The weight is the sum of a scenario's step durations, not its wall-clock
span. The span includes time spent waiting for a worker, which is a property
of the run's scheduling rather than of the scenario, and feeding it back would
let one crowded run's queueing distort the next split.What this gives up is exactly what hash mode was chosen for: a balanced split
is not stable under insertion. That is what balancing means, which is why the
flag is opt-in. A missing or malformed weights file is exit 2 — falling back
silently would hand back the unbalanced split the flag was passed to avoid. -
The editor tells proef's two variable tiers apart. A pack's
hurl: |block
is the centre of the authoring experience and, to every editor, a plain YAML
scalar — inside which${…}(resolved at lower time, by proef, before any
request exists) and{{…}}(resolved at run time, by hurl) look
identical. That distinction is ADR-0005's whole model and the thing authors
most often get wrong, and no generic grammar can see it: a YAML highlighter
sees a string, and a hurl highlighter never runs because the block is not a
file. proef is the only party that knows.The server now answers
textDocument/semanticTokens/full, lighting${…}as
macro — a substitution performed before execution, which is what a macro
is — and{{…}}as variable. Both are coloured differently by every
mainstream theme, so it works without anyone configuring anything. The$${
escape stays dark, because telling an author proef will substitute text it
will in fact leave alone is worse than no highlighting.The
${…}scan isproef_core::resolve::reference_spans, walking the same
first_referencethe resolver itself uses — a second implementation of the
escape rule would drift, and the drift would show as an editor confidently
colouring literal text. The{{…}}scan lives inproef-lsprather than
core, because that spelling is the engine's and ADR-0002's amendment is that
engine syntax does not accumulate in the core.Collapsing the seven-arm request dispatch behind a local macro came with it:
the chain crossed clippy's line limit the moment an eighth feature landed, and
the honest fix was to stop repeating an identical frame seven times rather
than to suppress the lint that noticed. -
The linear-validation claim is now a test, not a sentence. #138 made pack
validation linear and recorded the result as a shape: "the curve changed
shape — 4× per doubling before, ~2× after". That number lived only in the
changelog, where nothing could re-run it — so a future span locator scanning
the whole pack file again would have restored the quadratic behaviour
silently, a regression that costs seconds rather than correctness and which no
gate measured.The guard asserts the ratio between 1000 and 2000 macros, because the
claim is a ratio. It observes ~2.05× against a bound of 3.0; mutating
locate::MacroIndexto re-index per lookup — the exact pre-#138 shape —
measures 4.01×, matching the changelog's own prediction of 4× and turning
a 0.4-second test into a 73-second one. The failure message names the cause
rather than reporting a number.A ratio rather than a benchmark, for a reason now written into
TESTING-STRATEGY.md§7: load on a shared runner inflates both measurements
together and cancels, where an absolute threshold has to be loosened until it
means nothing.iai-callgrindwould be the better CI gate — instruction
counts ignore runner noise entirely — but it needs valgrind, so it would be a
gate the maintainer cannot reproduce on macOS;criterionanddivansit in
the same noise regime as this test while adding a dependency tree to a
workspace that audits every edge. No new dependency was added. -
Every diagnostic code is now named by a test, and a guard keeps it that
way.DIAGNOSTICS.mdcalls codes "a contract: they never change meaning".
Twenty-three of seventy-five had nothing holding them to it — reachable in
production, documented, exercised by nothing at all: not a seeded corpus
directory, not a unit test, not even an assertion on their message text. They
existed only at their definition site.The catalogue itself was found exactly honest — 75 codes defined, 75
documented, and its corpus column matched disk in both directions with zero
drift. The gap was never documentation; it was that a documented promise had
no enforcement.Nineteen new tests close it, each reaching its code through a real path rather
than constructing the diagnostic directly. Two ...
proef v0.16.0
Supersedes 0.15.0, which was cut (
release: v0.15.0, 2026-08-25) but never
tagged or published — its changes are all here, and crates.io goes 0.14.0 →
0.16.0 with nothing skipped.
Fixed
-
The record-size ceiling reached two of its four readers. 0.13.0 bounded
the run-record read at 256 MiB because records travel —diffreads a
downloaded baseline,flakyreads every retained run — and the read, the
line split and the parsedVec<Event>are resident at once, so a corrupt or
hostile file was an OOM rather than an error. The bound lives in
record::read_events, andexplainandreporteach opened
events.jsonlwith a bareread_to_stringinstead, so neither had it.
reporteven used the guarded reader for the base record two dozen lines
below the raw read of the primary one.Both now go through
read_events, which returns the parsed events — exactly
the read-once/parse-once its own comment asked for. A source-scanning test
makes the next reader go through the same door, the shape this project
already uses for the raw-print and malformed-plural rules: a guard added in
one place and left for the next call site to rediscover is how it went
missing the first time.
Added
-
explain,diffanddoctorspeak--format json. They were the three
commands with no machine output, and the three a consumer reaches for
after a run. A run directory isartifacts/ + events.jsonl + run.logand
carries no structured summary, so anything analysing a run it did not launch
— a CI job reading another job's artifact, a script, an agent — had to fold
events.jsonlitself. That is the fold proef's own two internal copies
disagreed on three ways beforereport::suite_totalsunified them; handing
the canonical answer over is cheaper than inviting everyone to re-derive the
one proef got wrong.Each object mirrors its prose field for field rather than modelling a richer
view — the prose is the contract a reader already knows, and a machine
surface that says something different is a second answer to one question.
diff'sflaky/slowerstay the rendered sentences for the same reason.
The flag is the existing single-variantjsonenum the listing commands
already use, renamed fromListFormattoJsonFormatnow that it serves
non-listing commands too. Machine mode owns stdout: notes whose content the
object already carries are suppressed rather than repeated on stderr.doctorneeded a real change to get there — it printed each check as it ran,
so the verdict was the only thing a caller could see. Checks are collected
before rendering now, which makes the JSON a second rendering rather than a
second walk: the failure mode where one surface gains a check the other never
learns about. -
--console failed— thefullBDD tree, but only for scenarios that
failed or warned. A clean run prints the run line and the summary; a dirty
one prints exactly whatfullwould. The gap it fills is the CI one:
fullis a wall of green on a large suite,dotteddrops the detail you
need when something breaks, andquietdrops everything.Warned scenarios are shown, which the name does not say and the code
explains: a warned scenario is one whoseoptional:step failed,
RunSummary::passedcounts it with the passes, and the summary line has no
warned column — so a mode that showed onlyFailedwould let a run in which
something did fail print exactly what a spotless one prints. A fourth
value on the existing flag rather than a new one.
Fixed
-
cargo denyfailed on a yanked transitive crate.rand 0.10.2resolved
chacha20 0.10.1, which was yanked from crates.io; the lock now takes
0.10.2. Not the secret store's copy —chacha20poly1305pins0.9.1,
which is unaffected — so nothing about encryption changed. Found by the
gate, which is what it is for. -
proef report -owrote the author's home directory into the file built to
be shared. With the report inside the run dir the artifact links are a
bareartifacts/…; with-opointing anywhere else they were made
absolute, which resolves only on the machine that produced them — and-o
exists to put the report somewhere it will be published, which is exactly
where that path is dead. 0.13.0 scrubbed machine identity out of the run
record (R12-1); this put it back, twelve times over, in the HTML uploaded
beside it. The href is now relative to the report, which resolves everywhere
the absolute one did plus wherever report and artifacts travel together,
and in the CI shape (-o public/report.html) names nothing outside the
workspace. The href is built from path components joined with/, not
fromPath::display— Windows renders\, which is not a separator in a
URL, so a Windows-generated report's links would have been dead either way
(the absolute path it replaces had the same flaw). A report written somewhere
sharing no ancestor with the run dir
still names the directories between them — that is what a correct relative
path from there is, and it is no worse than what it replaces. -
The report's
--skipcolour failed WCAG AA, and every status pill failed
it in dark mode.--skipwas the one palette token the dark block did not
redefine: a grey chosen against#0d1117(5.48:1 there) left carrying white
text on white at 3.45:1, against a 4.5:1 threshold — on the status a
reader scans for after an interrupted run. It is now#59636e(6.11:1).Writing the guard rather than the fix found a second defect nobody had
measured:.pillpaintedcolor:#fffon the status colour, and the dark
palette's colours are tuned as text on a dark ground, so all four dark
pills sat between 2.52:1 and 3.45:1. The pill foreground is now a palette
token — white on light, the page ground on dark — putting all four between
5.48:1 and 7.5:1. A test asserts the ratio rather than the hex, so a future
palette change is free to move a colour and not free to move it below AA,
and a second test pins that both palettes define the same token set (the
absence that caused this). -
The HTML report had one heading and no outline. The timeline carried an
<h2>; the tag table and the scenario list — the body of the page — had
none, so there was nothing to navigate by and no anchor to link a section
with. Both gained one, sharing the class the timeline already used (renamed
from.timeline-hto.section-h, since it now serves three). Pinned
structurally, so a section added without a heading fails the test. -
A step's
name:label reached the artifact and nothing else. A macro
with more than one step turns one feature sentence into several engine
steps, and they share aStepRefexactly — same file, same line, same
text. The emitter has always written the authoredname:into the
artifact's entry comment, which is why the.hurlcould tell them apart;
StepRefnever carried it, so the console, the HTML report,JUnit, TAP,
the job summary andexplainall printed the same sentence once per step,
with nothing but the status glyph to distinguish a warning from the failure
beside it. The reference corpus demonstrated it: threestep_finished
events for the cookie session is exercised, byte-identical in the pinned
snapshot, are nowobtain the session cookie,optional probe (forces a split)andcookie survives the split.StepOutcomeandstep_finishednow carrylabel, exactly as they carry
fragment— the two answer neighbouring questions (which file did this
request come from / which step of the sentence is this) and travel the
same channels. Oneproef_core::report::step_labelrenders it for every
sink, so the six cannot drift. Additive on the wire: absent when a step has
noname:, so every pre-existing record still parses and re-renders
unchanged, and the event schema stays1.This retires two claims that were not true when written:
AUTHORING.md's "they anchor artifacts, events, and failure output" and
LoweredStep::label's own "(events/console)". Same class as
reproduce_hintin the R18 wave — computed all along, printed all along,
dropped by the record. -
A fragment's text ran on into the comments introducing the entry below it.
hurl attaches the blank and comment lines above a request to that request,
which is exactly what makes the# @proefbinding reliable — but it also
means an entry has two different starts: where its lines begin and where its
request begins. The scanner used one value for both, ending each fragment at
the next entry's request line, so every comment a corpus author wrote to
introduce the next request was copied into the previous fragment and from
there into the emitted.hurl. An artifact could carry
# Destructive. Operators only.while containing no destructive request at
all, andtrim_endcould not help — a comment is not whitespace. The same
applied at the end of a file, where a trailing note became part of the last
fragment. A fragment now runs from its annotation to the end of its own
request and response; the gap between two entries documents the one below it
and belongs to neither. Nothing executed differently, because hurl permits
only comments and blanks between entries — which is why it survived: the only
damage was to what the durable record says a request is.The property covering this asserted one request line per fragment, which is
blind to comments; it now also asserts that no fragment holds any of the
generator's inter-entry filler. -
explainand the HTML report disagreed about a truncated run's totals.
A record with no tailrun_finished— a run killed mid-flight — is
reconstructed by counting, and each surface carried its own version of that
fallback. On the sa...
proef v0.14.0
Fixed
--rerunafter a cancelled run continues it, instead of a false green.
--max-fail(and Ctrl-C) stop a run early with the never-reached scenarios
honestly recorded as skipped — but--rerunfiltered to failures alone, so
stop → fix → rerun ran only the old failures and reportedexit 0with most
of the suite never executed in either run. Reproduced live before fixing
(found by round-15 external review): stop at 2 of 6, fix, rerun →
2 passed · 0 failed, green, four scenarios untested. On a cancelled
base record--rerunnow runs failures plus the cancellation-skipped
tail, and says so (note: the last run was cancelled before N scenario(s) ran…); scenario-level skips only exist under cancellation, so a completed
base keeps the old semantics exactly. This also changes--rerunafter
Ctrl-C — continuing the unfinished work is what stop → fix → continue always
meant. Mutation-tested: reverting the union fails the continuation test.
Added
-
proef test --shard I/N— stable hash-mode sharding (R3-3). A CI matrix
runs--shard 1/N…N/Non separate machines; scenarios are assigned by a
frozen FNV-1a hash of the run-wide(file, scenario)identity, so adding
a scenario never re-buckets the others — the measured stability argument
that rejected index-slicing at triage (inserting one scenario re-bucketed
the whole shifted tail under slicing, nothing under hashing; the shard tests
pin both directions, and the assignment itself is frozen by literals — the
hash is a published contract, and changing it would be breaking). Sharding
applies after every other selector (the pinned filter→shard order), so
each matrix job partitions one agreed-on set. An empty shard of a non-empty
selection is a note and exit 0 — a small suite over a big matrix is a fact,
not a mistake — while an empty selection keeps the loud typo'd-filter
refusal, sharded or not. -
proef flaky— flakiness verdicts over the retained run history (R3-2).
The 2026 discipline is detect → quarantine → resolve, and proef already
owned the middle step:@quarantineruns a scenario without gating the
exit code. This is the missing detect, a fold over the recordsruns-dir
already retains — the window is[run] keep-runs, and no new state is
written. Three signals from fields the record already carries (ADR-0008):
flapping (verdict changed between consecutive observed runs more than
once — transition-counting, not fail-rate, which is what separates flaky
from broken: a scenario failing every run is consistently broken, a
different problem), passes only on retry (green, but some step needed
more than one attempt — the latent flake pass/fail-history tools
structurally miss; the record keeps per-step attempts), and always
failing. A cancellation-skipped row is not evidence and does not count
toward a scenario's history; phases are excluded (ADR-0014).--output jsonemits one object per scenario with the counts behind each verdict.
Fewer than two runs is refused (exit 2), the same answerdiffgives. -
proef test --max-fail Nstops the run after N suite-scenario failures
(1= fail fast) — the convention Playwright (--max-failures), pytest
(--maxfail) and cargo-nextest (--max-fail) share, with the shared honest
semantics: in-flight scenarios finish, the never-run rest record as
skipped (not absent, never passed), and teardown still runs on its own
token. The stop rides the graceful-cancel path Ctrl-C already exercises, so
the record is a complete cancelled run — whichdiff --fail-on-regressionalready refuses to certify, exactly right for a
deliberately-partial one.[run] setup/teardownfailures never count
toward the threshold (a broken fixture is not a failing test, ADR-0014).
Documentation
- The R3 enhancement registry is triaged (OPEN-FINDINGS):
--max-fail
built; a flakiness verdict over the run history and hash-mode sharding
validated as build-next (the 2026 flaky pipeline is detect → quarantine →
resolve, and the@quarantinetag already owns the middle step); CTRF,
pack doc and the pre-M6 seam refactors deferred with named triggers; OTel
and Cucumber-Messages exporters declined under ADR-0008's one-record rule;
items defined only in the absent v1 research document held for a spec.
proef v0.13.0
Security
-
An encoded reflection of a secret is redacted (S1). Redaction was
exact-match on the raw secret bytes, and a server that reflects a bearer
token encoded — an OAuth introspection endpoint, a debug echo, a JWT claim
— defeated it: a failing assert quoted the base64 form in its detail, and a
string triviallybase64 -d-able back to the live credential reached the
console andevents.jsonl, the retained record CI uploads. Demonstrated
live against 0.12.0 by an external research pass and reproduced here before
fixing.Redactions::newnow derives each secret's common encoded forms as
additional needles — base64 (standard and URL-safe alphabets, with and
without padding), hex (both cases), RFC 3986 percent-encoding, and the
JSON-string escape — so every construction site (the CLI sink, the engine's
internal renderer, TAP) is covered by construction. This is the remedy
GitHub's own log-masking documents for the same limitation: register each
transformed value too. The needle set covers the reversible transforms that
occur at HTTP boundaries and does not claim completeness — a secret
reflected hashed or re-encrypted matches no needle list. Over-redaction is
the accepted failure direction. Property-tested over every derived form,
pinned end-to-end by a fixture route that echoes the bearer base64-encoded,
and recorded as an ADR-0005 amendment. -
The fragment corpus read is bounded.
[run] fragmentsnames a directory
proef did not write and does not control, and it was read with no per-file or
total cap: a 279 MB file cost 601 MB of resident memory onproef flows—
a command that never looks at a fragment — because the text is read whole and
then copied into anArc<str>. A file over 8 MiB is now skipped
(proef::pack::oversized_fragment_file) and the reader stops past 64 MiB
total (proef::pack::fragment_corpus_too_large). The size comes from the
directory entry, so an oversized file is never allocated at all; the same
bound applies inproef lsp, where the corpus is held between requests rather
than for one command. Skipped, never fatal — a corpus is foreign by design, so
one bad file must not sink the ones beside it. An unreferenced corpus still
costs nothing: the scan stays lazy, so nothing is reported unless a pack
actually names a fragment. Filed as R9-3.
Added
proef difftakes a path. Each side is now a run id, a record
directory, or an events.jsonlfile under any name — the stream is the
record (ADR-0008), so all three must mean the same thing. The file form is
the CI baseline flow an adopting suite asked for: download the base branch's
events.jsonlartifact andproef diff baseline.jsonl <new> --fail-on-regressiongates the PR, with no shared record store. Previously
every argument was joined ontoruns-dir, so a path produced
.proef-runs/<your path>/events.jsonl: No such file— the argument mangled
into the complaint. A path that does not exist now names itself; a--baseline
flag was considered and declined as a second spelling of the same positional.
Internal
-
A hung test is now a five-minute failure, not a five-day zombie. The
nextest config hadslow-timeoutwith noterminate-after, which only
labels a test SLOW and never kills it — anlsp_stdiotest wedged on an
unboundedchild.wait()ran for five days with itsproef lspchild alive.
Both layers fixed: the two barechild.wait()sites got the file's own
bounded-watchdog pattern (a server that fails to exit now fails the test in
10s, naming what did not exit), and the runner gainedterminate-after = 2
(120s), sized from a cold-cache census of the whole suite (slowest ordinary
test: 5.1s). Theharness_trio — which shellscargo testinside the test
and measured 216s on a fully cold cache — gets a per-test override to 600s,
the nextest docs' own tight-global-plus-overrides pattern. The
process-group kill (a spawned server dies with its test) was verified
empirically with a deliberately hung test holding a live child. -
Cleanup pass over this cycle's four PRs (reuse/simplification/efficiency/
altitude review). The corpus-bound decision moved into core as
pack::CorpusBudget— it was abstracted in the CLI and hand-copied in the
LSP, agreeing by copy rather than by construction; both readers now share it
and only measurement stays reader-local.Redactionsstopped allocating on
the miss path (nearly every call: per string field per event under the
reporter-stack mutex, with the needle list ~9× larger since the encoded
forms) — clean fields now hand back their originalArc. A relative source
path is left exactly as it arrived, per its documented contract — it had
been falling through to a per-file canonicalize that could rewrite a
../-typed spelling. The LSP's percent-encoder folded onto core's
(byte-identical copies, one character set to drift). The fixture's
hand-rolled base64 became the crate call — its dependency-surface rationale
died when this same cycle madebase64a workspace-wide compile.diff's
path-or-id resolution moved beside its sibling inrecord. Adeny.toml
home for the curl floor was tried and reverted by mutation test:
cargo-deny 0.19.8 mismatches build-metadata versions (curl-sys@<0.4.90
banned the good0.4.90+curl-8.21.0); the floor stays a unit test, now
scanning every lockfile entry rather than the first. -
The bundled libcurl cannot silently regress under the June-2026 CVE
batch.curl-sys 0.4.90+curl-8.21.0in the lockfile is past the batch —
but only as a transitive accident of resolution, and the usual gates are
structurally blind here: RUSTSEC carries no advisories for CVEs in a
*-sys-bundled C library, socargo audit/denystay green however stale
the bundled curl is. A test now asserts the lockfile floor, and each release
build prints the libcurl actually linked into that artifact (proef doctor
already reported it; the release log now carries it per target). The hurl-8.1
watch items —variables-file:'s missing sandbox first among them — are
recorded as a pin-bump checklist in the thin-fork runbook. -
Fuzzing reaches the fragment rules.
fuzz_pack_loadran against an empty
corpus, soref:resolution,bind:keys nothing reads, abind:colliding
with a variable the fragment supplies itself, and unbound placeholders were
covered on paper and unreachable in fact. The newfuzz_fragment_binding
target is structure-aware: it builds a well-formed pack and corpus and
spends its budget on the name space where those rules live. That shape was
chosen from measurement, not taste — a byte-oriented version never once
resolved aref:in 1.45 million runs, because reaching the rules meant
discovering valid YAML and a matching corpus at the same time. The corpus is
read by a synthetic scanner rather than hurl's, which is what keeps the fuzz
workspace free of native libraries: cargo dependencies are package-level, so
one engine-dependent target would compile hurl for all of them. -
Hurl's own annotation scanner is property-tested, in
proef-engine-hurl
where the native libraries already are. The properties pin what the
entry-boundary arithmetic is for: every reported line lies inside the file,
every entry is accounted for exactly once, the starts are ordered and
distinct, and — the one that matters — no fragment's text runs into the
entry after it. That last assertion exists because a first draft without it
passed while the boundary was deliberately broken. -
The fuzz target list comes from
cargo fuzz list. It had been spelled out
inci.ymlandnightly.yml, so a new target ran nowhere until both were
edited, and nothing failed to say so.
Fixed
-
A run record no longer names the machine that produced it.
[run] suite
resolves against the config directory (0.12.0), so a path-lessproef test
handed the front end an absolute path — and every emitter printed it: the
.hurl# source:header,.map.json'sfeature.file, every
step_finishedevent, the console, and pack diagnostics. Two checkouts of one
suite stopped producing equal artifacts, which is the property ADR-0010 exists
to guarantee; an adopting suite hit it as/Users/…in 133 artifact lines and
64% of its event stream by bytes.The resolution rule was right and stands. What was missing is its naming
dual: resolve against the project, then name against the project again.
front::SourceNamingis now the one boundary that answers "how is this path
spelled", for features, packs and fragments alike — replacing the fragment
corpus's separate cwd-relative strip, which was a second anchor for the same
question. The four ways to name one suite — derived from[run] suite, typed,
typed absolutely, or reached from a subdirectory — now emit one artifact, byte
for byte.A path that arrives relative is recorded exactly as it arrived; a suite or
corpus genuinely outside the project keeps its absolute name, there being no
project-relative spelling of it. Filed as R12-1, and it closes R9-6, which had
described the same defect as safe from the project root — it no longer was.Breaking, by the rule in
docs/RELEASING.md: it changes emitted artifact
bytes, which is inherently breaking and takes a MINOR bump. Migration: nothing
to do for a suite invoked with a typed relative path — those bytes are
unchanged. A tool readingstep.fileorfeature.fileout of a record
produced by a path-less run now sees a project-relative path where it saw an
absolute one; join it onto the directory holdingproef.toml. Records written
by earlier versions are not rewritten.
Added
[run] keep-runsbounds how many past run recordsruns-dirretains. Th...
proef v0.12.0
Fixed
-
A
runs-diredited mid---watchno longer feeds the loop its own output.
Reruns re-read the config (the fix below), so records went to the new
directory while the watcher's exclusion still named the one it had frozen at
startup — and every rerun'sartifacts/*.hurl, now under an unexcluded
directory, requeued the next run. One edit produced 39 runs in 12 seconds,
firing real traffic. This was the third outing for the watch-feedback class,
so the fix removes the second answer rather than resynchronising it: each
rerun registers where it is about to write, and the exclusion is derived from
the same config the run is. A directory a previous run wrote stays excluded
too, since its events can still be in flight. Filed as R11-8. -
--watch --config <relative path>retriggers on config edits. The watcher
compared the config by exact path whilenotifyreports events under the
spelling the OS resolved them to, so--config proef.tomlnever matched and
config edits produced nothing — silently, because feature edits kept working
and the loop looked alive. Symlinked and/tmp-style aliased paths failed the
same way and are also fixed: the flag is made absolute when it is stored, and
identity is settled by comparing canonical paths, which is a stricter question
than being absolute. The same relative-path flaw silently costproef lsp --config <relative>go-to-definition across the whole fragment corpus, since
documents::name_to_urlrefuses a relative name. Filed as R11-9. -
doctorreports aproef.tomlthat will not parse. The discovery arm had
become a silentunwrap_or_default, so a malformed config leftdoctor
reporting on invented defaults and printing "all checks passed", exit 0 — with
the parse error, which the previous code printed, discarded. It is aproject:
row now, so it reachesworstand the exit code a CI script actually reads.
Being absent is still not a finding:doctormust run outside a project.
Filed as R11-10. -
proef fragmentsexits non-zero when a[run] setup/teardownphase fails
to load. It printederror: setup feature failed to validate:and exited 0,
because the phase half flattened its failure to "not measured" while the suite
half kept its code. Withholding the counts was right; reporting success while
printing errors was not. -
proef.tomlhas one path rule. A path written in the config now resolves
against the directory holding the config; a path typed on the command line
still resolves against the working directory.[run] fragmentsalready worked
this way and everything else did not, so two keys in one table meant two
different roots: from a subdirectoryfragments = "hurl"resolved while
suite = "features"reported "neither a feature file nor a directory". With
--configthe split was worse than inconsistent — pointing at a config in
another tree randry-run OKover whatever suite happened to sit beside the
shell, and never looked at the configured one.The rule now covers
suite,setup,teardown,runs-dirand thetests/
convention probe, plus two files nothing had inventoried:.proef-state.json
(the persistent World) and.proef-secrets.json(the secret store), which
were anchored on the working directory — so two shells in one project were two
Worlds and two secret stores. It is the convention Cargo,tsconfig.jsonand
pytest's rootdir all follow. Absolute values are taken as written, and with no
proef.tomlin scope written paths stay relative to the working directory, so
the config-independent reference corpus is unaffected. Filed as R11-1. -
--watchrereads the config it retriggers on. Editingproef.toml
retriggered a run that still used the snapshot loaded at startup: changing
[url] baseproduced a rerun that dutifully called the old host, and the same
went stale forjobs,[env.*]andexclusive-tags. Watching a file whose
contents you then ignore is worse than not watching it, because the rerun
reports that the edit was taken. Each rerun now re-reads the file and
re-resolves the suite from it; a config that no longer parses fails that rerun
and leaves the loop watching, since half-typed TOML is the normal state of a
file being edited. Which directories the loop watches is still fixed at
startup, so changing[run] fragmentsor[run] suiteneeds a restart to be
watched. Filed as R11-2. -
--configis honoured or refused by every subcommand.doctorprinted the
error for a missing named file and then reported on defaults, exit 0 — the
"fall back to defaults"CONFIG.mdforbids — whilefmt,init,schemaand
secretaccepted a nonexistent path silently, against the "global to every
subcommand" claim inCONFIG.md,README.mdand this file. A named file that
is not there is now exit 2 everywhere, including where nothing reads it;
doctorstays lenient about discovery, which is a different claim.secret
additionally uses the flag, since the store is the project's. Filed as R11-3.
Breaking: the secret store, the persistent World and the run records
move with the config rather than with the shell. What decides whether this
reaches you is where you invoked proef, not whereproef.tomlsits: runs
started from the project root are unchanged, but a run started from a
subdirectory used to write.proef-state.json,.proef-secrets.jsonand
.proef-runs/beside the shell, and now writes all three beside the config.Nothing is migrated, and none of it announces itself. A World written from a
subdirectory reads as empty, sosaveAs: globalvalues start over on the
first run after upgrading; stored secrets read as absent; and the old run
records are simply invisible toexplain,reportanddiff, which say "no
run records" rather than erroring. To carry them over, move
.proef-state.json,.proef-secrets.jsonand.proef-runs/from the
directory you used to run from into the one holdingproef.toml. Otherwise
re-runproef secret setand take a fresh baseline.Breaking (library):
proef_cliis not a published library surface, but for
the recordfront::runtakes the state-file path,ProjectConfig::runs_dir
returns aPathBuf,setup/teardownreturnOption<PathBuf>,suiteis
gone (fold intodefault_suite_path), and thesecretstoreentry points take
the store path.proef_coregains one item:
pack::FragmentCorpus::unreadable_file.
[run] exclusive-tagsvalidates itself.--dry-rundid not parse the
expression at all, so a malformed one exited 2 fromproef testand passed
dry-run OK … 0 warning(s)from the gate CI runs. And a well-formed
expression matching no scenario was silent:@solozagainst a@solosuite
put every scenario back in the shared pool, exit 0, nothing said — the exact
silent degradation the key was designed as a config expression to prevent, and
one that reads as flakiness rather than as a typo. Both paths now parse it, and
a zero-match expression warns, naming it and pointing atproef flows. Judged
over every scenario the suite loaded rather than the ones selected, so a
--tagsfilter that removes the matches from one run is not reported as a
broken setting. Filed as R11-4 and R11-5.
Changed
-
proef fragmentssays which half it could not measure.--checkreported
"needs a suite that binds" when the suite had bound perfectly well and a
[run] setup/teardownfeature was the thing that failed to load, sending the
reader to inspect the half that was fine. The degraded listing also now carries
the notemacrosprints, so withheld counts read as "not measured" rather than
as a corpus nothing uses. -
proef fragments --checkrefuses to pass with no corpus configured. With
[run] fragmentsunset it printed0 entriesand exited 0, indistinguishable
from a fully-used corpus — so a CI gate disarmed silently the day the key left
the config. The listing still works; only the gate is now a user error. -
proef fragments --output jsoncarriesannotatedon both row shapes. The
annotated and unannotated rows differ in eight fields, and consumers had to
probe for the absence of one to tell them apart.
Documentation
CONFIG.md's "everything else keeps running atjobswidth" was false:
queueing is strict FIFO, so nothing new starts while an exclusive scenario
waits at the head. The cost is bounded, not absent, and is now described.- The one caveat
[run] exclusive-tagscarries is written down inCONFIG.mdand
ADR-0007: exclusivity is enforced against the dispatcher's active set, which a
watchdog-abandoned scenario leaves while its detached thread is still issuing
requests (hurl cannot be cancelled mid-entry). TECH-SPEC§10 gainedproef fragmentsand the global--config; §11's
[run]inventory listed three of seven keys.DIAGNOSTICS.mdcarried apack::loadrow nothing emits — a reader who
grepped it found a plausible cause that could never be one — and filed
lower::multiline_bindunderproef::pack::*. Both fixed, and the two-way
agreement between the file and the emitted codes is now a test, since this
drifted twice.OPEN-FINDINGSR9-2 still saidfuzz_tag_expr"sits in neither fuzz loop"
three sections after recording that it is in both.
proef v0.11.1
Fixed
-
An output path creates the directories it names.
--junit,--sarifand
report -ofailed when the parent directory did not exist, while
artifacts -oand the run directory created theirs — no rule, four sites
deciding separately, with the two used most in CI on the failing side. Every
adopter paid the samemkdir -p.pytest --junitxml,jest-junit,
cargo-nextest's JUnit store and thehurlproef embeds all create them.
This does not weaken the "side effects should be explicit" principle: that is
about writing files the user did not name, and here they named exactly this
path. -
proef fragmentscounts[run] setup/teardownusage. A fragment only a
phase feature reached was reportedUNREACHABLE — no macro refs it, which was
false, and failed--check— a false CI failure in the workflow--check
exists for. The verdict also depended on where the phase file sat: inside the
suite directory it was discovered as an ordinary feature and counted. The
listing's universe now matches the runner's, and a phase that fails to load
withholds every count rather than guessing. Filed as R10-2. -
One predicate answers "is this a fragment file?" (
FragmentSupport::claims).
Three answered it before — CLI discovery viaPath::extension, the core scan
viarsplit('.'), and the LSP's corpus invalidation case-insensitively — so
they disagreed aboutapi.HURL(the editor rebuilt its corpus for a file
nothing would scan) and about a dotfile named.hurl. Filed as R10-3. -
--configreachesproef lspand--watch. The flag bypasses the
upward search so aproef.tomlbeside the suite becomes usable — but the
editor re-discovered its own config and--watchwatched whatever a fresh
search found. So in exactly the layout the flag exists for,
proef test --config …ran green while the editor reported everyref:as
unknown, and editing the config driving the run never retriggered it.
ProjectConfignow keeps the file it was read from (withrootderived from
it rather than stored beside it), and both consumers use the config actually
in force. Forproef lspthe flag also outranks the client-announced
workspace root. Filed as R10-1.
proef v0.11.0
Added
-
[run] exclusive-tags— a tag expression selecting scenarios that run
with the pool to themselves. Real suites contain scenarios that cannot run
beside anything: one asserting absolute positions (items[0]) needs a store
no concurrent scenario writes to, and the only workaround was several CLI
invocations driven by tag discipline in a Makefile, each producing its own run
record, JUnit file and exit code to aggregate in shell.A matching scenario waits for the pool to drain, runs alone, and the pool
refills after it; everything else keeps running atjobswidth, and discovery
order is unchanged so an exclusive scenario never loses its place. A config
expression rather than a reserved tag name, because with a bare convention a
scenario added months later lands untagged in the parallel pool and breaks
isolation intermittently — which reads as flakiness rather than as a missing
declaration. A malformed expression is a user error, never a silently-ignored
key.This is exclusion, not ordering: a scenario that must run before the
rest belongs in[run] setup, which already runs once before the pool exists.
Deliberately one axis of the twocargo-nextestsettled on — per-group
concurrency limits (rate-limiting a shared dependency) are a real future need
that nobody has asked for, and a group table can be added later without
breaking this key. -
proef fragments— the corpus listing, symmetric withmacros. Until now
no proef output stated how many fragments there were, so neither way a
fragment can die had a denominator to be noticed against: one no macro
references was unobservable, and one reached only through a macro no scenario
binds looked covered because the macro was flagged. Both are now named
apart, unannotated entries are listed by line (they have no name to list by),
and--checkexits 1 when something never runs.--require-annotatedextends
that to unannotated entries and is deliberately opt-in: an unannotated entry is
inert by design (ADR-0018), so "not done yet" is a porting team's meaning, not
every adopter's. Reachability is read off the lowered scenarios, so a fragment
reached through a chain ofuse:counts as reached. -
--config <path>, global to every subcommand, naming theproef.tomlto
read instead of searching up from the working directory. Discovery only goes
up, so a config beside the suite is unreachable from the repository root — a
layout an adopting team planned and abandoned after it failed. A named file
that does not exist is a user error rather than a fall back to defaults:
discovery finding nothing means "no project here", but a named path that is
not there is a typo, and a silently unconfigured run is what that used to buy. -
proef doctorsees the fragment corpus — a row reporting how many
fragments loaded from[run] fragments, warning when the configured root is
not a directory. A misconfigured path used to surface much later as
pack::unknown_ref: an error about a name when the cause is a path. -
proef initscaffolds both body forms — a one-entry.hurlfile with a
# @proefannotation,[run] fragments, and a pack macro of each kind. The
newcomer with most to gain fromref:is the one who already owns a hurl
corpus, and a scaffold teaching onlyhurl: |reads as "proef wants your
files transcribed into YAML".
Fixed
-
A
bind:key nothing reads is refused (proef::pack::unread_bind_key),
with did-you-mean over the names actually in scope.bind_without_refonly
caught a table with noref:at all, sobind: { token: …, toekn: … }
validated clean — the one authoring mistake in the fragment path that produced
no signal whatsoever. Checked as a union over the scope, never against one
fragment: a pack-scope table is the plumbing every macro in the file needs, so
a key serving one macro and not its siblings stays correct. -
duplicate_fragmentno longer says "in bothxandx" for two entries
in one file, and stops offeringfile.hurl#nameas the remedy there — that
qualifies by file and cannot separate two entries inside one. Annotating a
corpus adds many names to few files, which makes same-file the likely
collision. -
unbound_placeholdernames all three supply routes. The omitted one was
the fragment's own[Options] variable:— the route that makes a corpus file
runnable standalone, which is the property ADR-0018 exists to preserve. -
A fragment's
[Options]escaped the ADR-0007 value caps.retry: -1,
repeat: -1and an unboundeddelay:were rejected in an inlinehurl:
block and accepted in aref:fragment — byte-identical text, exit 2 one
way and "dry-run OK, 0 warning(s)" the other, then written verbatim into the
executed input. The scan lived inside the inline-only linter; only the
twinned-option half of pass 6 had crossed to fragments. It reads the text
alone, so it now runs against a fragment's too, anchored on theref:line
and naming the fragment file and line. This is the case the caps exist for:
hurl has no cancellation, so an infinite retry makes the batch budget
unestimatable and leaves the watchdog abandoning a thread it cannot stop. -
A step declaring both
ref:and a payload was told, falsely, that its pack
had noref:at all. The conflicted step is reported and dropped, so the
loaded bodies stop showing everyref:the author wrote — and the pack-scope
bind_without_refcheck then drew a conclusion from the gap. It now infers
nothing from a pack whose steps did not all normalize. -
A pack-scope
bind:with noref:anywhere was silently dropped.
AUTHORING.mdsaidbind_without_refapplies "at every scope" while only the
macro and step scopes were checked — and a setting ignored in silence is the
bug those two exist to refuse. The check was the better half of the
disagreement, so the pack scope now has it too. -
A multi-line
bind:value blamed the artifact. A hurl
[Options] variable:value is a single-line scalar, so a newline could never
reach the entry — but it surfaced one stage later asemit::invalid_artifact,
pointing at generated text the author never wrote. Refused by name at lower
time aslower::multiline_bind, naming the inlinehurl: |form that is
what splices a multi-line body (ADR-0018's splicing-versus-binding boundary,
enforced where it can be explained).
Changed
-
Breaking (library):
AnalyzeCtxtakes the fragment corpus instead of
building one. Building it internally meant a fresh scan memo per call, so the
LSP re-read and re-hurl-parsed the whole corpus on every request — each
completion popup, each go-to-definition, each debounce tick. The server now
holds one and rebuilds it only when a fragment file changes; editing a pack or
a feature, which is nearly every keystroke, leaves it alone. It is also what
core purity already required: the caller does the IO. -
Breaking (library):
StepKindSpecgainedoptions, an engine-contributed
recogniser mapping a raw option key to what ADR-0007's budget rules should make
of it. The fragment half of that rule already crossed the seam while the inline
half matched"retry-interval:"as a literal insideproef-core— one rule at
two altitudes, and a second engine would have had its fragments linted and its
inline blocks not. A kind contributing no recogniser is not linted, since the
core has no way to know what its option keys mean. -
Breaking (library):
proef_core::engine::FragmentScannerreturns
ScannedFile { fragments, unannotated }rather thanVec<ScannedFragment>.
An engine's scanner now also reports the 1-based lines of entries carrying no
annotation — lines only, never built-then-discarded fragments, so a foreign
corpus still costs a push per unannotated entry. Without it "which entries did
I forget to annotate?" is unanswerable: a missing annotation produces a green
run and a silently absent test, and the entry that would prove it was never
built.FragmentCorpusgainsfragments(),unannotated()and
diagnostics(), because the scan is gated on some pack naming a fragment —
soPackSet::fragmentsis empty for exactly the suite a listing has most to
say about.
Documentation
-
Config discovery is a requirement, not a convention.
proef.tomlis found
by searching up from the working directory, so a config beside the suite
(tests/proef/proef.toml) is never found from the repository root — an
adopting team planned that layout and discovered it by failure. CONFIG.md now
says so, and notes that keeping the file at the root collapses the one place
[run] fragments(config-relative) andsuite/setup/teardown/runs-dir
(cwd-relative) differ. -
The release runbook could not work as written.
mainis a protected
branch, and step 4'sgit push origin main --follow-tagsfails in the
dangerous direction:--follow-tagsis not atomic, so the branch is rejected
while the tag still lands — and the tag is whatrelease.ymltriggers on,
starting a release build from a commit that is not onmain. It happened
cutting 0.10.0. The runbook now routes the release commit through a PR and tags
the merged commit, and thecargo publishsection carries the dry-run,
tag-check and--lockedsequence plus why only four crates go
([workspace.package] publish = falseis the default). Also drops step 1's
reference to changelog "bottom links", which do not exist.