Repository navigation
proef v0.17.0
Added
-
[http] cookie-store = falseruns the whole suite cookie-less — hurl
8.0's--no-cookie-store, surfaced through the table built for exactly this
class of setting. NoSet-Cookieis retained and none is replayed, which is
how a stateless API is proven stateless: the fixture-backed test is green
only because its steps assert the 403 a missing session cookie earns.This is the one
[http]key with no per-entry[Options]spelling at
all (OptionKindhas no cookie variant — verified against the enum), so
run-wide is not a compromise but the only place it can be said. With the
store off, the engine also skips both halves of the batch-split cookie
round-trip: hurl reads acookie_input_fileonly when enabling the engine,
so injecting one would be silently ignored — and there is nothing to write.
hurl's own FIXME (a handle once given cookie storage cannot lose it) never
reaches proef, becauserun_entriesbuilds its client per call (TECH-SPEC
§5) — a handle never transitions on → off.Breaking (library):
HttpDefaultsgains thecookie_storefield, so a
struct-literal construction needs the new line (..Default::default()sites
are untouched, and an absent[http] cookie-storekey changes nothing). -
--ctrf <path>— the run's verdicts as a CTRF report. CTRF
(https://ctrf.io) is the emerging JSON successor toJUnitXML for CI
dashboards, and it models in the schema whatJUnitcan only smuggle
through extensions — which is exactly the data proef already tracks: a
pass-after-retry carriesflaky,retries, andretryAttemptslisting
the real failed attempts with their (redacted) messages; every test carries
its tags and file path. One serializer off the same fold asJUnit, so the
two files cannot disagree — most visibly for a quarantined failure, which
both report as skipped with a message (ADR-0019), because a dashboard
reading "failed" beside exit 0 would contradict itself. AUser/System
fault staysfailedeven under a quarantine tag: quarantine is for flaky
tests, not broken input.The R12-3 contract applies from day one: a
[run] setupabort still writes
the file, carrying the setup scenario itself — a job gating on the report
must never see no file at all. The schema's required wall-clock
start/stopare measured at the CLI edge like every other clock read
(ADR-0015); the sans-IO core and the JSONL record are untouched — the
record remains the only record (ADR-0008). -
The HTML report answers "what is slowest". After "what failed", it is the
question a test report is most often asked, and the page could not answer it:
the timeline showed that workers were busy, never which scenarios to
attack. Every number needed was already in the fold.A ranked section, slowest first, each row linking to its own block, with the
heading reporting the share of run time the listed scenarios account for —
"3 of 40 · 71% of run time" is a decision, where a column of durations is
homework. Capped at eight: a ranking long enough to scroll has stopped
answering the question.Cost is the sum of a scenario's step durations, the same definition
timings.jsonuses for shard weights — one notion of what a scenario costs
across the whole tool. Not the wall-clock span, which includes time waiting
for a worker: a property of how the run was scheduled, and not something the
reader can go and fix.Absent when there is nothing to rank — fewer than two timed scenarios, or a
record with no injected durations at all. -
--shard-weightsbalances a shard matrix by measured duration.--shard
assigns by a frozen hash, which guarantees that adding one scenario never
re-buckets the others but cannot balance by time — and a CI matrix finishes
when its slowest shard does, so a count-split routinely leaves runners idle.
Every run that reaches its suite now writes a smalltimings.jsoninto its
run directory; CI archives that one file and each matrix job points
--shard-weightsat the same copy.The obvious design is silently wrong, and the module says so at length.
proef already retains records carrying every step's duration, so "weight by
the newest local record" looks free. But matrix jobs run on different
machines, each with its own (usually empty)runs-dir— every job would
compute a different weight table, therefore a different assignment, and
scenarios would run twice or not at all while the suite reported green.
Nothing about that announces itself. One named file shared by every job is
what makes the split a pure function of (selected scenarios, that file).Two rules place scenarios and they partition rather than compete: a
scenario the file mentions goes through longest-processing-time-first
placement, and one it does not mention falls back to the frozen hash. So a
test added after the timings were captured still runs exactly once. That is
pinned by a test that runs a whole three-way matrix — with a weights file
covering only five of nine scenarios, so both rules are exercised at once —
and asserts set equality both ways; mutating the placement by one bucket drops
two scenarios and the test names them.The weight is the sum of a scenario's step durations, not its wall-clock
span. The span includes time spent waiting for a worker, which is a property
of the run's scheduling rather than of the scenario, and feeding it back would
let one crowded run's queueing distort the next split.What this gives up is exactly what hash mode was chosen for: a balanced split
is not stable under insertion. That is what balancing means, which is why the
flag is opt-in. A missing or malformed weights file is exit 2 — falling back
silently would hand back the unbalanced split the flag was passed to avoid. -
The editor tells proef's two variable tiers apart. A pack's
hurl: |block
is the centre of the authoring experience and, to every editor, a plain YAML
scalar — inside which${…}(resolved at lower time, by proef, before any
request exists) and{{…}}(resolved at run time, by hurl) look
identical. That distinction is ADR-0005's whole model and the thing authors
most often get wrong, and no generic grammar can see it: a YAML highlighter
sees a string, and a hurl highlighter never runs because the block is not a
file. proef is the only party that knows.The server now answers
textDocument/semanticTokens/full, lighting${…}as
macro — a substitution performed before execution, which is what a macro
is — and{{…}}as variable. Both are coloured differently by every
mainstream theme, so it works without anyone configuring anything. The$${
escape stays dark, because telling an author proef will substitute text it
will in fact leave alone is worse than no highlighting.The
${…}scan isproef_core::resolve::reference_spans, walking the same
first_referencethe resolver itself uses — a second implementation of the
escape rule would drift, and the drift would show as an editor confidently
colouring literal text. The{{…}}scan lives inproef-lsprather than
core, because that spelling is the engine's and ADR-0002's amendment is that
engine syntax does not accumulate in the core.Collapsing the seven-arm request dispatch behind a local macro came with it:
the chain crossed clippy's line limit the moment an eighth feature landed, and
the honest fix was to stop repeating an identical frame seven times rather
than to suppress the lint that noticed. -
The linear-validation claim is now a test, not a sentence. #138 made pack
validation linear and recorded the result as a shape: "the curve changed
shape — 4× per doubling before, ~2× after". That number lived only in the
changelog, where nothing could re-run it — so a future span locator scanning
the whole pack file again would have restored the quadratic behaviour
silently, a regression that costs seconds rather than correctness and which no
gate measured.The guard asserts the ratio between 1000 and 2000 macros, because the
claim is a ratio. It observes ~2.05× against a bound of 3.0; mutating
locate::MacroIndexto re-index per lookup — the exact pre-#138 shape —
measures 4.01×, matching the changelog's own prediction of 4× and turning
a 0.4-second test into a 73-second one. The failure message names the cause
rather than reporting a number.A ratio rather than a benchmark, for a reason now written into
TESTING-STRATEGY.md§7: load on a shared runner inflates both measurements
together and cancels, where an absolute threshold has to be loosened until it
means nothing.iai-callgrindwould be the better CI gate — instruction
counts ignore runner noise entirely — but it needs valgrind, so it would be a
gate the maintainer cannot reproduce on macOS;criterionanddivansit in
the same noise regime as this test while adding a dependency tree to a
workspace that audits every edge. No new dependency was added. -
Every diagnostic code is now named by a test, and a guard keeps it that
way.DIAGNOSTICS.mdcalls codes "a contract: they never change meaning".
Twenty-three of seventy-five had nothing holding them to it — reachable in
production, documented, exercised by nothing at all: not a seeded corpus
directory, not a unit test, not even an assertion on their message text. They
existed only at their definition site.The catalogue itself was found exactly honest — 75 codes defined, 75
documented, and its corpus column matched disk in both directions with zero
drift. The gap was never documentation; it was that a documented promise had
no enforcement.Nineteen new tests close it, each reaching its code through a real path rather
than constructing the diagnostic directly. Two of them exercise guards that
are unreachable in normal operation and were therefore the most valuable to
test:lower::kind_unroutedfires only when the engine registry and pack
validation disagree, so the test makes them disagree on purpose; and
lower::expansion_too_deepsits behind pack validation's identical depth
limit, so the test bypasses validation withload_collecting— the only way
to hand lowering a graph validation would have stopped, and therefore the only
way to prove the second line of defence is still there.Two codes are exempted by name, with reasons recorded in the guard:
source::unreadableandconfig::unreadableneed a file the process may stat
but not read, a permissions state CI runners do not reproduce because they run
as root. The guard also checks its own exemption list, failing if an exempted
code is deleted or renamed — an exemption that outlives its code silently
excuses nothing.The guard joins the four in
source_guards.rsand is mutation-verified:
rewriting one test to match a code by suffix instead of naming it turns the
guard red, which is the point — a test that matches the prose pins the
wording, and only one that names the code pins the contract. -
[http]now carries the settings that describe an environment: TLS, proxy
and mTLS. The table exposed two of hurl's runner options —timeout-msand
follow-location— while the embedded engine has supported the rest all
along;TECH-SPEC.md:235even listedinsecureamong whatRunnerOptions
carries. So a suite that had to run against staging's self-signed certificate,
or through a corporate proxy, or against an mTLS-protected API, could not say
so anywhere: the only route was repeating an[Options]block inside every
macro's raw hurl, which defeats environment profiles exactly where they are
most useful, since these settings are the difference between environments.Eight new keys —
insecure,proxy,no-proxy,cacert,client-cert,
client-key,max-redirs,user-agent— each merging field-wise through the
existing[http]<[env.<name>.http]chain, so a staging profile turns
verification off without production inheriting it. No new concept: only more
of one that already worked.Three deliberate edges.
insecure = truewarns on every run, naming the
profile that set it — a suite that goes green without verifying a certificate
has not proved what a green suite normally proves, and since the run record
carries no config by design, the warning is the entire audit trail. A
client-keywithout aclient-certis exit 2 rather than a pass-through:
libcurl accepts the pair and then presents nothing, so the failure would
otherwise surface at the server as an authentication error naming nothing
about the cause. And credentials are excluded on purpose — there is no
userornetrckey, because a password belongs in the secret store where it
is encrypted at rest and masked out of every sink.The three path-valued keys resolve against
proef.toml, the one-path rule
every other config path follows; core still reads no filesystem and receives
them already resolved (ADR-0012). Each option is applied to hurl's builder
only when actually set, so a project with no[http]table runs
byte-identically to one built before the keys existed — pinned by a test.
Per-entry[Options]still override all of them exceptuser-agent, for
which hurl has no per-entry option at all; that exception is documented rather
than papered over.Breaking (library):
proef_core::engine::HttpDefaultsgains eight fields and
losesCopy— it now carriesStrings.Defaultstays hand-written, and
the reason is now stated in the type: a derive would maketimeout_mszero,
which libcurl reads as no timeout at all, silently converting ADR-0007's
budget into an unbounded wait at every existingdefault()call site.
Changed
-
The toolchain policy is stated in the spec that
rust-toolchain.toml
cites. R18-2 corrected the policy to latest stable, adopted at its
x.y.1point release, and the correction reachedRELEASING.mdand
CLAUDE.mdwhile TECH-SPEC §15 — named byrust-toolchain.tomlas its
authority — still said "always latest stable". R18-2's own conclusion was
that an unwritten policy contradicting the written one is a docs defect;
fixing it in two files and leaving the source of truth contradicting itself
reproduced the defect one level down. Now consistent across all four. -
An artifact is named by its feature's path, not its stem — two scenarios
can no longer claim one file. Slugs were{stem}--{scenario}, dropping the
directory, sofeatures/x.featureandfeatures/sub/x.featureeach with a
same namescenario both producedx--same-name: the second artifact
silently overwrote the first while the CLI reported writing two. Silent
loss of the hand-off ADR-0010 calls a contract — and the project already
treats same-named scenarios across files as real, which is what
--scenario-fileexists for. The same slug drives the HTML report's anchors
and artifact links,reproduce:lines, and harness trial names, so all of
them move together off the one helper.Names are now
features-sub-x--same-name. Derived from the path rather than
disambiguated on collision, deliberately: a counter or hash appended only
when two names clash would make one scenario's artifact name depend on
whether some other file exists, so adding a feature would rename an
unrelated artifact — the instability--shard's frozen hash exists to
avoid. The path fed in is the portable suite-relative name the record
carries, never a path off the running machine.Breaking, and quietly so for library callers:
emit::artifact_slugkeeps
its(&str, &str) -> Stringsignature while its first argument changes
meaning from stem to feature path, so the API gate cannot see it —
passing a stem still compiles and now yields a different name.
emit::emit's second parameter changes the same way, andemit::feature_stem
is removed (it had no remaining consumer). Artifact filenames and report
anchors change for every suite; the snapshot corpus was regenerated under the
new names and reviewed.The unification that made that a one-line change came first: the stem
expression (file_stem, falling back to"feature") had existed four times
across both crates — the emitter's caller, the dispatcher's spec naming, the
HTML report's anchors, the editor's analysis — and thestem--scenario
composition twice, with the report's links to artifact files resolving only
because both sides happened to derive the same name. Worklist item Q6 called
the four sites a future-drift risk; collapsing them to one helper is what
let the collision above be fixed in a single place instead of four. In the
same pass, Q2 (the editor's per-request walks) was found already closed by
the #146 analysis cache, and its entry now says so with the evidence. -
"What a scenario costs" is defined once, as
ScenarioOutcome::cost. The
sum of a scenario's step durations was computed in three places on the same
type —JUnit's per-suite time,JUnit's per-case time, and the new
timings.jsonweights — plus a fourth over the record-fold shape in the HTML
report. Four surfaces free to drift apart about a number they are supposed to
agree on, and the argument for summing steps rather than taking a wall-clock
span was written out twice.Now a method on the type that owns the steps, with the rationale stated there
and referenced from the rest. The one behaviour change is a fidelity gain: the
weights file used to truncate each step to whole milliseconds before summing
and now truncates the sum, so its numbers agree with the timesJUnithas
always reported. Additive to the library surface. -
A run whose setup aborted no longer leaves shard weights behind.
timings.jsonwas written from inside the CI-report block, which a setup
abort also reaches — with the setup phase's summary. The file that came out
named setup scenarios, and a weights file naming them is worse than no file:
those identities never appear in a suite run, so they absorb bucket load on
behalf of scenarios that never run and skew the very split--shard-weights
exists to balance, silently. The write moved to the one site where the summary
is the suite's, pinned by a test that reproduces the old file. -
lower.rsstops threading the same three values through twelve
functions.out,refsandsinkstravelled as separate parameters
everywhere, and five functions —expand_macro,expand_step,
expand_ref_step,expand_payload_step,finish_step— carried 8 to 11
parameters each behind individual arity suppressions. Adding one piece of
lowering state meant editing five signatures and five call sites, which is the
shape of change that drops a parameter at one site.Two bundles, both of them types that were already implied by the code:
Emit { out, refs, sinks }(the mutable outputs, always passed together and
never independently),StepScope { step_ref, ctx, at }(what stays fixed for
one authored step however deep expansion recurses), and a smallFinished
for the four values that describe a step being completed.What was not done matters as much. The obvious refactor — hoist the
state into aselfand make the five methods — would have broken the reason
they are parameters at all:resolve_inand friends take them explicitly so
they remain callable while other state is mutably borrowed, and a method on
&mut selfcannot be called whileselfis borrowed elsewhere. The threading
discipline is load-bearing, so it stays; only the arity changes.Arity suppressions across the workspace: 13 → 6, with
lower.rsat zero.
No behaviour change, and the 241 core tests say so. -
ADR-0002 now names the core's hurl entry grammar, and a guard keeps it
closed. "Adding an engine leavesproef-corediff-empty" was true of
engine-types and never of engine-syntax: the core does text surgery on
entries — splicing[Options]in, merging anexpect:block's asserts into
the previous entry — so it has to find an entry boundary in text hurl will
later parse. The worklist carried the gap for two rounds as "~290 lines of
hurl grammar in core", a figure that counted#[cfg(test)]fixtures.Measured: twelve literals across four files. Seven the core writes, four it
recognises to find a boundary, and one it quotes — a hurl snippet inside
a did-you-mean help string inbind.rs, which generates nothing and parses
nothing but drifts like any other copy. The four boundary recognisers are
already one sharedpub(crate)set. proef's own pack keys are shaped like
option lines and are excluded by name rather than listed as sanctioned rows.The guard lexes whole files. The first version scanned line by line and so
could not see a literal that spans lines — which is where a larger piece of
engine syntax would naturally be written, and where the one entry above that
nobody had counted was in fact sitting.The amendment sanctions that set and closes it. Deferred with a named
trigger — a second engine being scheduled — is moving the written half behind
the seam, where the reading half already lives:
StepKindSpec::optionsexists precisely so an engine's option spellings stay
out of the core, and it covers recognising them only, soretry:,
retry-interval:,delay:andvariable:are still core literals. Until a
second engine exists that migration relocates seven literals that exactly one
implementation will ever supply, at the cost of a public-API break.crates/proef-cli/tests/source_guards.rs(renamed fromstderr_hygiene.rs,
which had not been only about stderr for two rules now) pins the set: a new
token, or an existing one spreading to another core module, fails the test
and names both remedies. A claim of this shape decays the moment it is only
prose — this one already had, by an order of magnitude, in the direction that
made it look worse than it is.
Fixed
-
A
file,…;body in aref:fragment resolves where its author put it.
hurl resolves a file body against the directory of the file that wrote the
reference — its--file-rootdefault, and the same rule Karate, pytest and
Jest use for fixtures. proef resolved every asset against the feature,
and a fragment lives in another tree entirely ([run] fragments), so the
same bytes passed under stockhurland failed under proef, as exit 2,
blaming the author for a path that was correct. Nothing worked around it:
moving the file beside the feature breaks the standalone run ADR-0018
exists to guarantee, a reaching../path is refused by hurl's own sandbox,
and the advice that refusal prints — "check --file-root option" — names a
flag proef does not expose.Each asset is now staged from beside the source that referenced it, feature
or fragment, into that scenario's own asset root, which is what the engine
gets as its context dir. Staging rather than two roots because hurl offers
one context dir per run of entries and no per-entry override, while a
single batch may mix both body forms — measured, not assumed: an inline step
and aref:step in one macro lower to one batch. Copying fixtures into the
build output is the standard answer to exactly this, and it adds no copy
operation: the record already copied these files once per scenario, just
into a shared directory instead of the right one. What it does change is the
footprint — an asset N scenarios read is now N files in the run record
rather than one, which is the same fact as the collision below, seen from
the disk's side rather than the reader's. -
Two scenarios' assets no longer overwrite each other. Staging was flat
and keyed by the asset's bare name, so two features that each keep a
data.jsonbeside them staged to one file — last writer wins, with "0
warning(s)" — and the loser's artifact replayed against the other's bytes.
artifact_slugalready refuses that trade for the.hurltext, deriving
from the feature's whole path so two same-named scenarios cannot collide;
the files it reads now get the same treatment. An artifact that reads a
file says so in its replay line (--file-root assets/<slug>); one that does
not is byte-identical to before. Two sources claiming one name inside a
single scenario — the case a per-scenario root cannot separate — is refused
rather than narrowed.A missing asset is also an error now instead of a silent skip. It had to
become one: the staged root is what the engine reads, so a file that quietly
failed to arrive is no longer an incomplete record but a request reading
nothing.[Options] output:resolves through the same root, so the root is created
for every scenario rather than by the staging loop — which never runs for a
scenario that reads no file body. A response written that way now lands
inside the run record, where a run's outputs belong, instead of in the
feature's own directory.Breaking (library):
emit::file_referencesis replaced by
Artifact::assets, aVec<AssetRef>carrying each reference with the
source that wrote it — the provenance a whole-artifact text scan destroys,
and the whole reason the bug was expressible.emit::asset_rootnames the
staging directory for the three call sites that must agree on it, and
pack::split_qualifiedis now the one reader of thefile.hurl#nameform
Fragment::qualifiedwrites — there were two, resolving aref:and a
use:, and staging assets was about to make a third in another crate. New
diagnostic:proef::run::asset_unstageable. -
A
--run-idrecord is findable again (ADR-0021).--run-id pr-1234
writes a perfectly good record, and every command that resolves the latest
run —explain,diff,flaky,report,--rerun— enumerated by the
uuid shape, so that record was invisible to all of them.--rerunwas the
sharp edge: it silently continued some older run instead of the one just
produced.One predicate had been answering two questions whose risks point in opposite
directions — may I delete this? is unsafe when broad, is this a run I can
show you? is unsafe when narrow — so the deletion-safety choice had
silently become a visibility choice. They are now separate: a directory is a
record because it holds anevents.jsonl, while rotation still deletes
only uuid-named directories, so a custom-id run is discoverable and still
never reclaimed by[run] keep-runs. Ordering stopped riding on the name
too — uuid-v7 sorted chronologically until a custom-id directory joined the
set and sorted by its first letter — and now takes the timestamp a uuid-v7
name carries (48 bits of unix milliseconds, which is why the lexical sort
worked), falling back to directory mtime for a name that carries none. -
A run with a failed
optional:step no longer prints exactly like a
spotless one.ConsoleMode::Failed's own doc comment states the
requirement and theWarnedarm implementing it was unreachable: a
scenario's aggregate status was only everFailed | Skipped | Passed, so a
real optional failure was invisible under--console failed, showed a.
rather than the documentedwunder--console dotted, and left the HTML
report's warned count and its filter-bar warned button permanently empty —
four consumers and three docs describing something that could not occur.
Steps carriedWarned; scenarios never did. The aggregate now promotes,
which changes what a run says and never whether it gates:Warnedcounts
as passing in the exit code, the totals andJUnit. -
explainprints a step's authoredname:, like its five siblings.
step_label's own doc enumerates the six surfaces that must render it —
console, HTML,JUnit, TAP, the job summary,explain— andexplainwas
the one that never called it, so the post-mortem tool showed one sentence
repeated where the live console had told the steps apart. -
The HTML "Slowest" section no longer counts
[run] setup/teardowninto
"% of run time". Every other aggregate on the page excludes phases
(ADR-0014), including the tag table directly above it, so the share meant
something different in that one section. A slow phase stays visible in the
timeline and in its own block. -
.cargo/audit.tomlno longer suppresses advisoriesdeny.toml
deliberately un-suppressed. It carried the quick-xml pair
(RUSTSEC-2026-0194/0195) with a comment claiming it mirroreddeny.toml—
which had removed them, precisely because the reason had expired
(quick-junit0.7 moved to the patchedquick-xml0.41, which the lockfile
is on). So the nightlycargo auditjob was suppressing for no reason, and
would also have silenced any new advisory filed against that line. -
A merged report covers the whole suite again, and its headline agrees
with its page. Two independent failures in--reruncomposition, against
docs/CI.md's promise that "one report stands for the composed result". The
overlay followed only the immediatererun_of, so the ordinary
fix → rerun → fix → rerun loop — the workflow the feature exists for —
silently dropped everything from before the last link, with no banner saying
so; the page just got smaller. It now walks the chain, newest verdict
winning, with a cycle guard becausererun_ofis a string read out of a
record and records travel. And the headline took its numbers from the tail
totals, which belong to the re-run, so one page read2 passed · 0 failed
above a tag table summing to eight and a siblingJUnitsayingtests="8".
The composed stream now declines those totals rather than inventing new
ones, so the headline counts the scenarios actually rendered. -
Run metadata reaches the two ADR-0020 §5 consumers that never received
it. The GitHub job summary — named in the ADR, and the page a CI reader
actually opens from the job — carried none, so the commit under test was in
the record and the HTML report but not there. Anddiff --format json
carriedenvbut notmetadatawhile diff's human output printed
metadata differences, leaving the machine surface a CI gate reads missing
exactly the context the ADR was written to provide. Still handed over, never
harvested. -
The ADR-0002 grammar guard can now see the shapes the ADR names. The
amendment claims the core's hurl vocabulary is closed and pinned; the guard
classified four shapes, and method lines — one of the four boundary
recognisers the amendment's own Measurement section names — was not among
them. Teaching it surfaced one unenumerated token immediately:
GET ${url:base}/PATH, sitting inbind.rsin the same literal as the
already-pinnedHTTP 200. The ADR's table and the pinned set both now carry
it, and the vocabulary is thirteen literals rather than twelve. The scan also
stopped truncating at a file's first#[cfg(test)] modand now excises every
test module: production code placed after one was silently unscanned, and two
core files already carry a second test module. -
--rerunon a truncated record no longer reports success over a suite
that never ran.Record::scenariosis built fromscenario_finished
events alone, so a run killed mid-flight — SIGKILL, OOM, a full disk, a
container eviction — leaves its unreached scenarios absent rather than
recorded. The candidate list built from such a record named nothing, the
"no failures" branch fired, and--rerunexited 0 having executed no
scenario at all.explainsaw the truncation the whole time;--rerundid
not, and CI is exactly where truncation happens. The same class as the
cancelled-run bug fixed in 0.14.0, which this code's own comment describes.A truncated base inverts the question: not "what did the record say to
re-run" but "what can the record prove finished" — everything else in the
selected front runs, announced with a warning naming the truncation. That
distinction now lives in aRerunFilterpredicate rather than a list,
because only the record reader knows which of the two questions applies. -
diffno longer reports a scenario skipped in both runs as "now skipped
… (was passing)". Both halves were false — it did not become skipped, and
it was not passing — and it fired for every@skipscenario on every diff,
including two runs of an unchanged suite, handing--format jsonconsumers
the same wrong pair. The bucket exists for transitions (ADR-0019 §7); the
guard makes that true of the code and not only of its name. -
A tab in a bound value is refused where every other control character
already was. The lower-time guard exempted\t, which hurl'svariable:
grammar rejects like any other control character, so exactly one character
kept taking the late path the guard exists to close — dying as
emit::invalid_artifactagainst generated text the author never wrote,
rather than as a refusal naming their ownbind:. -
--shard-weightsno longer piles every zero-cost scenario into shard 0.
Costs are whole milliseconds, so anything sub-millisecond stores as0—
routine for a fast suite — and adding0never moved a shard's load, so
shard 0 stayed the minimum forever. An all-zero weights file put the entire
suite in one shard and left the others selecting nothing: the flag doing the
exact opposite of its purpose, silently, with the partition still exact so
nothing complained. Assignments are now a tie-break alongside load, which
also gives the right answer when weights genuinely cannot separate
scenarios: equal cost, equal share. -
A disk filling mid-run now reaches the exit code. A stdout that was
already broken at start has failed loudly since the correctness series — but
the human report's own writes go through the console reporter, which
swallows write errors (a reporter cannot report its own channel dying), so a
disk filling during the run truncated the report while the run still
exited by its verdict. TheTeeunder the reporter is the last place the
failure is visible; it now latches the same stdout-failure flagoutln!
uses, and the exit funnel turns lost output into exit 3. Same closed-pipe
exemption as ever —proef … | headis the reader ending the pipeline, not
a failure — and a stderr console (machine mode) does not claim stdout
failed. Pinned by a three-case test, mutation-checked. -
The complexity guard added moments earlier was itself flaky, and now runs
alone. It shipped in the ordinary suite on the reasoning that a ratio
cancels out runner load. Measurement disagreed on its second full-suite run:
2.05× isolated, 3.09× under nextest's full parallelism, against a bound of
3.0. The larger input has the larger working set, so memory-bandwidth
contention penalises it more than the smaller one — the ratio drifts rather
than cancelling, and interleaving the samples cannot fix a systematic effect.nextest's
test-groupsbound concurrency within a group and do not isolate
one from the rest of the suite, so the only mechanism that actually delivers
isolation is#[ignore]plus a dedicated invocation: a CI step of its own and
just perf. The samples are interleaved as well, which removes the one skew
that ordering alone creates.TESTING-STRATEGY.md§7 previously asserted the opposite in as many words —
that a ratio "survives a shared runner" — and is corrected with the numbers.
The claim was reasoning, not measurement, which is the failure this whole
section of the changelog exists to record.
Internal
-
A fixture that spells the record by hand can no longer drift off the
schema.explain's truncated-record test wrote its stream as three JSON
string literals, and all three had drifted: ascenarioscount on the head,
aschemaon the body events, alineon the close.Eventcarries none of
them. Nothing failed and nothing could — the reader has no
deny_unknown_fields, so a stale key parses cleanly and is dropped, and a
fixture built to assert "a record holding one passed scenario" was
three-quarters describing a format proef has never written. It is typed now,
through the helpers its two neighbours already use.The class is closed by a sixth
source_guards.rsrule: every string literal
in the workspace that parses as a JSON object taggedeventmust
deserialize as anEvent, and every key in it must matter — a key is
phantom when deleting it yields the sameEvent. Inertness rather than an
inventory, so it stays correct through renames,#[serde(default)]and
skip_serializing_if, none of which a key-set comparison survives.
Substring assertions against records proef actually emitted
("event":"run_finished","passed":1) are skipped by construction — they are
not objects, and they check the opposite direction. -
Each doc check now lives in the half of the gate that its own rule
names.tests/docs.rsholds the checks that need a built binary (they ask
clap, rather than parsing help text into a model that could drift);
xtask docs-checkholds the ones that read files. The changelog-heading
check added moments earlier read one file and parsed headings, so it sat in
the wrong half — and the cost was concrete rather than tidy: it never ran in
the fast doc-only CI step, only under a full nextest that had to build a
binary it did not use. -
A PR that changes source now has to record itself.
RELEASING.mdhas
always said that every landed change adds an[Unreleased]line in the
commit series that lands it, and nothing checked it — this very entry is the
one that was missed. Measured before being written: across the previous 21
source-touching merges the rule would have fired exactly once, on exactly the
commit that broke it, so the check earns its place by count rather than by
argument.