Skip to content

v1.4.0 - Trustworthy verdicts

Choose a tag to compare

@ericodx ericodx released this 23 Sep 12:10
· 222 commits to main since this release
2faea74

Trustworthy verdicts

Five separate ways the tool could report a wrong number without saying so are fixed. A failing suite no longer makes every mutant look killed, cached verdicts no longer outlive the code they were measured against, --no-cache no longer leaves a cache behind, and cleanup from a timed-out mutant no longer kills the next one. SPM packages now test their mutants in parallel.


What's new

Baseline validation (SPM)

  • The unmutated suite runs once before any mutant, and the run ends unless it passes
  • Anything else throws BaselineError, naming the tests that failed, the timeout that stopped the suite, or the output it died with
  • The failing-tests message points at the sandbox under $TMPDIR, since a test deriving paths from #filePath passes in place and fails there
  • The baseline runs with an explicitly empty mutant selection, so a stray __SWIFT_MUTATION_TESTING_ACTIVE in the environment cannot select one for the run that must have none

--keep-logs <directory>

  • Writes what each mutant's tests printed to <directory>/<mutant-id>.log
  • Each log opens with the mutant, its location, the mutation, the verdict and the duration, so it reads on its own without the JSON report
  • Unviable mutants get a log too, holding the output of the build that failed — the verdict that explains itself least, and on some packages the majority of a run
  • Failing to write a log never fails a run

Parallel execution for SPM packages

  • Mutants are tested by running the compiled .xctest bundle directly rather than through swift test, which takes a lock on .build and forced workers to queue
  • --concurrency is honoured for SPM again
  • The build still happens exactly once
  • Measured on a real package: 10 concurrent test processes at 499% CPU, against 196% for the same tool run sequentially

Bug fixes

Verdicts

  • A failing suite no longer makes every viable mutant report Killed. validateSPMBaseline ran the unmutated suite and discarded the result; one reported run showed 260 killed and 0 survived, with the true survivors invisible (#66)
  • Cleanup after a timed-out mutant no longer kills the next mutant's test process. The escalation searched every process on the machine for one whose arguments mentioned the sandbox, which cannot tell one mutant's run from another's when both share a sandbox — the truncated output was read as a crash (#69)
  • The SIGKILL that follows a timeout is now tied to the run's lifetime: a process that stops when asked has its descendants cleaned up at once, rather than on a timer that outlives it

Cache

  • Editing the code under test now invalidates the verdicts measured against it. fileContentHash fell back to the file path for every schematizable mutant — a constant — so stale verdicts were replayed (#79)
  • A Killed verdict is now dropped when the test that killed it changes. The killer file was stored absolute and compared against project-relative keys, so the comparison was never true (#67)
  • --no-cache no longer writes a cache. It guarded only the read side, so a run told not to use the cache still left one behind for the next run to replay (#68)
  • Crash verdicts are re-measured instead of being kept forever. They were grouped with .unviable and skipped by every diff, so a spurious one could never be cleared (#80)

Configuration

  • --concurrency is resolved to what a run can actually use. Xcode schemes targeting platform=macOS get one worker, as do XCTest runs, rather than reporting a number the simulator pool will not honour (#70)
  • The readiness line no longer claims simulators that do not exist: a run with --concurrency 9 against an SPM package printed ✓ 9 simulators ready

Architecture changes

New type Responsibility
BaselineError Why the unmutated suite cannot serve as a baseline, in the three shapes the failure takes
ProjectRelativePath The one way a path is written down when it has to be compared with another
ProcessTree Descendants of a process, read from the kernel's process table by parent pid
TimeoutEscalation The SIGKILL after a timeout, tied to the run's lifetime rather than a timer
TestBundleInvocation How to run a package's compiled test bundle without going through swift test
MutantLogWriter One log per mutant, when --keep-logs is given

MutantDescriptor gains sourceContentHash — the hash of the unmutated file, computed at discovery where the contents are already in hand rather than read back once per mutant. The cache key is built from it, alongside the file path: content alone collides for two byte-identical files, whose mutants are not interchangeable because they compile into different places.

RunnerEvent.simulatorPoolReady(size:) becomes workersReady(count:usesSimulators:) — the reporter can no longer call a worker a simulator when there is no simulator.

SimulatorPool hands out every configured slot when it has no simulators to clone, instead of a single one.

Both testing libraries are run for SPM mutants, in the order swift test runs them and under one shared deadline. A package may hold XCTest classes and Swift Testing functions at once, and running only the configured one silently skips the other's tests.


Upgrade notes

Existing caches are invalidated. The cache key changes shape, so the first run after upgrading re-measures everything. That is the point: the old keys were built on something that could not detect the change they existed to detect.

A failing suite now ends the run. Previously the run continued and produced a report. If your suite does not pass inside the sandbox — most often because a test derives paths from #filePath, which resolves to $TMPDIR there — the run will stop and name the tests to fix.

--concurrency behaves differently. SPM packages honour it, where before it was accepted and ignored. Xcode schemes on platform=macOS, and XCTest runs, resolve it to 1 rather than reporting a figure the pool will not honour.


Validated against a real project

Run against a 116-file package with 944 tests and compared mutant-by-mutant against a report from v1.3.0.

Discovery is unchanged: the same 949 mutants at the same positions, none exclusive to either version. 50 verdicts (5.3%) differ.

status v1.3.0 v1.4.0
Killed 238 254
Survived 10 16
Crash 99 98
Timeout 15 10
Unviable 587 571

Scoped to one directory and run at matching concurrency, 32 of 63 verdicts changed — 20 Unviable → Killed and 2 Unviable → Survived. Those last two matter most: real surviving mutants that were hidden behind "not testable", which is a test gap the tool was not reporting.

Four Crash → Survived are false kills from the cleanup defect, now gone.


Test coverage

  • 692 tests in 80 suites
  • 4693 of 4693 source lines covered

Known issues

  • Swift Testing failures are reported as Crash rather than Killed, because the parser looks for a line shape the library does not emit. The score is unaffected, but the killer test file is not resolved (#83)
  • SwapTernary produces mutants that cannot compile, so none is ever testable (#82)
  • Builds are bounded by the test timeout, so a slow build under parallel load can be recorded as Unviable (#84)
  • A high proportion of mutants are reported Unviable and the cause is not yet diagnosed (#85)

Thanks

To @jwp23, who reported #66, #67, #68, #69 and #70 — the five defects this release is built around. Each came with measurements, a list of what had been ruled out, and in #66 a negative control: an assertion removed on purpose so a known true survivor could be shown still coming back Killed. That is what made them actionable without having to reproduce them first.


Requirements

  • macOS 15+
  • Swift 6.2+
  • Xcode project with a valid scheme and test target, or an SPM package with a test target

Installation

See the Installation Guide for Homebrew, pre-built binary, and build from source instructions.