v1.4.0 - Trustworthy verdicts
Trustworthy verdicts
Five separate ways the tool could report a wrong number without saying so are fixed. A failing suite no longer makes every mutant look killed, cached verdicts no longer outlive the code they were measured against, --no-cache no longer leaves a cache behind, and cleanup from a timed-out mutant no longer kills the next one. SPM packages now test their mutants in parallel.
What's new
Baseline validation (SPM)
- The unmutated suite runs once before any mutant, and the run ends unless it passes
- Anything else throws
BaselineError, naming the tests that failed, the timeout that stopped the suite, or the output it died with - The failing-tests message points at the sandbox under
$TMPDIR, since a test deriving paths from#filePathpasses in place and fails there - The baseline runs with an explicitly empty mutant selection, so a stray
__SWIFT_MUTATION_TESTING_ACTIVEin the environment cannot select one for the run that must have none
--keep-logs <directory>
- Writes what each mutant's tests printed to
<directory>/<mutant-id>.log - Each log opens with the mutant, its location, the mutation, the verdict and the duration, so it reads on its own without the JSON report
- Unviable mutants get a log too, holding the output of the build that failed — the verdict that explains itself least, and on some packages the majority of a run
- Failing to write a log never fails a run
Parallel execution for SPM packages
- Mutants are tested by running the compiled
.xctestbundle directly rather than throughswift test, which takes a lock on.buildand forced workers to queue --concurrencyis honoured for SPM again- The build still happens exactly once
- Measured on a real package: 10 concurrent test processes at 499% CPU, against 196% for the same tool run sequentially
Bug fixes
Verdicts
- A failing suite no longer makes every viable mutant report
Killed.validateSPMBaselineran the unmutated suite and discarded the result; one reported run showed 260 killed and 0 survived, with the true survivors invisible (#66) - Cleanup after a timed-out mutant no longer kills the next mutant's test process. The escalation searched every process on the machine for one whose arguments mentioned the sandbox, which cannot tell one mutant's run from another's when both share a sandbox — the truncated output was read as a crash (#69)
- The SIGKILL that follows a timeout is now tied to the run's lifetime: a process that stops when asked has its descendants cleaned up at once, rather than on a timer that outlives it
Cache
- Editing the code under test now invalidates the verdicts measured against it.
fileContentHashfell back to the file path for every schematizable mutant — a constant — so stale verdicts were replayed (#79) - A
Killedverdict is now dropped when the test that killed it changes. The killer file was stored absolute and compared against project-relative keys, so the comparison was never true (#67) --no-cacheno longer writes a cache. It guarded only the read side, so a run told not to use the cache still left one behind for the next run to replay (#68)- Crash verdicts are re-measured instead of being kept forever. They were grouped with
.unviableand skipped by every diff, so a spurious one could never be cleared (#80)
Configuration
--concurrencyis resolved to what a run can actually use. Xcode schemes targetingplatform=macOSget one worker, as do XCTest runs, rather than reporting a number the simulator pool will not honour (#70)- The readiness line no longer claims simulators that do not exist: a run with
--concurrency 9against an SPM package printed✓ 9 simulators ready
Architecture changes
| New type | Responsibility |
|---|---|
BaselineError |
Why the unmutated suite cannot serve as a baseline, in the three shapes the failure takes |
ProjectRelativePath |
The one way a path is written down when it has to be compared with another |
ProcessTree |
Descendants of a process, read from the kernel's process table by parent pid |
TimeoutEscalation |
The SIGKILL after a timeout, tied to the run's lifetime rather than a timer |
TestBundleInvocation |
How to run a package's compiled test bundle without going through swift test |
MutantLogWriter |
One log per mutant, when --keep-logs is given |
MutantDescriptor gains sourceContentHash — the hash of the unmutated file, computed at discovery where the contents are already in hand rather than read back once per mutant. The cache key is built from it, alongside the file path: content alone collides for two byte-identical files, whose mutants are not interchangeable because they compile into different places.
RunnerEvent.simulatorPoolReady(size:) becomes workersReady(count:usesSimulators:) — the reporter can no longer call a worker a simulator when there is no simulator.
SimulatorPool hands out every configured slot when it has no simulators to clone, instead of a single one.
Both testing libraries are run for SPM mutants, in the order swift test runs them and under one shared deadline. A package may hold XCTest classes and Swift Testing functions at once, and running only the configured one silently skips the other's tests.
Upgrade notes
Existing caches are invalidated. The cache key changes shape, so the first run after upgrading re-measures everything. That is the point: the old keys were built on something that could not detect the change they existed to detect.
A failing suite now ends the run. Previously the run continued and produced a report. If your suite does not pass inside the sandbox — most often because a test derives paths from #filePath, which resolves to $TMPDIR there — the run will stop and name the tests to fix.
--concurrency behaves differently. SPM packages honour it, where before it was accepted and ignored. Xcode schemes on platform=macOS, and XCTest runs, resolve it to 1 rather than reporting a figure the pool will not honour.
Validated against a real project
Run against a 116-file package with 944 tests and compared mutant-by-mutant against a report from v1.3.0.
Discovery is unchanged: the same 949 mutants at the same positions, none exclusive to either version. 50 verdicts (5.3%) differ.
| status | v1.3.0 | v1.4.0 |
|---|---|---|
| Killed | 238 | 254 |
| Survived | 10 | 16 |
| Crash | 99 | 98 |
| Timeout | 15 | 10 |
| Unviable | 587 | 571 |
Scoped to one directory and run at matching concurrency, 32 of 63 verdicts changed — 20 Unviable → Killed and 2 Unviable → Survived. Those last two matter most: real surviving mutants that were hidden behind "not testable", which is a test gap the tool was not reporting.
Four Crash → Survived are false kills from the cleanup defect, now gone.
Test coverage
- 692 tests in 80 suites
- 4693 of 4693 source lines covered
Known issues
- Swift Testing failures are reported as
Crashrather thanKilled, because the parser looks for a line shape the library does not emit. The score is unaffected, but the killer test file is not resolved (#83) SwapTernaryproduces mutants that cannot compile, so none is ever testable (#82)- Builds are bounded by the test timeout, so a slow build under parallel load can be recorded as
Unviable(#84) - A high proportion of mutants are reported
Unviableand the cause is not yet diagnosed (#85)
Thanks
To @jwp23, who reported #66, #67, #68, #69 and #70 — the five defects this release is built around. Each came with measurements, a list of what had been ruled out, and in #66 a negative control: an assertion removed on purpose so a known true survivor could be shown still coming back Killed. That is what made them actionable without having to reproduce them first.
Requirements
- macOS 15+
- Swift 6.2+
- Xcode project with a valid scheme and test target, or an SPM package with a test target
Installation
See the Installation Guide for Homebrew, pre-built binary, and build from source instructions.