v4.0.0
kernelcheck 4.0.0 closes the blind spots that the seeded mutant evaluation of 3.0.0 found in the fuzzer, and measures the result on the same mutants. m1 had missed 12 real defects of five kinds, and each kind is fixed in its own commit. Every case is now checked three times, stopping at the first failure: with the NaN canary as before, with every buffer filled with the finite value 2^100 instead, which a fmaxf, a comparison or a value computed from the canary cannot hide, and in a tight placement where each buffer ends on a page boundary with at most three elements of slack, so that a read past a view faults even when its value is thrown away, while every view keeps its alignment and so every kernel its code path. A campaign file can now add extras drawn from a second stream of each case, which leave every case they do not touch exactly as before: tall shapes past the 65535 tiles a grid may have in y, where transpose and gemm loop over their tile rows; a distribution of values from -128 down to just above -2^21, for softmax rows whose maximum lies far below zero; and non-finite rows that are masks of -inf at a high rate, as attention masks are. Both implementations of the input data, the runner's and the Python one, gained the new distribution and the mask and still agree bit for bit, and each new check has a planted defect in the fuzzer's own tests that only it detects.
The rerun m1-v4 was pre-registered before it ran, with the same 171 seeded mutants of the v1.0.0 kernels, the same 2000 finite and 500 non-finite cases for each kernel a mutant changes, and the same shrink budget; the cases come from campaign c2, which is c1 with the extras, and the timeout per case is three times c1's because a case now runs up to three checks. The unmutated kernels passed every one of the 15000 control cases through all three checks. The fuzzer detected 156 of the 171 mutants, against 144 in m1: every one of the 156 real defects, with no mutant lost, and none of the 12 equivalent and 3 contract-equivalent ones, which stay undetected as they must. The tight placement caught the seven discarded reads, the finite fill the two that the canary had hidden, the tall shapes the missing barrier in transpose's tile-row loop, the negative data the softmax maximum that started at zero and the masked rows the online sum that started at one. The unit tests still detect 130, so the fuzzer now catches 26 mutants they miss and they catch none it misses, and by root cause 22 of its detections trigger only when a width is not a multiple of 32. Two evaluations were repeated after the host stalled, a test discovery timeout and four unrelated unit tests killed after 1042 seconds, and their first attempts are kept in the artifact.
The runner also builds with nvcc now. Its device build copies each checked buffer, guards included, into device memory, launches the kernel there and copies everything back, so the same checks judge what the kernel left on the device; the cuda-compile CI job compiles and links it against the CUDA runtime with toolkits 12.9.1 and 13.4.1 and starts it with --version, which touches no GPU. docs/fuzzing.md describes how to run a campaign with it on a machine that has one. It has not been run on a GPU by this repository, and as before every result here comes from the CPU build under the execution model, with no timing figure published. 216 kernel and model tests and 72 fuzzer tests at this commit.
Correction (2026-09-29, v5.0.1): with three mutants argued equivalent reclassified as real defects past the fuzzer's shape limits, m1-v4's fuzzer detected 156 of 159 real defects, not all 156, an in-sample figure since the fixes were written for these mutants' misses; under the rules as pre-registered the unit tests detect 131 and the fuzzer alone 25 (130 and 26 with the repeated evaluation).