v3.4.5 — Production-Ready Test Density Scoring
What this release fixes
v3.4.4 introduced test density scoring but had a ship-blocking bug: the grep for counting test cases scanned node_modules, dist, and build artifacts. On a real React project (forge-ui), this inflated 4,187 real matches to 14,130 phantom matches — a 237% over-count.
Bug fixes
- P0:
node_modulesinflation — Test case grep now usesfindwith the sameBUILD_ARTIFACT_EXCLUDESas source file counting. No more phantom matches from dependency packages. - Double-counting — v3.4.4 used two grep passes that counted
*.test.*files insidetests/twice. Replaced with singlefind | xargs greppipeline. - Missing patterns — Now detects:
it.each(),test.skip(),test.only(),it.concurrent()(data-driven/skipped tests)__tests__/directory convention (Jest without.test.in filename)*.cy.*files (Cypress E2E tests)#[tokio::test],#[rstest],#[actix_web::test](async Rust)
Validated against 3 real projects
| Project | Stack | Tests | Source files | Density | Score |
|---|---|---|---|---|---|
| forge-ui | React 19 | 4,187 | 851 | 4.9/file | 15/20 (solid) |
| forge-orchestrator | Rust | 362 | 43 | 8.4/file | 20/20 (thorough) |
| game.clicker | React + Vite | 37 | 3 | 12.3/file | 20/20 (thorough) |
Known limitations (documented, not bugs)
- Distribution blindness — 100 tests in 1 file + 0 tests in 99 files averages to 1.0 density. Real line coverage (Tier 1) catches this; density (Tier 2) does not.
- Parameterized tests —
test.each()with 10 rows counts as 1 test case. Go table-driven tests witht.Run()subtests are also undercounted. - BDD/Cucumber —
.featurefiles withScenario:are not detected.
These limitations are inherent to any fast proxy. The scoring clearly recommends running --coverage for Tier 1 accuracy.
Upgrade
/forge:updateThen restart Claude Code.
Full changelog: v3.4.4...v3.4.5