Skip to content

Easy ASR Bench v0.3.9

Pre-release
Pre-release

Choose a tag to compare

@rollingedit rollingedit released this 09 Jun 04:46
· 225 commits to main since this release

Easy ASR Bench v0.3.9

Release status: preview / prerelease. This is a much larger validation and product-path release than v0.3.8, with real model/runtime smoke coverage, expanded bootstrap repair behavior, a redesigned batch report, and explicit blocked evidence for the rows this machine cannot prove. It is not promoted as an all-pass stable release because clean Windows Sandbox/VM, NVIDIA CUDA, and Intel provider rows still require external hardware/OS evidence.

What changed

  • Added real end-to-end ASR/LLM validation paths around the normal product flow instead of only mocked unit coverage. The v0.3.9 smoke artifact now carries 110 pass / 12 blocked / 0 not_run manual rows, with blocked rows recorded as explicit external requirements.
  • Added full real-smoke doctor support for the major runnable paths: faster-whisper/CTranslate2, OpenAI Whisper .pt, whisper.cpp GGML, Hugging Face Safetensors ASR, Generic ONNX CTC, ASR GGUF+mmproj, SmolLM GGUF reference/correction, and same-media multi-model benchmark rows.
  • Added real public folder-batch validation with multiple speech files, multiple ASR backends per file, SmolLM 135M GGUF correction/scoring, and generated final batch reports that show corrected references, transcripts, ranking, timing, RAM, and VRAM/GPU memory.
  • Redesigned the batch report into the main user-facing result view: final_results.html opens to an overall model ranking, simple file picker, per-file winners, full model transcripts, corrected-reference editor, direct two-model comparison, and advanced raw paths hidden behind details.
  • Added live corrected-reference editing in the report. Pasted ChatGPT/Claude text rescoring/highlights immediately, shows a visible update confirmation, and can be reset per file.
  • Improved transcript diff readability: true word changes are red, punctuation/case differences are yellow and separate from the winner score, missing reference words are distinct, and very high WER values get a compact non-clipping warning.
  • Hardened setup/bootstrap repair behavior for missing, stale, corrupt, incompatible, and provider-specific dependency states, including Python packaging tools, Visual C++ runtime, faster-whisper/CTranslate2 native load failures, ONNX provider conflicts, Transformers/Torch dependency groups, and llama.cpp/MTMD native runtime discovery.
  • Fixed the v0.3.8 audit blocker in standalone one-file setup: setup.bat now downloads install.ps1 through a precomputed temp path instead of using stale %INSTALLER_PS1% expansion inside a batch block.
  • Added lonely one-file setup coverage to public asset smoke and GitHub Actions, so release verification runs from a folder containing only setup.bat, matching the normal-user download path.
  • Fixed Run.bat --doctor --strict parity by passing --strict through app.main into app.doctor.
  • Expanded Hugging Face downloader/package handling for sequential link entry, invalid-link retry, sharded Safetensors indexes, ONNX sidecars, split GGUF parts, GGUF ASR mmproj pairing, and quant/precision classification.
  • Added Windows Sandbox clean-bootstrap deployment scaffolding. On supported Windows Pro/Enterprise/Education hosts it generates a .wsb plus startup script that runs setup repair, model-layout repair, clean VM bootstrap proof, full real-smoke doctor, release-smoke generation, evidence merge, and validation.
  • Updated GitHub Actions to front-load setup/workflow contract tests before the full pytest run, reducing expensive failures from YAML/script syntax or bootstrap regression.

Automated Packaging Checks

  • Built from commit: dc5c895a349f924693c1a50eec8b057a7abab3ed.
  • Local release file validation passed.
  • Local physical source validation passed.
  • Local ZIP build/layout validation passed for Easy-ASR-Bench-v0.3.9-win.zip.
  • Staged setup verification passed with setup.bat --dry-run --verify-release --asset-dir dist\release-assets.
  • Release smoke validation passed with required rows, log hashes, and environment summaries.

Manual Smoke Rows Marked Pass

  • 110 manual release rows are marked pass in release-smoke-v0.3.9.json.
  • Coverage includes setup/doctor repair contracts, dependency repair isolation, model-layout repair, media handling, OpenAI Whisper .pt safety, whisper.cpp, faster-whisper/CTranslate2 repair, HF Safetensors, Generic ONNX CTC, ASR GGUF+mmproj, SmolLM correction/scoring, same-media CPU/DirectML benchmark rows, report/reference validation, and batch report UX proof.

Still Blocked / Pending External Evidence

  • win11_clean_no_python_setup: requires a clean Windows 11 VM/Sandbox where Python is not visible before setup.
  • windows_sandbox_clean_bootstrap_deploy: this laptop is Windows Home/Core and does not expose Windows Sandbox.
  • clean_vm_zero_dependency_bootstrap: requires explicit clean VM/Sandbox proof with EASY_ASR_BENCH_CLEAN_VM_BOOTSTRAP_PROOF=1.
  • win10_existing_python_setup: requires a Windows 10 VM state.
  • NVIDIA CUDA rows remain blocked on this machine: hardware detection, Torch CUDA tensor smoke, ONNX Runtime CUDA, faster-whisper/CTranslate2 CUDA, llama.cpp CUDA/SmolLM, and the combined CUDA row.
  • Intel provider rows remain blocked on this machine: Intel DirectML hardware and OpenVINO provider smoke.

Release assets

  • Easy-ASR-Bench-v0.3.9-win.zip: sha256:4f76a8c8c56eed16e27ed27bf8309f42b8c12acb6d490bf68dff4a8c9b2dca97
  • setup.bat: sha256:ff28e044d4212f54186e705b06020c34df43cc6bd2f70047489d2ca3ed55886a
  • install.ps1: sha256:21701f93b0b662962a5dae3df3f4f134de4e0b3fc52d2b2c82cda0ae4b496aa4
  • manifest.json: sha256:02a9ab2a9b58df9651b52cd1c8005b3d0bae6bfe3a0e246331b891e7be607e1d

Known limits

  • This is still a preview/prerelease until the external clean Windows Sandbox/VM and NVIDIA CUDA evidence is merged and strict all-pass validation succeeds.
  • CUDA package installation remains conservative; CPU fallback is explicit when CUDA is requested but not verified.
  • VRAM/GPU memory telemetry is reported separately from RAM and depends on Windows GPU counters or backend-specific metrics.
  • Unsafe pickle-backed .pt checkpoints remain blocked unless explicitly trusted through the supported safety path.