Skip to content

vhs-decode-dotnet-v0.4.0-1.4.3

Choose a tag to compare

@JunliangRen JunliangRen released this 31 Jul 15:33
Immutable release. Only release title and notes can be modified.
f35fc47

Highlights

  • Specializes the Exact PocketFFT radix-8 direction outside the hot loop while preserving the pinned expression and operation order.
  • Adds AVX2/SSE4.1 VHS chroma UInt16 conversion with the original finite, saturation, truncation, and scalar-tail behavior.
  • Adds the verified 16-lane AVX/FMA current chroma burst fit and deterministic 16+16 radix narrowing for current sync quantiles, with scalar and IEEE-special-value fallbacks.
  • Keeps worker scratch local and bounded; no cross-field state, output ordering, data type, or compatibility-visible numerical sequence moves.
  • Refreshes the English, Chinese, and Japanese Python/.NET performance matrix and raises the standard xUnit v3 suite and CI floor to 1,169 tests.

On the same private local RF sample and bounded range, six balanced Exact current/20-worker pairs reduced median wall time from 9.637 to 8.627 seconds (10.48% lower, 11.71% more throughput). Four single-worker pairs reduced the median from 50.377 to 37.300 seconds (25.96% lower, 35.06% more throughput). Across two opposite-order 1,000-frame pairs, mean wall time fell from 72.405 to 62.752 seconds (13.33% lower, 15.38% more throughput). Mean integrated allocation changed by only +0.18%, candidate working set stayed below 706 MiB, and no progressive slowdown was observed.

Validation

  • PR #90 was merged through the protected main workflow.
  • An independent read-only local sub-agent reviewed all 19 changed files and found no actionable issue before merge.
  • Protected-main Windows Release workflow passed at f35fc471d0ddebbf947438d24a2366317262f69b.
  • Local Release build completed with 0 warnings and 0 errors.
  • Standard xUnit v3 suite: 1,169 passed with native hardware, AVX2 disabled, and all hardware intrinsics disabled.
  • Intel IPP native build and artifact verification passed in GitHub Actions.
  • Native and scalar Exact v0.4.0/current gates matched luma, chroma, raw JSON, stdout, normalized stderr/logs, and ordered fileLoc at serial, default-five, and 20-worker settings.
  • Both opposite-order 1,000-frame runs completed 2,000 fields with exact artifacts, diagnostics, and ordered fileLoc.
  • All 60 final Exact/IPP-fast, profile, and worker-count matrix runs matched their profile/backend references.
  • The packaged self-contained executable passed the vhs --help startup smoke test.

Python v0.4.0 g4315520 --threads 0 remains the strict oracle for the v0.4.0 profile. Merged Python PR341 remains the profile peer for current.

Package

Windows x64 binary-only package containing a single self-contained decode.exe.

  • decode.exe SHA-256: AC4554F7B2864B87126F226EDE95A368F0927B7725B285AD74B2A59690A08221
  • ZIP SHA-256: F7F7B2499185DB15448103E78594D8B7092C460D7CA670EBB147167D9B0FB851