Repository navigation
Vflash 0.3.2 — wide indices and dual-SM86 T2 native
Vflash 0.3.2 fixes signed 32-bit element-offset overflow in large fused H3 tensors and adds the fixed t2va-turbo4-exact-sm86 native profile for two RTX 3080 20 GB GPUs using sequence-head.
- Fifteen small/wide GPU operator checks passed bit-for-bit. Saved 243-frame conditioning completed eight Ref evaluations with finite FP32 audio/video latents, 240-frame 1344 × 768 media, stereo 32 kHz audio and confirmed cleanup. This is not a full-trajectory bitwise comparison.
- Dual-SM86 T2 compilation passed 52 official same-architecture timestep/modulation checks. Application-owned encoding/core/media completed five-second 928 × 512 and ten-second 640 × 352 requests. Thermal throttling limits performance conclusions.
- H3Pipeline remains five seconds at 24 fps. Its new dual-SM86 T2 wrapper was not rerun on GPU; these native integration checks do not add a ten-second wrapper API or establish general semantic/audio quality.
The wheel and two images contain the runtime built at 7ddfeee22e005f8172751cf6dc044e43ca5e758b; the release tag adds documentation and image identities without changing runtime source. Dependencies remain pinned to the 0.3.1 image layers. Model weights are separate licensed inputs.
English release notes · 中文说明 · Immutable image inventory
Images: hansimov/vflash:0.3.2 and hansimov/vflash:0.3.2-pipeline. Attachments include the exact wheel, the small-layer Docker recipe and SHA256 checksums.