Repository navigation
Releases: Hansimov/vflash
Release list
Vflash 0.6.22
2026-10-10 quality review: Temporal quality qualification for the optional Sage path is reopened. Consecutive native-resolution crops and framewise diagnostics of existing 2048-square clips show detail fluctuations in both Flash/Veda and Sage/Veda, with potentially stronger local variation in Sage. The cause has not been isolated; the earlier five-time-point review does not establish flicker-free output. Treat Sage as experimental, not a production quality recommendation. Measured speed and memory observations remain valid within their recorded scope; defaults and released binaries are unchanged.
Keep profiled native blocks on the ordinary path's activation lifetimes: do not retain attention temporaries or pre-gather FFN modulation across the next large allocation. CPU FP32/BF16 parity and weak-reference checks cover both paths; a seven-video RTX 5090 integration completes the previously failing profiled 1536×864/15 s request and 2048²/5 s output. This fixes diagnostic-path memory overhead, without broadening the existing canvas contract.
Add an explicit veda_dense_backend="sageattention2" option for single-SM120 Veda, also exposed by generate --veda-dense-backend sageattention2. The ten dense layers use separately installed, pinned SageAttention 2.2.0 with INT8 QK / FP8 PV; the other forty Veda layers are unchanged. Flash remains the default. Unsupported architectures, missing extensions and kernel failures raise errors instead of silently changing execution. See installation and measured scope.
Same-instance I2VA warm full-generation controls measured 1536×864/15 s at 257.718/223.718 s on RTX PRO 6000 Server (13.19% shorter) and 379.763/310.595 s on RTX 5090 (18.21%). PRO 2048²/5 s was 314.753/302.123 s (4.01% shorter). The 5090 also completed 2048²/5 s at 368.807 s and 26.982 GiB denoising allocation, without a matched Flash control. These times include conditioning, denoising and media but exclude initialization and delivery. The PRO's 1205.634 s initialization prevents claiming a cold-start improvement. Complete AV decode, application playback, five sampled times and native crops accompany the comparisons; changed motion, existing camera adherence defects and unassessed audio semantics remain. No general quality-equivalence or H100 Sage qualification is claimed.
Private integration evidence: video-gen 40745220; only aggregate measurements are public. Product deployment is separate from this package release.
v0.6.21 — owned SM120 hybrid FFN capacity
Reuse the exclusively owned SM120 hybrid FFN base projection for its activated value half, instead of allocating another full output. The gate half and adapter remain unchanged, strict BF16 boundaries are retained, and the next linear consumes a strided view. Other architectures retain the allocated-output path. The owned contract is explicit and rejects autograd inputs.
On a 32 GB, 600 W RTX 5090, a seven-video integration completed 2048²/5 s in 451.039 seconds (26.982 GiB peak denoising allocation), 1536×864/15 s in 364.943 seconds (24.143 GiB), and 1440²/10 s in 392.383 seconds. The earlier 4.28 GiB FFN output allocation failure is avoided. Same-host 1920×1088/5 s A/B/A2 produced identical complete MP4 bytes; its denoising peak stayed 16.960 GiB and no speed gain was established. This is a capacity improvement, not a general speed or image-quality claim. Separate-host 500 W and 600 W runs do not isolate this optimization's latency effect. Full media decode, multi-time visual review and application playback accompany the integration; existing camera motion and unassessed audio semantics remain. Product routing is unchanged.
Private integration evidence: video-gen 3e20f9f3; only aggregate results are public. The promoted kernel also passed four real-GPU shape checks, including subsequent strided linear output without a hidden clone.
Validation: focused clean-source and installed-wheel checks; real RTX 5090 kernel probes; CI and bilingual Pages passed for commit 7c2eb5c7983c1fd64d1ba3e015866de666f6c708.
v0.6.20: SM120 hybrid fusion and FFN activation lifetime
Adds strict BF16 FFN/QKV adapter fusion for the SM120 hybrid four-step profile and releases dead FFN modulation activations earlier.
A same-instance RTX 5090 comparison reduced denoising by about 2.1% and peak denoiser allocation by 2.14 GiB with identical complete MP4 outputs. Earlier activation release saved another 0.99 GiB without a measured speed gain. A previously failing 1536×864 fifteen-second case completed in 431.6 seconds with a 27.92 GiB denoiser peak. A 2048×2048 five-second case still exceeded capacity; these results do not qualify every shape or imply general video quality improvement.
Validation: 46 focused tests from committed source and the installed wheel, privacy checks, bilingual documentation build, and exact-commit CI/Pages. Full videos, multiple frames and real browser playback were inspected. See the bilingual hardware profile and release notes for scope and remaining limits.
v0.6.19 — Owned FP32 media scaling
The FP32 media encoder now obtains its owned work buffer from scaling, eliminating a separate copy while preserving lower-precision buffer reuse.
Eight-CPU full encoding of a decoded FP32 2048-square, 120-frame clip measured 15.145 seconds versus 17.342/17.627 second controls (about 13.4%). Complete MP4 bytes matched and full audio/video decoded. This is CPU encoding evidence; remote full-generation speedup remains unqualified.
No changes to model precision, codec, thread defaults, or hardware qualification. Focused media checks, clean source/wheel builds and bilingual documentation passed.
Vflash 0.6.18 — CPU media quantization
Reuse owned CPU RGB blocks during quantization, reducing temporary allocations while preserving input tensors and media output.
One 2048², 120-frame encoding comparison on eight CPU cores improved 15.7%; complete MP4 bytes were identical. This is encoding-stage evidence, not a GPU or complete-generation throughput claim. No new GPU qualification or thread-count default.
Validation: 17 focused media checks in a clean clone, source and wheel builds, standalone wheel import, privacy/lint hooks and bilingual documentation build.
v0.6.17: Full-memory Blackwell residency
Extend explicit hybrid/Veda trunk residency to single SM120 allocations with at least 90 GiB. A same-PRO-6000-Server A/B/A2 comparison completed nine videos: warm five-second latency improved 18.1% and a matched fifteen-second request 3.3% versus the faster return control, while initialization plus first output increased 18.4 seconds. Peak host RSS fell from 108.7 to 68.7 GiB. Use total batch time when selecting residency; the tested three-output batch did not amortize startup. Default block-ring behavior, I2VA limits, allocator requirement and smaller-card exclusions remain. See scope and evidence.
v0.6.16 — Explicit compact asset lineage
Support explicitly declared exact H3 consumed subsets through portable and prepared assets. Preserve original source identities without relabeling them as derived payload hashes. Ten complete PRO Server cases establish the storage reduction; no startup, host-memory or media-quality gain is claimed. Includes the resident I2VA admission fix from 0.6.15.
v0.6.15 — Resident I2VA admission fix
Fix the opt-in resident request guard to validate the temporal first frame. Real I2VA requests intentionally have no Ref2VA references. Hardware, allocator and capacity restrictions remain unchanged. Regression tests now use the actual VideoRequest contract.
v0.6.14 — H100 full residency
Adds explicit single-H100 Veda full residency for original v0.1 hybrid I2VA. The existing block-ring default is unchanged.
- Requires SM90 with at least 75 GiB visible VRAM and
PYTORCH_ALLOC_CONF=expandable_segments:Truebefore process startup. - Qualified request envelope: one I2VA image, at most 1536 × 864 area and 362 model frames (15 seconds). Other modes and larger canvases remain outside this residency option.
- Same-instance A/B/A2: warm 512² five-second output took 16.99 seconds versus 19.56 seconds after switching back (13.1% faster); initialization plus first output cost about 10.6 additional seconds. About six such requests amortize that measured startup tradeoff.
- Host peak dropped from about 103 GiB to 63 GiB for the small case; the resident long case reached 68 GiB host and 70.3 GiB GPU allocation. This is not a blanket cold-start or visual-quality improvement.
Seven complete outputs were decoded and played through the integration UI. English and Chinese profile guides document scope, evidence and rollback. Clean source, wheel build, focused regressions, full CI and Pages passed.
v0.6.13 — Heterogeneous I2VA preview and CPU preparation
Vflash 0.6.13 adds explicit heterogeneous H3 profiles and CPU preparation while preserving existing SM86/SM89 defaults.
- Explicit SM90, SM103 and SM120 profiles; CUDA-visible device and MIG partition discovery.
veda-tritonfor explicit original-v0.1 hybrid attention; the existingveda-sm89name and default paths remain intact.- CPU-prepared hybrid tables and portable asset views with pinned source provenance, without modifying or hardlinking source weights.
- Complete I2VA hybrid/Veda measurements on H100 SXM and RTX PRO 6000 Blackwell Server/Workstation. New profiles remain previews: other cards, MIG capacity and new-device FL2VA require separate complete-video qualification.
Validation includes a clean source checkout, focused CPU tests, the full CI suite, source and installed-wheel CLI checks, bilingual documentation, and complete GPU video batches. Measurements retain cold-start, host-resource and media-quality limitations; diagnostic success does not establish universal quality improvement. This release contains source and a wheel, no model payloads or new public container image.
中文:新增SM90/SM103/SM120显式配置、MIG可见设备识别、CPU hybrid准备和保留来源身份的资产视图。H100 SXM与PRO Server/Workstation已有完整I2VA实测;其他设备和新架构上的其他模式仍需独立资格。既有SM86/SM89默认行为保持。源码、wheel、完整CI和双语文档通过;冷启动、主机差异、媒体语义限制明确保留,不包含模型权重或新的公开容器镜像。