Releases: jaydemks/phrasync
Release list
Phrasync v0.4.4
Phrasync 0.4.4 fixes manually added lyric cues disappearing when they overlap transcribed lines. The fix covers local preview and MP4 export, including older projects with untimed manual cues. New manual cues also respect the playhead timing offset, and edits refresh the preview immediately. The full automated test suite passes.
Install the Windows app for free from the Microsoft Store. This GitHub release provides the source code, not a separate Windows installer. On Windows, set up Python 3.11 or 3.12 and run run_windows.bat; the README also explains how to open the same local editor in your browser.
The 0.4.4 Store package, with updated release notes and a demo trailer, has been submitted and is currently in certification. Until Microsoft approves and distributes it, the Store may still offer 0.4.3. Going forward, Phrasync updates will be maintained for both the GitHub source and the Microsoft Store app, with Store availability following its certification process.
See README.md and RELEASE_NOTES.md for details.
Phrasync v0.4.3
This source release brings GitHub up to date with Phrasync 0.4.3. Since the first release: a native Windows window, English/Italian UI, visible model-download progress, NVIDIA GPU support where available, longer subtitle projects, richer 3D scenes including Lost in the Ocean, better lyric timing, and improved MP4 export with a chosen destination and Japanese font fallback. Version 0.4.3 also fixes stale interface assets after an in-place Store update and adds in-app Store update checks and a diagnostics panel. The Microsoft Store build is still in certification; no installer is attached to this GitHub release. Existing Store users should back up important projects, uninstall an earlier version once, then install 0.4.3 from the Store when available. See RELEASE_NOTES.md for details and the remaining update-flow verification limit.
Phrasync v0.1.0
Phrasync v0.1.0
This is the first public release: useful today, intentionally honest about what still needs engineering.
Highlights
- Two first-class project modes: Lyric Video for music and Subtitles for spoken audio or video.
- A video selected in Subtitles mode becomes the transcription source, preview footage, render background, and final audio source without duplicate setup.
- Hierarchical local Faster-Whisper transcription: ten-second language mapping, confidence and silence filtering, stable span smoothing, quiet-boundary snapping, language-locked span decoding, phrase-level language verification, selective corrective re-reading, adaptive fallback, and persisted gauntlet diagnostics.
- Auto follows language changes across a track; one code locks a language, comma-separated codes restrict the candidate set, and
singlekeeps one automatically detected language. - Language metadata is preserved on cues and returned as a dominant language, ordered language list, editable API-level span map, confidence, and diagnostics.
- Seven kinetic 2D typography presets with browser/Python timing parity.
- Seven separately named and tuned 3D typography personalities with independent world-space X/Y offsets.
- Grounded world-space 3D text: phrases are planted on the scene floor and the camera travels past them like other meshes.
- Odyssey 3D environments for Japan, Italy, China, and the USA.
- Manual rain, snow, fog, storms, falling leaves, dawn, day, sunset, night, and four seasons.
- Automatic journey mode crossfades daytime and seasonal colour while selecting deterministic weather from the playback clock.
- Image, video, and dynamic backgrounds; local project save/load; critic preflight; cancellable render jobs; H.264 MP4 download and postflight decoding.
Verified for this release
- 50 automated tests pass in the current release environment.
- A real short MP4 render passes with one video file used for both picture and audio in Subtitles mode.
- Chrome smoke coverage passes for every 2D/3D text composition state over Odyssey 3D, a 2D dynamic visual, image, and video modes; actual MP4 decoding is covered by the integration render above.
- 3D text remains world-planted and completes its fly-past at scene speeds of 20%, 100%, and 260%.
- Weather, daytime, seasons, automatic state changes, project-mode copy, and the 3D preset catalogue pass in the live browser with zero captured exceptions.
- Repeated 244.9-second
large-v3Auto runs map Italian, Portuguese, French, Spanish, English, and Japanese; recover the opening lyrics; preserve Japanese writing; follow late line-level switches; and produce no false outro after the final vocal. - Development and physical testing were performed on Windows 11, Ryzen 9 3900X, 32 GB RAM, and RTX 3090 24 GB. Linux and macOS launchers are included but not physically release-tested yet.
Known limits
- Odyssey Flat, Odyssey 3D, and 3D typography are preview-only in v0.1.0. MP4 preflight blocks them explicitly because the Python renderer does not yet reproduce WebGL output faithfully.
- Full song-length output has not yet been rendered. Current verification uses short engineering renders; complete music videos will be added after publication.
- Projects reference media in the local Phrasync workspace rather than embedding it into a portable archive.
- Browser video preview is intentionally limited to MP4 and WebM; convert MOV, MKV, AVI, and unusual codecs before import.
- The transcription workflow performs strongly on the tested songs, but expressive singing, unusual pronunciation, effects, and closely related languages can still need human review.
- There is no signed installer, automatic updater, or hosted web version.
Evaluated next-step technology
- Troika Three Text for SDF-quality, multilingual 3D typography.
- JASSUB/libass for professional ASS subtitle parity in browser preview.
- Three.js Sky for physically based atmosphere when deterministic export parity is ready.
- three.quarks only if future VFX outgrow the lightweight deterministic weather system.
Phrasync is created by jaydemks and distributed under Apache-2.0 with attribution preserved through NOTICE. Output belongs to its creator, subject to the rights they hold in the source music, text, footage, images, and fonts.