Lumen Studio v0.1.3 - GPU backend switching fix
·
12 commits
to main
since this release
GPU backend switching fix
This patch fixes GPU inference and fast response streaming when changing backends.
Fixed
- Vulkan and CUDA responses no longer appear frozen or stop printing when chunks arrive quickly.
- Changing CPU, Vulkan, CUDA, or Auto now safely stops any active generation and automatically reloads the same model.
- The loaded-model panel and confirmation message now report the backend actually in use.
- Failed backend switches restore the previous working backend instead of leaving stale UI state.
- Chat completion and error events stay attached to the correct conversation.
Verified
- Automated CPU -> Vulkan -> CPU switching test passed.
- CPU mode used 0 GPU layers.
- Vulkan mode offloaded all 23 layers of the test model and streamed 28 chunks successfully.
- Windows package includes CPU, Vulkan, and CUDA x64 runtimes.
The Windows installer is currently unsigned, so Microsoft Defender SmartScreen may display a warning. Download it only from this GitHub release or the official Lumen Studio website and verify its SHA-256.
SHA-256: dd16bb840c5b11e0bc4caf59a3e4faeb0928cf638588991f28d00e52ba83f4a9