Local Idea Studio v0.1.6 - Performance and visible reasoning
0825-Local-Idea-Studio-YouTube.mp4
Performance and visible reasoning
- Auto mode now uses every logical CPU thread available to the app.
- Accelerated backends use a 2048 prompt batch and request full GPU layer offload when memory permits.
- Runtime info reports actual CPU threads, batch size, backend, GPU layers, and offload percentage.
- Thinking-capable GGUF chat templates stream their real model-generated reasoning in a collapsible live panel.
- Standard models show only their answer and never display a fake empty Thinking panel.
- Backend switching tests cover CPU -> Vulkan -> CPU without stale state.
Verified on the development system
- Ryzen 5 8600G: 12/12 logical CPU threads available.
- Radeon 760M Vulkan: all 29/29 Qwen3 1.7B Q2_K model layers offloaded.
- Measured generation: about 46.8 tokens/second with the tested model and settings.
- Qwen reasoning output and TinyLlama non-reasoning behavior both passed.
Performance varies with model, context, quantization, memory bandwidth, drivers, temperatures, and hardware. A constant 100% on both CPU and GPU is not guaranteed or necessarily optimal.
The Windows installer is currently unsigned. Verify its SHA-256 before running it:
235f81dd38efd91600175e2a4b28d7f0a1d3ae40c5f64ac848a6778a44233f45