Skip to content

Local Idea Studio v0.1.6 - Performance and visible reasoning

Choose a tag to compare

@workyulian-ship-it workyulian-ship-it released this 25 Aug 04:23
· 8 commits to main since this release
0825-Local-Idea-Studio-YouTube.mp4

Performance and visible reasoning

  • Auto mode now uses every logical CPU thread available to the app.
  • Accelerated backends use a 2048 prompt batch and request full GPU layer offload when memory permits.
  • Runtime info reports actual CPU threads, batch size, backend, GPU layers, and offload percentage.
  • Thinking-capable GGUF chat templates stream their real model-generated reasoning in a collapsible live panel.
  • Standard models show only their answer and never display a fake empty Thinking panel.
  • Backend switching tests cover CPU -> Vulkan -> CPU without stale state.

Verified on the development system

  • Ryzen 5 8600G: 12/12 logical CPU threads available.
  • Radeon 760M Vulkan: all 29/29 Qwen3 1.7B Q2_K model layers offloaded.
  • Measured generation: about 46.8 tokens/second with the tested model and settings.
  • Qwen reasoning output and TinyLlama non-reasoning behavior both passed.

Performance varies with model, context, quantization, memory bandwidth, drivers, temperatures, and hardware. A constant 100% on both CPU and GPU is not guaranteed or necessarily optimal.

The Windows installer is currently unsigned. Verify its SHA-256 before running it:

235f81dd38efd91600175e2a4b28d7f0a1d3ae40c5f64ac848a6778a44233f45