Chat with Sources with Ollama local model too slow to work #1261
|
I configured open notebook with Ollama local models Gemma4:12b and Qwen3.8 on my Mac M5 Pro 24GB laptop. I tested "Chat with Sources" to find generative AI usage in a 45-page research article. The 'thinking" icon animates to no end for several hours. It happened to both models. Are there any optimization tips to make it work? Thanks. |
Replies: 2 comments
|
A multi-hour spinner is not normal generation latency; first update Open Notebook before tuning the models. Current
After updating, I would test in this order:
The relevant changes and configuration names are recorded in the project's CHANGELOG, including the 8192 Ollama default and per-credential On 24 GB unified memory, start with the smaller of the two models, 8k context, insights-only context, and no other loaded Ollama model. Increase one dimension at a time after a successful response. |
|
@Michael-WhiteCapData thank you, this is the answer I'd have written. I checked each point against the changelog and it holds: the 8192 default and per-credential @chris-opendata one more cause that fits "spins for hours" specifically: #1327. Source chat's stream sent nothing while a slow model was generating, so the connection went idle and got dropped. The answer was actually produced, but only showed up after leaving and reopening the source. If that's what you see (the reply appears after navigating away), it's that bug, and #1332 fixes it. Combined with Michael's checklist (8k context, insights before full content, a smaller non-thinking model first), that should get you to a working baseline. I'm closing this as answered; reopen it with the API log from one request if it still hangs. Status: answer. Tuning per the checklist above; idle-connection drop tracked in #1327 / #1332. |
A multi-hour spinner is not normal generation latency; first update Open Notebook before tuning the models. Current
maincontains several fixes directly related to this case:num_ctxwas reduced from 128k to 8192 to avoid consumer-GPU OOM/slowdown.num_ctxcan now be set per Ollama credential.After updating, I would test in this order:
num_ctxto 8192 initially.ollama ps. If it re…