Skip to content
Discussion options

You must be logged in to vote

A multi-hour spinner is not normal generation latency; first update Open Notebook before tuning the models. Current main contains several fixes directly related to this case:

  • Ollama's default num_ctx was reduced from 128k to 8192 to avoid consumer-GPU OOM/slowdown.
  • num_ctx can now be set per Ollama credential.
  • thinking-model output handling was fixed.
  • context-length failures stop retrying instead of looping through repeated background attempts.
  • the frontend request timeout is configurable.

After updating, I would test in this order:

  1. In Settings → API Keys, set the Ollama credential's num_ctx to 8192 initially.
  2. Use one model only and confirm it is fully resident with ollama ps. If it re…

Replies: 2 comments

Comment options

You must be logged in to vote
0 replies
Answer selected by lfnovo
Comment options

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
3 participants