Light-Mode, Keep-Alive
For those using a bright luminance terminal background, this update should help with that
-
--light-modewill load up an appropriate color wheel for all the pretties bring printed to the screen. Also, the System Promp is being told about your preference. Some models actually seem to obey. Your mileage may vary (they will/should attempt to use higher contrast emoji that work will with light backgrounds). -
Keep Alive Hack I found Ollama would unload my models even after explicitly telling it to do otherwise. And while using another chat application Enchanted (do check them out!!!), I noticed their app routinely "pings" the Ollama server. I quickly noticed how much quicker their app would garner a response from LLMs. And I want the same. So, I am adding a threaded operation to ping Ollama. Poof. We can now enjoy 10 second reduction in LLM first token response times depending on model size (Qwen 235B, reduced by 10/s).