Context Management
This release marks an improvement to contextual token management.
-
Staggered chat history retrieval based on users
--chat-history-session. e.g--chat-history-session 5
Would retrieve the following history slices from all recorded history given there were 20 responses in all to draw from:
[9, 13, 16, 18, 19]
On top of that, there is also a cut-off in tokens at a rate of
550 * 5 = 2750 tokens(--chat-history-session). -
Specialized LLM prompt files. I am allowing for separate prompt files to coexist based on the model in use. So far, There are special prompts for Gemma, Llama, and Owen.
-
Improvements to field-filtering matching with the RAGs. Could be better still #4