Repository navigation
Replies: 1 comment
|
my yesterday summary.. Going to test the system prompt wrapping and some more kv cache today.. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
athena_cheatsheet.docx
FINDINGS.md
I have found how to activate thinking and believe to have found "thinking limitations in qwen 3.5 -.8b and 2b models.. I feel for developers this might be useful informations. I will say the thinking answers arent always better, but the prompt location is different and the model supports lone system prompt injection that holds the config outside of kv cache..
anyways I am still running more data to find the most accurate "thinking band" and "non thinking" band as well as the Thought to reply prefered ratio band too..
I was wondering if anyone else can confirm.. I have for the 0.8b I did the code in straight python.. then I used chat template raw in llama.cpp to confirm myself with the 2b version..
now I would say it was the abliterated model, however I have confirmed it is with any gguf or model of the size.. like the docs say to use <|think|> and maybe that works on the bigger models but "" gives you the one token entrance and exit with "</thin k>" the think should be together, but the chat makes the < t h i n k > dissappear..
i would be happy to answer any questions and will share the new matrix I am running through it right now.. or share the raw data. Being my first post, I'm not sure of file limits for the raw data.. but I would be happy to share it all.
All reactions