Hi, I downloaded DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf from antirez HF repo, it works quite well aside the facts that it thinks forever.
I'm using it through pi agent using the exact configuration from antirez which I previously used for the preview version, but thinking now has gone mad, I currently set it to low (but I seen no difference from medium or high) and to just update a bunch of files on my llm-wiki it generated 9k tokens of reasoning!
It goes all the the time Mhh wait let me reconsider... it became qwen3.5 all of a sudden lol
Have you observed the same problem? Are there any workaround ?
Like this it is unusable, I'm on a m2 max with 96GB and between the low inference speed and infinite thinking it is very painful to use ....
Thanks
Hi, I downloaded DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf from antirez HF repo, it works quite well aside the facts that it thinks forever.
I'm using it through pi agent using the exact configuration from antirez which I previously used for the preview version, but thinking now has gone mad, I currently set it to low (but I seen no difference from medium or high) and to just update a bunch of files on my llm-wiki it generated 9k tokens of reasoning!
It goes all the the time Mhh wait let me reconsider... it became qwen3.5 all of a sudden lol
Have you observed the same problem? Are there any workaround ?
Like this it is unusable, I'm on a m2 max with 96GB and between the low inference speed and infinite thinking it is very painful to use ....
Thanks