Skip to content
Discussion options

You must be logged in to vote

Both the above answers were helpful. I needed to use a smaller quant, and I needed to use the new settings.
New settings on their own didn't help, the Q4_K_XL seems to be just too big, one of the AIs suggested it's to do with the size of the lookup table. So I went with a much smaller quant IQ2_M plus new settings, and this seemed to fix the issue. I did some tests, I wouldn't trust they are consistent from run to run, since one example shows different results for two different runs.

Context   n-cpu-moe   vram  sysram  t/s
No LLM    -           0.5*  10.0    -
65K 1st   40          5.3   56.0    11.2
65K 2nd   40          12.1  48.6    14.2
65K       36          15.5  45.3    15.2
65K    …

Replies: 3 comments

Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Answer selected by mkultra333
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
3 participants