You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
KV cache precision tail – very interesting new feature in v0.4.0, where latest X tokens in KV cache are stored in (B)F16. Shows massive precision improvement with KLD, and can be very useful in real tasks to preserve details from prompt, code, etc.: https://anbeeld.com/articles/kv-cache-precision-tail-implementation-and-benchmarks
BeeLlama also closely follows upstream these days, although I gave up speculative decoding stack of favour of upstream implementation. But there's more stability now.
I also added more standard quants: q6_0, q6_1, and low-bit from q2_0 to q3_1, replacing options that were limited to TurboQuant.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Hello there! I updated articles on my website with up-to-date info about BeeLlama features.
Mature KVarN implementation, with great results on Qwen, kvarn6 consistently looks better than q8_0: https://anbeeld.com/articles/kvarn-kv-cache-implementation-and-benchmarks
KV cache precision tail – very interesting new feature in v0.4.0, where latest X tokens in KV cache are stored in (B)F16. Shows massive precision improvement with KLD, and can be very useful in real tasks to preserve details from prompt, code, etc.: https://anbeeld.com/articles/kv-cache-precision-tail-implementation-and-benchmarks
Massive research involving 413 pairs of KV cache tested in KLD with both Qwen and Gemma, including standard quants, KVarN, precision tail: https://anbeeld.com/articles/kv-cache-quantization-benchmarks-kvarn-precision-tail
BeeLlama also closely follows upstream these days, although I gave up speculative decoding stack of favour of upstream implementation. But there's more stability now.
I also added more standard quants: q6_0, q6_1, and low-bit from q2_0 to q3_1, replacing options that were limited to TurboQuant.
Would be grateful if you checked it out!
All reactions