Optional BF16 scales for Q4_0 for Gemma 4 QAT #26788
Replies: 2 comments
|
I agree, this seems to be very worth discussing. GGUFs should be accurate as possible, especially considering the purpose of QAT is specifically to preserve quality while reducing memory footprint. Yet, I have noticed some degradations with Gemma 4 QAT 26B compared to regular Q4_k in some of my locked real world tests, which is unexpected. And I am far from the only one. The reason is likely that exact one and the data Unsloth has gathered proves it, KLD and Top-P matching both unexpectedly high/low versus the BF16 QAT model. Another reason could be that crucial layers of the model are quanted to q4_0, like token embeddings and attention layers, QAT might be unable to counteract the huge loss in numerical precision. @JohannesGaessler @ggerganov Is there any reason why we can't inference the model in BF16? I am aware there are is hardware that does not support it (mine included), but an optional pathway for accuracy can't hurt. Plus, BF16 KV Cache already seems to work well, even on my machine without hardware support for BF16. @osanseviero and @danielhanchen are also welcome to join the discussion. Apparently the Google Team has already been contacted by Unsloth regarding this matter. |
Uh oh!
There was an error while loading. Please reload this page.
"The main issue is converting from QAT BF16 to llama.cpp's Q4_0 format is not lossless. llama.cpp uses F16 scales, whilst QAT BF16 uses BF16 scales, and the scales are not determined optimally in llama.cpp land."
Even with unsloth's conversion workaround, they still report up to 0.13288 mean kld and as low as 85.63% top-1 matching. Ideally we could store and inference GGUFs with BF16 scales.
All reactions