π§ͺ The Lobotomized Attention Head Bugβ’ β where one head does all the work while the others stare into the void #6
ToddThomson
started this conversation in
Show and tell
Replies: 1 comment
-
|
πͺ¦ Memorial Style In loving memory of Attention Heads 1β11. |
Beta Was this translation helpful? Give feedback.
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
-
π§ The Lobotomized Attention Head Bug β A Transformer Debugging Rite of Passage
I just finished building the fast prefill β decode inference path in my Mila DNN library.
Everything seemed fine β the model produced coherent text, KV caching worked, and decode mode looked solid.
Butβ¦ my transformerβs residuals were way off compared to π€ HF GPT-2, and the hidden states just felt wrong.
Not exploding, not NaNβing β just wrong.
After hours of combing through attention math, KV cache, QKV packing, layernorm, and positional encodingsβ¦
I found it.
π§ͺ Root Cause
In the MHA prefill path, my unpermute_output kernel was wrong.
It needed a padded variant (unpermute_output_padded), and instead of writing all attention heads back into the output tensorβ¦
It only wrote back ONE head.
All the other heads?
Nowhere.
Silent.
Forgotten.
Lobotomized.
π€‘ Symptoms (that still produced coherent text!)
Hidden states completely misaligned from HF
Residuals with huge swings
Prefill corrupted β Decode still worked (go figure!)
Yet⦠model still produced coherent sentences
(Transformers are absurdly resilient.)
π Why it still βworkedβ
Decode path was correct (so per-token incremental attention was fine)
LayerNorm aggressively stabilized everything
MLP + embeddings carried most of the workload
Attention became βSingle-Head Attention + Moral Supportβ
π Lesson
If you ever see:
Prefill mismatch
HF vs your model drifting hard
Residuals acting hyperactive
Yet decode produces intelligible sentencesβ¦
Check your unpermute logic.
A single bad stride or head offset can quietly βturn offβ most of attention.
Beta Was this translation helpful? Give feedback.
All reactions