-
Notifications
You must be signed in to change notification settings - Fork 10
Dead Latents
I have a bunch of dead latents. I'm training width 16K SAEs on DINOv3 ViT-L activations on FishVista (very homogenous images). I am comparing TopK and ReLU + L1 sparsity penalties and I realized that many, many of my latents are dead.
Here are graphs of the number of dead latents during training, grouped by the relative sparsity penalty:
ReLU + L1 sparsity penalty dead units throughout training on FishVista.
TopK dead units throughout training on FishVista.
My options, based on prior work, include:
- Better W_dec/enc initialization (datapoint initialization, see https://transformer-circuits.pub/2025/october-update/index.html#data-point-init)
- [maybe] Better b_enc initialization (according to Anthropic, see https://transformer-circuits.pub/2025/january-update/index.html#DL )
- [maybe] Gao et al's 'AuxK' loss
- [maybe] Anthropic's pre-act loss (again, see https://transformer-circuits.pub/2025/january-update/index.html#DL)
Starting with (1) seems super easy and reasonable. Doing (2) seems like I need to do a bunch of forward passes, which seems bad. 3 and 4 seem reasonable as well, but I don't really like throwing additional loss terms into my objective.
So I am going to:
- Train ReLU and TopK on ADE20K
- Add better W_dec/W_enc initalization using datapoint init + transposes
- Train ReLU and TopK on ADE20K and FishVista with the better initialization. For all checkpoints trained on layer 24/24:
| Activation | Initialization | Dataset | Dead Unit % |
|---|---|---|---|
| ReLU | Kaiming | ImageNet-1K | 38.89 |
| ReLU | Kaiming | FishVista | 46.96 |
| ReLU | Datapoint | ImageNet-1K | 31.44 |
| ReLU | Datapoint | FishVista | 54.28 |
| TopK | Kaiming | ImageNet-1K | 53.08 |
| TopK | Kaiming | FishVista | 79.09 |
| TopK | Datapoint | ImageNet-1K | 47.22 |
| TopK | Datapoint | FishVista | 77.64 |
For all checkpoints across layers 14, 16, 18, 20, 22 and 24:
| Activation | Initialization | Dataset | Dead Unit % |
|---|---|---|---|
| ReLU | Kaiming | ImageNet-1K | 17.98 |
| ReLU | Kaiming | FishVista | 36.80 |
| ReLU | Datapoint | ImageNet-1K | 17.88 |
| ReLU | Datapoint | FishVista | 51.33 |
| TopK | Kaiming | ImageNet-1K | 29.12 |
| TopK | Kaiming | FishVista | 55.20 |
| TopK | Datapoint | ImageNet-1K | 13.35 |
| TopK | Datapoint | FishVista | 51.46 |
We've had success fixing dead neurons using the Muon or Signum optimizers, or by adding a linear k-decay schedule (all available in EleutherAI/sparsify). The alternative optimizers also seem to speed up training a lot (~50% reduction).
Another approach ... has been adding a small penalty term on the standard deviation of the encoder pre-act (i.e., before the top-k) means across the batch dimension. This has basically eliminated my dead neuron woes.
Update 12/08/2025: Muon alone did not fix our dead latent problem. I'm adding AuxK and will report back.