Interpreting the ultra-low density cluster in sparse autoencoders from Anthropic.
Interesting findings:
- Features in ultra low cluster have very high loss
- Average loss vs frequency shows phase transition
- Low loss features more likely to be interpretable
Successfully interpreted several low frequency features. Built upon opensourced work from Neel Nanda.