Skip to content

Repository files navigation

Sparse Autoencoder Features

Interpreting the ultra-low density cluster in sparse autoencoders from Anthropic.

Interesting findings:

  • Features in ultra low cluster have very high loss
  • Average loss vs frequency shows phase transition
  • Low loss features more likely to be interpretable
Screenshot 2023-12-05 at 10 37 37 PM

Successfully interpreted several low frequency features. Built upon opensourced work from Neel Nanda.

Notebook | Report

About

Interpreting the ultra-low density cluster in sparse autoencoders from Anthropic's Towards Monosemanticity paper

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages