SAE Lens

SAE Lens

SAELens exists to help researchers:

Train sparse autoencoders.
Analyse sparse autoencoders / research mechanistic interpretability.
Generate insights which make it easier to create safe and aligned AI systems.

Please refer to the documentation for information on how to:

Download and Analyse pre-trained sparse autoencoders.
Train your own sparse autoencoders.
Generate feature dashboards with the SAE-Vis Library.

SAE Lens is the result of many contributors working collectively to improve humanities understanding of neural networks, many of whom are motivated by a desire to safeguard humanity from risks posed by artificial intelligence.

This library is maintained by Joseph Bloom and David Chanin.

Loading Pre-trained SAEs.

Pre-trained SAEs for various models can be imported via SAE Lens. See this page in the readme for a list of all SAEs.

Tutorials

Join the Slack!

Feel free to join the Open Source Mechanistic Interpretability Slack for support!

Citations and References

Research:

Reference Implementations:

Name		Name	Last commit message	Last commit date
Latest commit History 477 Commits
.github		.github
.vscode		.vscode
content		content
docs		docs
sae_lens		sae_lens
scripts		scripts
tests		tests
tutorials		tutorials
.flake8		.flake8
.gitignore		.gitignore
.pre-commit-config.yaml		.pre-commit-config.yaml
.pylintrc		.pylintrc
CHANGELOG.md		CHANGELOG.md
LICENSE		LICENSE
README.md		README.md
__init__.py		__init__.py
makefile		makefile
mkdocs.yml		mkdocs.yml
pyproject.toml		pyproject.toml
requirements.txt		requirements.txt

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

SAE Lens

Loading Pre-trained SAEs.

Tutorials

Join the Slack!

Citations and References

About

Releases

Packages

Languages

License

Lewington-pitsos/SAELens

Folders and files

Latest commit

History

Repository files navigation

SAE Lens

Loading Pre-trained SAEs.

Tutorials

Join the Slack!

Citations and References

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages