An open-source PyTorch library for interpreting and validating large vision models.
Read the paper now as part of Nature Machine Intelligence (Open Access).
SemanticLens is a universal framework for explaining and validating large vision models. While deep learning models are powerful, their internal workings are often a "black box," making them difficult to trust and debug. SemanticLens addresses this by mapping the internal components of a model (like neurons or filters) into the rich, semantic space of a foundation model (e.g., CLIP or SigLIP).
This allows you to "translate" what the model is doing into a human-understandable format, enabling you to search, analyze, and audit its internal representations.
Overview of the SemanticLens framework as introduced in our research paper.
The core workflow of SemanticLens involves three main steps:
-
Collect: For each component in a model M, we identify the data samples that cause the highest activation (the "concept examples"). We provide a suite of
ComponentVisualizersthat implement different strategies, from simple activation maximization to relevance-maximization and attribution-based cropping. -
Embed: These examples are then fed into a foundation model (like CLIP), which creates a meaningful vector representation for each component. SemanticLens includes built-in support for OpenCLIP and can be easily extended with other foundation models (see base.py).
-
Analyze: These vector representations enable powerful analyses. The
Lensclass is the main interface for this, orchestrating the preprocessing, caching, and evaluation needed to search and audit your model using its new semantic embeddings.
You can install SemanticLens directly from PyPI:
pip install semanticlensTo install the latest version from this repository:
pip install git+https://github.com/jim-berend/semanticlens.gitExample usage:
import semanticlens as sl
... # dataset and model setup
# Initialization
cv = sl.component_visualization.ActivationComponentVisualizer(
model,
dataset_model,
dataset_fm,
layer_names=layer_names,
device=device,
cache_dir=cache_dir,
)
fm = sl.foundation_models.OpenClip(url="RN50", pretrained="openai", device=device)
lens = sl.Lens(fm, device=device)
# Semantic Embedding
concept_db = lens.compute_concept_db(cv, batch_size=128, num_workers=8)
aggregated_cpt_db = {k: v.mean(1) for k, v in concept_db.items()}
# Analysis
polysemanticity_scores = lens.eval_polysemanticity(concept_db)
search_results = lens.text_probing(["cats", "dogs"], aggregated_cpt_db)
...- Full quickstart guide: quickstart.ipynb
- Package documentation: docs
We welcome contributions to SemanticLens! Whether you're fixing a bug, adding a new feature, or improving the documentation, your help is appreciated.
If you'd like to contribute, please follow these steps:
- Fork the repository on GitHub.
- Create a new branch for your feature or bug fix (git checkout -b feature/your-feature-name).
- Make your changes and commit them with a clear message.
- Open a pull request to the main branch of the original repository.
For bug reports or feature requests, please use the GitHub Issues section. Before starting work on a major change, it's a good idea to open an issue first to discuss your plan.
@article{dreyer_mechanistic_2025,
title = {Mechanistic understanding and validation of large {AI} models with {SemanticLens}},
copyright = {2025 The Author(s)},
issn = {2522-5839},
url = {https://www.nature.com/articles/s42256-025-01084-w},
doi = {10.1038/s42256-025-01084-w},
language = {en},
urldate = {2025-08-18},
journal = {Nature Machine Intelligence},
author = {Dreyer, Maximilian and Berend, Jim and Labarta, Tobias and Vielhaben, Johanna and Wiegand, Thomas and Lapuschkin, Sebastian and Samek, Wojciech},
month = aug,
year = {2025},
note = {Publisher: Nature Publishing Group},
keywords = {Computer science, Information technology},
pages = {1--14},
}