A mechanistic interpretability workstation that maps neuron activations and polysemantic features across LLaMA and GPT-2 weights using Sparse Autoencoders (SAEs) and TransformerLens. Patch-isolates the Indirect Object Identification circuit to verify linear Taylor approximations through an interactive dashboard. (326 characters)
postgresql telemetry python3 pytorch circuit-analysis sparse-autoencoders ai-safety sae asyncpg linear-probing ai-alignment fastapi mechanistic-interpretability representation-engineering activation-steering gemma-scope transformer-lens
-
Updated
Aug 7, 2026 - JavaScript