Repository navigation
v0.2.0
Pre-release
Pre-release
0.2.0 - 2026-07-23
Release notes
- SAE steering with portable feature artifacts, decoder-direction additions, and latent clamp/ablate operators.
- Sequence-level conditional gates and plan-composition diagnostics for controlled multi-intervention execution.
- Qwen2.5-VL and InternVL adapters with semantic vision/projector sites, image-token and patch selectors, and processor-driven generation.
- Multimodal provenance, compatibility checks, evaluation schemas, and SAE/VLM guides for reproducible steering workflows.
Added
- A backend-independent SAE protocol, lazy SAELens provider, custom/functional
adapters, common SAELens hook-name-to-site inference, and portable
SAEFeatureArtifactbundles. - Supervised SAE feature ranking and causal-effect ranking, including a
high-level contrastive capture path and supervised shortlisting before a
user-supplied causal evaluator. - SAE decoder-direction steering through the ordinary
Addoperator,
direct-latentClamp/Ablate, and residual-preserving
SAEClamp/SAEAblate. ProbeGate,SAEActivationGate, cosine/callable gates, and logical gate
composition. Explicit language-site gates are evaluated in prefill and
cached per sequence for decode.- Plan composition diagnostics for stable execution order, direction norms,
same-target pairwise cosine, effective rank, condition number,
orthogonalized bases, additive-fusion candidates, and non-commutative pairs. - Qwen2.5-VL and InternVL architecture adapters with language, vision-residual,
projector-input, and projector-output semantic sites. - An adapter-produced
ModalityMapplus serializableImageTokens,
ImagePatches, and normalized-XYXYObjectPatchesselectors. Qwen mapping
accounts for dynamic static-image grids, spatial merge, and window order;
InternVL mapping accounts for CLS offsets, pixel shuffle, and supported tile
grouping. - Static-image processor-driven generation, including native placeholder
insertion for the unambiguous one-prompt/one-image case and fail-closed
cardinality checks for batched or multi-image input. - Processor and modality provenance in artifacts, exact preprocessing
fingerprint compatibility, and a finite-only multimodal evaluation schema
split into representation, causal-behavior, and capability/quality metrics. - SAE/VLM guides and a compensatory-steering example that combines a
vision-side prefill control and language-side decode control in one
SteeringPlan.
Changed
- Hugging Face loading now selects an image-to-text model and
AutoProcessor
for recognized VLM configs; generation preserves its validated modality map
across language decode. - The compiler resolves and validates gate read dependencies, rejects
non-language explicit sequence-gate sites, preserves priority/declaration
order, and exposes composition data throughdiagnostics()and
explain(). - A statically zero
ConstantorNormRelativeschedule is compiled as a
strict no-op: no target/gate hooks run and random-number-generator state is
unchanged.
Fixed
- Sequence-gate caches are reset at every new generation and cannot leak into a
first decode call backed by pre-populated KV state. - Compiled plans cannot be rebound after their source model wrapper has been
collected, even if Python later reuses the same object identity. - Composition JSON no longer emits non-finite numbers or compares directions
from different concrete targets. - Exact artifact/processor compatibility includes preprocessing
configuration, not only a declared processor id and revision. - Public
dtype=loading is translated to the Transformers 4.x-compatible
torch_dtypekeyword across the declared dependency range. - InternVL
ObjectPatchesfails closed for multi-tile source images when the
processor does not expose tile-to-original-image crop transforms;
ImagePatchesremains explicitly tile-local.