Releases: binhu02/repsteer
Releases · binhu02/repsteer
Release list
v0.2.4
v0.2.2
0.2.2 - 2026-09-10
Release notes
- Artifact-first steering recipes, deterministic selection, and checksummed artifact bundles improve reproducibility and auditability.
- Versioned adapter capability declarations make supported representation surfaces explicit and fail closed for unsupported controls.
- LAT now follows RepE-style pair handling, and Hugging Face Qwen3/Qwen3-MoE models are supported.
- Hardened artifact I/O, expanded CI, and broader regression coverage improve reliability.
Added
rs.recipes.directional_ablation(...)andrs.recipes.gated_direction(...)for composing common artifact-backed steering plans.repsteer.selectionwith deterministic holdout splits, JSON-only candidates, max/min selection, grid search, and checksummed replayableSelectionReportrecords.ArtifactBundlewith role-keyed components, nested artifact verification, provenance, optional selection evidence, and component-level compatibility diagnostics.- Versioned
AdapterCapabilities,SampleMappingCapability, andModalityMappingCapabilitydeclarations for text and VLM adapters. Qwen3Adaptersupport for standard Qwen3 decoder-only and Qwen3-MoE architectures.- CPU CI across Python 3.10–3.14, pre-commit checks, and expanded tests for artifacts, selection, bundles, adapters, hooks, recipes, and multimodal runtime behavior.
Changed
- LAT now supports deterministic pair shuffling, explicit
pair_signs, configurableshuffle_pair_order, pairwise sign voting, and RepE-compatible raw-difference scaling. - Built-in adapters expose read-only capability records for tested residual, sample-mapping, and modality-mapping surfaces.
- Sequence-gate decisions are scoped to a single generation call, require one decision per batch item, and fail closed for targets without an explicit sample-mapping contract.
- Hugging Face Transformers, PyTorch, and safetensors are now base dependencies; SAE, VLM, and development dependencies remain optional extras.
- Artifact and selection records use stricter JSON schemas, checksum validation, and deterministic serialization.
Fixed
- Exact artifact compatibility now rejects explicitly known architecture mismatches.
- Artifact reads and writes reject symlinks, unsafe paths, unexpected files, malformed manifests, and checksum or digest mismatches.
- Sequence-gate caches are cleaned up on both successful and exceptional generation exits.
- Unsupported head-result and attention-bias surfaces now produce explicit fail-closed diagnostics instead of being approximated.
v0.2.1
0.2.1 - 2026-07-24
Release notes
- Native Hugging Face chat-template support for text instruction models and static-image VLMs.
- Structured messages, conversations, and conversation batches now work directly with generation, capture, learners, and sweeps.
- Template-specific inputs, rendering choices, and template hashes are recorded in capture fingerprints and artifact provenance.
- Safer instruction prompting through assistant-prefix defaults, duplicate special-token protection, and fail-closed validation for unsupported inputs.
Added
- Hugging Face chat-template generation for structured message mappings,
conversations, and conversation batches.model.generate(messages=...)is
supported alongside positional andprompt=inputs. chat_template_kwargson generation,CaptureRequest, and contrastive
learners for template-specific inputs such as tools, documents, or an
explicitly selected template.- Processor-preferred chat rendering for image generation, with tokenizer
fallback, so instruction-style VLM prompts can use their native template.
Changed
- Chat generation adds an assistant generation prompt by default; capture and
learner capture retain their explicit default ofFalse. - Rendered chat text is tokenized with
add_special_tokens=Falseby default
to avoid duplicating template-owned BOS/EOS/control tokens. An explicit
tokenizer or processor option may still override that default. - Capture cache keys include template kwargs, while learner artifact provenance
records the resolved chat-rendering choice. Exact artifact compatibility now
checks learned tokenizer-template hashes; sweep evaluation recognizes a
one-message chat mapping as a prompt rather than generation keyword args.
Fixed
- Single message mappings and raw strings explicitly requesting a template are
normalized before callingapply_chat_template; mixed raw/chat batches and
disabled templates for structured messages fail with actionable errors.
v0.2.0
0.2.0 - 2026-07-23
Release notes
- SAE steering with portable feature artifacts, decoder-direction additions, and latent clamp/ablate operators.
- Sequence-level conditional gates and plan-composition diagnostics for controlled multi-intervention execution.
- Qwen2.5-VL and InternVL adapters with semantic vision/projector sites, image-token and patch selectors, and processor-driven generation.
- Multimodal provenance, compatibility checks, evaluation schemas, and SAE/VLM guides for reproducible steering workflows.
Added
- A backend-independent SAE protocol, lazy SAELens provider, custom/functional
adapters, common SAELens hook-name-to-site inference, and portable
SAEFeatureArtifactbundles. - Supervised SAE feature ranking and causal-effect ranking, including a
high-level contrastive capture path and supervised shortlisting before a
user-supplied causal evaluator. - SAE decoder-direction steering through the ordinary
Addoperator,
direct-latentClamp/Ablate, and residual-preserving
SAEClamp/SAEAblate. ProbeGate,SAEActivationGate, cosine/callable gates, and logical gate
composition. Explicit language-site gates are evaluated in prefill and
cached per sequence for decode.- Plan composition diagnostics for stable execution order, direction norms,
same-target pairwise cosine, effective rank, condition number,
orthogonalized bases, additive-fusion candidates, and non-commutative pairs. - Qwen2.5-VL and InternVL architecture adapters with language, vision-residual,
projector-input, and projector-output semantic sites. - An adapter-produced
ModalityMapplus serializableImageTokens,
ImagePatches, and normalized-XYXYObjectPatchesselectors. Qwen mapping
accounts for dynamic static-image grids, spatial merge, and window order;
InternVL mapping accounts for CLS offsets, pixel shuffle, and supported tile
grouping. - Static-image processor-driven generation, including native placeholder
insertion for the unambiguous one-prompt/one-image case and fail-closed
cardinality checks for batched or multi-image input. - Processor and modality provenance in artifacts, exact preprocessing
fingerprint compatibility, and a finite-only multimodal evaluation schema
split into representation, causal-behavior, and capability/quality metrics. - SAE/VLM guides and a compensatory-steering example that combines a
vision-side prefill control and language-side decode control in one
SteeringPlan.
Changed
- Hugging Face loading now selects an image-to-text model and
AutoProcessor
for recognized VLM configs; generation preserves its validated modality map
across language decode. - The compiler resolves and validates gate read dependencies, rejects
non-language explicit sequence-gate sites, preserves priority/declaration
order, and exposes composition data throughdiagnostics()and
explain(). - A statically zero
ConstantorNormRelativeschedule is compiled as a
strict no-op: no target/gate hooks run and random-number-generator state is
unchanged.
Fixed
- Sequence-gate caches are reset at every new generation and cannot leak into a
first decode call backed by pre-populated KV state. - Compiled plans cannot be rebound after their source model wrapper has been
collected, even if Python later reuses the same object identity. - Composition JSON no longer emits non-finite numbers or compares directions
from different concrete targets. - Exact artifact/processor compatibility includes preprocessing
configuration, not only a declared processor id and revision. - Public
dtype=loading is translated to the Transformers 4.x-compatible
torch_dtypekeyword across the declared dependency range. - InternVL
ObjectPatchesfails closed for multi-tile source images when the
processor does not expose tile-to-original-image crop transforms;
ImagePatchesremains explicitly tile-local.
v0.1.0
- Initial static text activation-steering release.
- Hugging Face causal-LM runtime and semantic architecture adapters.
- ActAdd, DiffMean/CAA, PCA, LAT, and linear-probe learners.
- Typed sites, token selectors, schedules, operators, artifacts, and sweeps.