Skip to content

Releases: binhu02/repsteer

v0.2.4

Choose a tag to compare

@binhu02 binhu02 released this 11 Sep 23:43

0.2.3 - 2026-09-11

Release Notes

Fixed bugs in v0.2.3 of the ITI implementation.

v0.2.2

Choose a tag to compare

@binhu02 binhu02 released this 11 Sep 01:28

0.2.2 - 2026-09-10

Release notes

  • Artifact-first steering recipes, deterministic selection, and checksummed artifact bundles improve reproducibility and auditability.
  • Versioned adapter capability declarations make supported representation surfaces explicit and fail closed for unsupported controls.
  • LAT now follows RepE-style pair handling, and Hugging Face Qwen3/Qwen3-MoE models are supported.
  • Hardened artifact I/O, expanded CI, and broader regression coverage improve reliability.

Added

  • rs.recipes.directional_ablation(...) and rs.recipes.gated_direction(...) for composing common artifact-backed steering plans.
  • repsteer.selection with deterministic holdout splits, JSON-only candidates, max/min selection, grid search, and checksummed replayable SelectionReport records.
  • ArtifactBundle with role-keyed components, nested artifact verification, provenance, optional selection evidence, and component-level compatibility diagnostics.
  • Versioned AdapterCapabilities, SampleMappingCapability, and ModalityMappingCapability declarations for text and VLM adapters.
  • Qwen3Adapter support for standard Qwen3 decoder-only and Qwen3-MoE architectures.
  • CPU CI across Python 3.10–3.14, pre-commit checks, and expanded tests for artifacts, selection, bundles, adapters, hooks, recipes, and multimodal runtime behavior.

Changed

  • LAT now supports deterministic pair shuffling, explicit pair_signs, configurable shuffle_pair_order, pairwise sign voting, and RepE-compatible raw-difference scaling.
  • Built-in adapters expose read-only capability records for tested residual, sample-mapping, and modality-mapping surfaces.
  • Sequence-gate decisions are scoped to a single generation call, require one decision per batch item, and fail closed for targets without an explicit sample-mapping contract.
  • Hugging Face Transformers, PyTorch, and safetensors are now base dependencies; SAE, VLM, and development dependencies remain optional extras.
  • Artifact and selection records use stricter JSON schemas, checksum validation, and deterministic serialization.

Fixed

  • Exact artifact compatibility now rejects explicitly known architecture mismatches.
  • Artifact reads and writes reject symlinks, unsafe paths, unexpected files, malformed manifests, and checksum or digest mismatches.
  • Sequence-gate caches are cleaned up on both successful and exceptional generation exits.
  • Unsupported head-result and attention-bias surfaces now produce explicit fail-closed diagnostics instead of being approximated.

v0.2.1

v0.2.1 Pre-release
Pre-release

Choose a tag to compare

@binhu02 binhu02 released this 24 Jul 15:58

0.2.1 - 2026-07-24

Release notes

  • Native Hugging Face chat-template support for text instruction models and static-image VLMs.
  • Structured messages, conversations, and conversation batches now work directly with generation, capture, learners, and sweeps.
  • Template-specific inputs, rendering choices, and template hashes are recorded in capture fingerprints and artifact provenance.
  • Safer instruction prompting through assistant-prefix defaults, duplicate special-token protection, and fail-closed validation for unsupported inputs.

Added

  • Hugging Face chat-template generation for structured message mappings,
    conversations, and conversation batches. model.generate(messages=...) is
    supported alongside positional and prompt= inputs.
  • chat_template_kwargs on generation, CaptureRequest, and contrastive
    learners for template-specific inputs such as tools, documents, or an
    explicitly selected template.
  • Processor-preferred chat rendering for image generation, with tokenizer
    fallback, so instruction-style VLM prompts can use their native template.

Changed

  • Chat generation adds an assistant generation prompt by default; capture and
    learner capture retain their explicit default of False.
  • Rendered chat text is tokenized with add_special_tokens=False by default
    to avoid duplicating template-owned BOS/EOS/control tokens. An explicit
    tokenizer or processor option may still override that default.
  • Capture cache keys include template kwargs, while learner artifact provenance
    records the resolved chat-rendering choice. Exact artifact compatibility now
    checks learned tokenizer-template hashes; sweep evaluation recognizes a
    one-message chat mapping as a prompt rather than generation keyword args.

Fixed

  • Single message mappings and raw strings explicitly requesting a template are
    normalized before calling apply_chat_template; mixed raw/chat batches and
    disabled templates for structured messages fail with actionable errors.

v0.2.0

v0.2.0 Pre-release
Pre-release

Choose a tag to compare

@binhu02 binhu02 released this 24 Jul 15:49

0.2.0 - 2026-07-23

Release notes

  • SAE steering with portable feature artifacts, decoder-direction additions, and latent clamp/ablate operators.
  • Sequence-level conditional gates and plan-composition diagnostics for controlled multi-intervention execution.
  • Qwen2.5-VL and InternVL adapters with semantic vision/projector sites, image-token and patch selectors, and processor-driven generation.
  • Multimodal provenance, compatibility checks, evaluation schemas, and SAE/VLM guides for reproducible steering workflows.

Added

  • A backend-independent SAE protocol, lazy SAELens provider, custom/functional
    adapters, common SAELens hook-name-to-site inference, and portable
    SAEFeatureArtifact bundles.
  • Supervised SAE feature ranking and causal-effect ranking, including a
    high-level contrastive capture path and supervised shortlisting before a
    user-supplied causal evaluator.
  • SAE decoder-direction steering through the ordinary Add operator,
    direct-latent Clamp/Ablate, and residual-preserving
    SAEClamp/SAEAblate.
  • ProbeGate, SAEActivationGate, cosine/callable gates, and logical gate
    composition. Explicit language-site gates are evaluated in prefill and
    cached per sequence for decode.
  • Plan composition diagnostics for stable execution order, direction norms,
    same-target pairwise cosine, effective rank, condition number,
    orthogonalized bases, additive-fusion candidates, and non-commutative pairs.
  • Qwen2.5-VL and InternVL architecture adapters with language, vision-residual,
    projector-input, and projector-output semantic sites.
  • An adapter-produced ModalityMap plus serializable ImageTokens,
    ImagePatches, and normalized-XYXY ObjectPatches selectors. Qwen mapping
    accounts for dynamic static-image grids, spatial merge, and window order;
    InternVL mapping accounts for CLS offsets, pixel shuffle, and supported tile
    grouping.
  • Static-image processor-driven generation, including native placeholder
    insertion for the unambiguous one-prompt/one-image case and fail-closed
    cardinality checks for batched or multi-image input.
  • Processor and modality provenance in artifacts, exact preprocessing
    fingerprint compatibility, and a finite-only multimodal evaluation schema
    split into representation, causal-behavior, and capability/quality metrics.
  • SAE/VLM guides and a compensatory-steering example that combines a
    vision-side prefill control and language-side decode control in one
    SteeringPlan.

Changed

  • Hugging Face loading now selects an image-to-text model and AutoProcessor
    for recognized VLM configs; generation preserves its validated modality map
    across language decode.
  • The compiler resolves and validates gate read dependencies, rejects
    non-language explicit sequence-gate sites, preserves priority/declaration
    order, and exposes composition data through diagnostics() and
    explain().
  • A statically zero Constant or NormRelative schedule is compiled as a
    strict no-op: no target/gate hooks run and random-number-generator state is
    unchanged.

Fixed

  • Sequence-gate caches are reset at every new generation and cannot leak into a
    first decode call backed by pre-populated KV state.
  • Compiled plans cannot be rebound after their source model wrapper has been
    collected, even if Python later reuses the same object identity.
  • Composition JSON no longer emits non-finite numbers or compares directions
    from different concrete targets.
  • Exact artifact/processor compatibility includes preprocessing
    configuration, not only a declared processor id and revision.
  • Public dtype= loading is translated to the Transformers 4.x-compatible
    torch_dtype keyword across the declared dependency range.
  • InternVL ObjectPatches fails closed for multi-tile source images when the
    processor does not expose tile-to-original-image crop transforms;
    ImagePatches remains explicitly tile-local.

v0.1.0

v0.1.0 Pre-release
Pre-release

Choose a tag to compare

@binhu02 binhu02 released this 23 Jul 15:45
  • Initial static text activation-steering release.
  • Hugging Face causal-LM runtime and semantic architecture adapters.
  • ActAdd, DiffMean/CAA, PCA, LAT, and linear-probe learners.
  • Typed sites, token selectors, schedules, operators, artifacts, and sweeps.

v0.0.0

v0.0.0 Pre-release
Pre-release

Choose a tag to compare

@binhu02 binhu02 released this 16 Jul 18:00
Create python-publish.yml