Skip to content

v0.2.3

Choose a tag to compare

@binhu02 binhu02 released this 11 Sep 23:18
· 1 commit to main since this release

0.2.3 - 2026-09-11

Release Notes

This release adds first-class ITI (Inference-Time Intervention) support and fixes two steering defaults so they match the original papers.

Added

  • ITI profile learning with one deterministic linear probe per layer/query head, held-out head ranking, mass-mean or probe-weight directions, and per-head projected-standard-deviation calibration.
  • ITIArtifact, ITIAdd, sites.head_result(), and recipes.iti(). The runtime applies alpha * sigma * theta at the selected query-head slots in self_attn.o_proj input.
  • Safe artifact I/O and role-keyed artifact-bundle support for ITI profiles, plus local profile, serialization, adapter-contract, and runtime regressions.

Changed

  • ActAdd/CAA now default to normalize=None (the raw mean(positive - negative) vector), matching the official CAA (Rimsky et al., nrimsky/CAA) and ActAdd (Turner et al., montemac/activation_additions) implementations, neither of which normalizes the mean-difference vector. Previously both defaulted to a unit-length direction, which silently changed the effective magnitude of any strength/multiplier value carried over from the papers. Pass normalize="l2" to keep the previous unit-length behavior.
  • PCA components are now sign-oriented toward the positive group. Each component is flipped, if needed, to align with mean(positive) - mean(negative), instead of the previous label-independent "largest-magnitude coordinate positive" convention. This makes a rank-1 PCA result safe to use directly with Add-style steering. This changes PCA's numeric output for existing callers: a component's sign may now differ from prior versions. Recorded in artifact metadata as sign_alignment: "positive_negative_mean_difference".