Repository navigation
Releases: wrynx/undercurrent
Releases · wrynx/undercurrent
Release list
v0.1.0
The first public release of Undercurrent, Wrynx's activation-probing
platform for LLM inference. Earlier versions were internal development builds
and were never published.
Planned work that is not part of this release (an OpenAI-compatible
undercurrent serve server, Hugging Face Hub integration, a published
container image, and more) is tracked in the
roadmap.
Added
- One installable package.
pip install undercurrentinstalls the
undercurrentpackage for Python 3.10+, with the Hugging Face
transformersbackend in the base install. vLLM is never installed by
default. Tested on Python 3.10–3.13 with torch 2.1+ and transformers
4.40–5.x; see Compatibility. ProbedModel, the high-level API (undercurrent.ProbedModel):
ProbedModel.from_pretrained(model, spec=..., probes=...), then
generate(), which returns aGenerationOutputwith the text and each
extraction point'sProbeResult. Backends"hf"(default) and"vllm".
Passon_result=callbackfor a simple result hook alongside the sink API.- Extraction-point specs (
undercurrent.spec). Declare what to capture
in YAML (load_yaml_file,parse_yaml,parse_dict) or construct
ExtractionPoints in Python. Specs are validated by pydantic models and
round-trip back to YAML (to_yaml).- Tensor types
residual_stream,attn_out,mlp_outandfinal_norm,
selected per layer. - A position-selector grammar: absolute indices (
5), prompt and generated
indices (prompt[-1],generated[0]), every generated token
(generated[*]), generated slices (generated[5:],generated[2:8]),
and offsets (prompt[-1]+1). Malformed selectors raise
PositionSyntaxErrorwith the accepted forms. untilconditions for continuous selectors (generation_end,
fixed_count,stop_token).- Example specs in
examples/specs/.
- Tensor types
- Probe interface (
undercurrent.core). Probes implement the lifecycle
spawn→on_start→on_activation* →on_endand return
ProbeSignals and a finalProbeResult. There are two probe kinds,
single_shotandtrajectory. A fresh instance is spawned for each
request and extraction point, and probe classes that declare mutable
class-level attributes are rejected when the class is defined, so instances
can't share state across requests. Reference probes are in
undercurrent.core.examples:MLPClassifierProbe(single-shot) and
TrajectoryScoreProbe(trajectory). - Function probes and a probe registry. Decorate a plain function
(record) -> float | bool | ProbeSignal | Nonewith@probe("name", ...)
for stateless single-shot probes, and register probe classes by name with
@register_probe("name"). Third-party packages can expose probes through
theundercurrent.probesentry-point group. - Router (
undercurrent.router). Dispatches activation records to
per-request probe instances. Each extraction point picks an execution mode:inline: the probe runs on the generation path and its signal is
returned to the adapter, so it can abort generation.async: the probe runs out of band and doesn't block the response. Each
binding has a bounded queue (DEFAULT_QUEUE_DEPTH), served by a shared
worker pool, with a configurableOverflowPolicy(drop_oldest,
drop_newestorblock).
- Interventions. An inline probe can return
ProbeAction.ABORTto stop
generation. With theblock_until_signalintervention policy, the router
waits up totimeout_msfor the probe, then applieson_timeout
(continueorabort). A circuit breaker downgrades an extraction point
to observe-only aftercircuit_breaker_thresholdconsecutive timeouts or
probe errors, so a failing probe can't keep stalling generation. - Metrics. Per-binding queue depth, drops, activation latency and probe
errors go to a pluggableMetricsSink.InMemoryMetricsRegistryprovides
snapshots. - Observation sinks (
undercurrent.sinks) for async probe output:FileLogSinkwrites NDJSON.WebhookLogSinkPOSTs JSON from a background thread, so a slow endpoint
never adds latency to the caller. Failed deliveries are retried with
exponential backoff, records that exhaust their retries are
dead-lettered to a file, and the send queue is bounded.- A
LogSinkbase class for writing your own sink.
- Sink redaction.
redact_keys(),drop_keys()andchain()control
what a sink writes.WebhookLogSinkredacts prompt text by default. - Router context managers.
Routerandrouter.request(...)can be used
inwithblocks so requests end and workers shut down cleanly. - Hugging Face adapter (
undercurrent.adapters.hf, in the base install).
HFEngineAdaptercaptures activations fromtransformersmodels with
forward hooks duringgenerate(), and aborts generation through a
StoppingCriteria(ProbingStoppingCriteria). - vLLM adapter (
undercurrent.adapters.vllm). Install Undercurrent into
your existing vLLM environment or image (the recommended production path),
or use theundercurrent[vllm]extra for a vLLM from the tested range.- Capture is aware of continuous batching: a worker extension installs the
hooks, andSeqIdMappertranslates batch rows to
(request_id, token_pos). Abort signals are wired back to the engine. - A
vllm.general_pluginsentry point registers the worker extension. - Constructing the adapter checks the installed vLLM version against the
supported range and raisesVLLMAdapterLimitationErrorwith the
installed version and the supported range. Set
UNDERCURRENT_ALLOW_UNSUPPORTED_VLLM=1to downgrade the error to a
warning. See
docs/compatibility.md. - Tensor and pipeline parallelism: results from every worker are merged,
and events that tensor-parallel ranks duplicate are deduplicated.
Multi-worker topologies are refused unless you pass
allow_unsupported_executor=True, because cross-rank ordering under
pipeline parallelism and Ray/distributed executors haven't been
validated. See the
tensor and pipeline parallelism notes.
- Capture is aware of continuous batching: a worker extension installs the
- Command-line tool
undercurrent:inspect-modellists a model's
hookable layers and supported tensor types,validatechecks spec files,
andschemaprints the spec JSON Schema (also in
schema/probe-spec.schema.jsonandundercurrent.spec.json_schema()). - Curated top-level API. The common names import from
undercurrent
directly; see API stability
for what is public.load_spec()
loads a spec from a path, YAML string or dict. The package shipspy.typed. - Error hierarchy. Errors Undercurrent raises on purpose derive from
ProbingError(and the matching built-in type), with messages that say
how to fix the problem. - Clear errors for missing backends. Using an adapter whose backend isn't
installed raisesMissingDependencyError, which says how to install it. - Examples (in the repository, not installed with the package, and covered
by the test suite):examples/content_safety/:
a demo content-safety probe in bothsingle_shotandtrajectory
variants, with YAML specs, a dummy-checkpoint maker, a vLLM pipeline
script and a Dockerfile. It is a demonstration of the probe API, not a
trained safety classifier.examples/openai_server/:
a reference OpenAI-compatible wire format for probe verdicts (completion
bodies and SSE chunks) and a reference HTTP server built on it.
- Probe-training example (
examples/train_probe/) and Colab notebooks
(notebooks/quickstart.ipynb,notebooks/train_probe.ipynb).
Fixed
- vLLM
residual_streamon fused-residual layers. On Llama-family models
(and the other vLLM architectures whose decoder layer returns
(hidden_states, residual)), aresidual_streamextraction point captured
the layer's MLP output. It now captureshidden_states + residual, the
residual stream after the layer, matching the HF backend. Decoder-layer
outputs the adapter can't classify raiseVLLMAdapterLimitationError
instead of being captured silently. Verified on GPU (NVIDIA L4, vLLM 0.28.0)
bytests/adapters/vllm/test_residual_stream_gpu.py, which compares the
vLLM and HF captures token by token.
Known issues
See Known issues (v0.1)
for details and workarounds.
- No GPU CI. The GPU tests run by hand with
scripts/gpu_check.shbefore
each release. 0.1.0 passed on an NVIDIA L4 with vLLM 0.28.0, torch 2.13.0
and CUDA 13.0; other GPUs and multi-GPU topologies haven't been run. - Only vLLM 0.28 is supported (
vllm>=0.28,<0.29), and vLLM 0.30 is
already out.UNDERCURRENT_ALLOW_UNSUPPORTED_VLLM=1lets you try a newer
vLLM; the range widens only after a GPU validation run.
Security
- Vulnerabilities can be reported privately to security@wrynx.com or through
GitHub private vulnerability reporting. See
SECURITY.md. - The content-safety example loads classifier checkpoints only with
torch.load(..., weights_only=True), accepts only plain tensor
state_dicts, and refuses to load on torch older than 2.6 (the first
release with a fix for CVE-2025-32434).
v0.1.0rc1
What's Changed
- Check repo, Pages, Colab and PyPI links now that they are live by @alizishaan in #7
- Finalize the 0.1.0 changelog by @alizishaan in #8
- Fix vLLM adapter shutdown, GPU test hangs, and GPT-2 downloads by @alizishaan in #9
- Record the 0.1.0 GPU verification (NVIDIA L4, vLLM 0.28.0) by @alizishaan in #10
New Contributors
- @alizishaan made their first contribution in #7
Full Changelog: https://github.com/wrynx/undercurrent/commits/v0.1.0rc1