Proposal: activation capture, probes, and model steering in EDSL, with a Modal-backed service #2651
johnjosephhorton
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Proposal
Add support for activation capture, probe training, and inference-time steering to EDSL, with an optional Expected Parrot-hosted backend. Modal looks like a plausible infrastructure provider for that backend.
The motivating paper is Interpreting and Steering LLM Agents for Social Simulations (PDF). It compares prompt-based manipulation, sparse autoencoder (SAE) feature steering, and probe-direction steering in social simulation tasks. The paper already uses EDSL, including a custom Goodfire controller bridge for SAE steering; its probe experiments use PyTorch/Hugging Face hooks and scikit-learn logistic regression. Strong prompting can outperform SAE steering on creativity tasks, so strong prompt baselines should be part of the design.
The proposed user workflow is:
This is a design proposal, not an implemented feature. The new APIs below are illustrative. No hosted deployment or performance benchmark has been completed.
What probes actually do
A probe is a small model trained to read information from a larger model's internal activations. The base language model stays frozen.
For a lottery experiment:
hat a specified intermediate layer and token position, such as the final prompt token before generation.score = wᵀh + b.Only the probe's weights change during fitting. If the activations and labels have already been saved, fitting the probe does not require loading the original language model.
Steering is a separate operation. During a new generation, intercept the activation and replace it with:
Then let the remaining layers continue. Positive and negative strengths move the state toward opposite sides of the probe's decision boundary. This modifies the current computation, not the prompt or the model's stored weights.
Reading activations is enough for probe fitting and scoring. Steering additionally requires write access during inference. These are usually intermediate-layer activations; output text and token log probabilities are not substitutes.
Increasing the linear probe's score does not guarantee a corresponding behavioral change. A probe might recognize payoff size, wording, or an answer already present in the context instead of a general behavioral trait. Prediction, intervention effects, and cross-task generalization therefore need separate evaluation.
SAEs offer a related interface: inspect sparse features and intervene on selected features. They require a compatible SAE as well as access to the model runtime. Training new SAEs would be a separate, larger project.
What writing this in EDSL could look like
Ordinary experiment: existing EDSL functionality
Here
disable_remote_inference=Truemeans EDSL contacts the configured endpoint directly instead of using Expected Parrot inference. It does not mean EDSL loads weights locally. A hosted Expected Parrot integration would use its remote execution path instead.Prompt manipulation already fits the existing interface: replace the neutral agent with, for example,
Agent(traits={"risk_attitude": "Enjoys taking risks"}), and rerun the survey. Few-shot examples and reasoning prompts can also be expressed through ordinary survey/prompt construction.Steering: proposed functionality
The survey and neutral agent stay the same. The model configuration tells a compatible server where and how to intervene. An ordinary chat-completions server would not support this automatically.
with_intervention()should return a separate configuration so the baseline model remains unchanged. Multiple strengths could become aModelListsweep. The intervention is a model-runtime setting; any higher-level agent interface would ultimately need to resolve to such a setting for each request.Preparing a probe: proposed functionality
Layer selection, regularization, and strength calibration use development data. Final evaluation uses untouched scenarios, ideally including new task formulations. The first artifact should be described as a risky-choice probe until stronger evidence supports interpreting it as a general risk-preference measure.
The
.probesuffix is illustrative. Whether probes should become first-class.eppackages is an open design question.Integration points and missing pieces
EDSL already has an OpenAI-compatible inference service. In the checkout inspected for this proposal, its inherited request builder forwards a fixed parameter set. Simply adding a
controllerorsteeringfield would not send it to the server. The Goodfire bridge described in the paper was not present in the inspected inference-service implementation.The additions would be:
Could Expected Parrot host this on Modal?
Yes. Modal supplies custom GPU containers, multi-GPU execution, HTTP serving, and autoscaling. Expected Parrot would build and operate the instrumented runtime and experiment integration on top. See GPU support and serving guidance.
The minimum backend needs three operations:
Probe fitting can run separately, often on CPU. Workers should load weights once per container and reuse them across requests. Large captures should be written as artifacts, with compact references returned through the API.
Start with one instruction-tuned model around 8B parameters and a straightforward PyTorch/Hugging Face runtime. Its BF16 weights alone occupy approximately 16 GB; a 48 GB GPU is a plausible starting point, subject to actual context, batch, and runtime memory requirements. Larger models and optimized serving engines can follow measured demand.
Accept structured operations initially rather than arbitrary user Python. This keeps the public service's execution and isolation requirements bounded. Users would need neither a Modal account nor their own GPU setup.
Experimental correctness and operational requirements
Modal GPU snapshots may help some cold-start workloads, but they remain an alpha feature with limitations, including multi-GPU compatibility. They should be an optional optimization, not a requirement for the first version. Snapshot documentation
Alternatives and possible partners
These findings come from provider documentation reviewed on September 16, 2026; availability and access conditions should be rechecked before integration.
The intervention specification and artifact formats should be independent of provider so a hosted Expected Parrot backend and external research backends can share the EDSL-facing interface.
Scope, economics, and a first milestone
The recommended first milestone is one model, one lottery survey, prompt baselines, activation capture, a fitted probe, and a held-out steering-strength sweep. Deliver the answers, artifacts, and comparison curves as a reproducible EDSL workflow.
Rough engineering estimates, assuming an experienced developer and available GPU access:
Modal's listed GPU rates at the time of review are approximately $1.95 per L40S GPU-hour and $3.95 per H100 GPU-hour, before CPU, memory, storage, and other charges. Thus 100 L40S GPU-hours would be about $195 in GPU charges. This is not a throughput estimate or a customer price; we need to benchmark model loading, capture, generation, and artifact handling. Pricing
Questions for discussion
.epobjects?All reactions