Skip to content

Releases: deepfates/imp

Imp v0.5.0

Choose a tag to compare

@deepfates deepfates released this 27 Sep 18:02
068409a

Imp v0.5.0

Imp is a framework for typed, optimizable language-model programs on the BEAM.
Declare a task as named inputs and outputs, call it like any other Elixir
program, measure it on examples, compile it with an optimizer, and run the
selected program under OTP.

This release puts Imp on Hex. It is 0.5.0 rather than a patch because the
install line changes, an OTP release that uses the protocol adapters lists one
more application, and ReActV2 and Imp.MCP.OAuth change shapes a program may
depend on.

Install

{:imp, "~> 0.5"}

Every dependency comes from Hex. Use a path dependency only while developing
against a local checkout.

ExMCP and erlexec are declared runtime: false, so an OTP release that uses
Imp.ACP or Imp.MCP must list applications: [ex_mcp: :load, erlexec: :load]
in its release definition; see releases that use MCP or
ACP
.
Ordinary Imp startup starts no protocol endpoint.

mix deps.get and mix hex.audit report two cowlib advisories
(CVE-2026-43966, CVE-2026-43969). cowlib arrives only through ExMCP's
Cowboy server, and Imp's HTTP goes through Req, Finch and Mint. The first is
fixed one layer up: Cowboy 2.16.0 and later refuse a response header
containing CR or LF, and a fresh mix deps.get resolves Cowboy 2.19.0. The second is in the encoder
for an outgoing Cookie request header, which nothing in Imp's dependency
tree calls, and no cowlib release fixes it yet.

Headline changes

  • GEPA works on agents. Optimizing an Imp.react agent, the reflection model
    reads the whole run (tool calls, tool results, the final answer) and the
    agent's tools, and GEPA rewrites the agent's instruction. By default
    Imp.Optimizer.GEPA behaves as DSPy's GEPA does; Imp's own search is
    execution_profile: :beam_native.
  • An Imp.Deadline reaches the work Imp starts for you: Imp.parallel/3,
    evaluation rows, optimizer workers and runs inherit the caller's deadline,
    and Imp.start_run/3 takes deadline:.
  • Imp depends on ExMCP 1.5 from Hex, unpatched. What Imp needed from the
    deepfates/ex_mcp fork now lives in Imp: stdio MCP servers that end with
    their connection, children included; a clean PATH for them inside a
    release; trust for authorized remote servers; the connection options public
    servers need; and the browser OAuth flow.
  • ReActV2 offers submit only to a signature that needs one. A task with
    exactly one text output ends its turn on a step that answers in text, and an
    interrupted turn makes one last request whose text is the answer instead of
    failing.
  • ReActV2 gains finish_on for tools whose call is the answer.
  • :model_request events record the whole request, and tool definitions are
    emitted once per run as :tools_sent.
  • An MCP tool call that got no answer says whether it was refused, had its
    credential refused, was never sent, or may have run (Imp.MCP.CallFailure,
    Imp.Tool.outcome/1), an error result that declares its outcome is read as
    declared, and a failed tool call reaches the model as plain text.
  • A ReAct prediction's fields are its outputs; how the turn ended is metadata,
    in one vocabulary, with Imp.Prediction.complete?/1.
  • Imp.Run and Imp.ACP refuse options they do not know, and
    Imp.Run.Event.kinds/0 lists every event kind.
  • A host names its own run pool and limit (Imp.Run.start/3's :admission),
    and a failing run event sink is reported to the run's owner.
  • Imp.MCP.connect/2 takes pool_size:, so several calls to one HTTP server
    run at once, and an HTTP call can take as long as its :timeout allows.
  • A run no longer outlives its control process, and a cancellation that never
    returns no longer holds a run.

Breaking changes from v0.4.0

  • Replace {:imp, github: "deepfates/imp", tag: "v0.4.0"} with
    {:imp, "~> 0.5"}. EX_MCP_PATH is no longer read.
  • A release that uses Imp.MCP or Imp.ACP adds erlexec: :load beside
    ex_mcp: :load.
  • Trust for an authorized remote MCP server is VM-wide. While a connection to
    it is open, its exact origin (scheme://host:port) is in ExMCP's
    trusted_origins, so any ExMCP client in the same VM may send credential
    headers to that origin without consent. In 0.4.0 the trust belonged to the
    one connection. No other origin is trusted, the origin is removed when the
    last connection to it closes, and origins the host configured are left
    alone. A host that runs other ExMCP clients it does not trust with those
    origins should know this.
  • Imp.Optimizer.GEPA defaults to DSPy's GEPA, execution_profile: :gepa_v0_1_4_merge: merge on, no evaluation cache, perfect minibatches
    skipped, the pinned RNG, and :generations turned into a metric budget when
    :max_metric_calls is not given. Options the DSPy profiles fix (ComBee,
    :feedback_fn, :module_selector, :candidate_selection_strategy,
    :proposal_concurrency, :reflection_strategy, the frontier, sampling,
    selection, evaluation and acceptance policies, :max_reflection_calls, and
    reflection_record_mode: :beam_native) raise unless execution_profile: :beam_native is given, which is the 0.4.0 behaviour. Resuming a checkpoint
    written by a 0.4.0-default run raises under the new default; resume it with
    execution_profile: :beam_native.
  • Imp.MCP.OAuth.begin/3 no longer takes :flow; a pre-registered client is
    client_registration: {:pre_registered, client_id, client_secret} with
    client_issuer: naming the authorization server it belongs to. A server
    with no OAuth metadata at all is refused instead of given guessed endpoints.
  • An Imp.Tool named with a string keeps the string, and tools imported from
    an MCP server are named by the server's string. Code that compared an
    imported tool's name to an atom compares it to the string.
  • An MCP tool call that got no answer returns
    {:error, %Imp.MCP.CallFailure{}} instead of
    {:mcp_tool_call_failed, server, reason} or
    {:mcp_connection_unavailable, server, reason}. A call that reaches its
    :timeout is %Imp.MCP.CallFailure{outcome: :unknown, reason: :timeout},
    answered at the timeout while the request runs on; a call to an HTTP server
    whose connections all stay busy until the timeout is :not_sent with
    reason: :no_idle_connection.
  • "type" => "sse" is MCP's deprecated HTTP+SSE transport, and its "url"
    is the event stream's. In 0.4.0 it was Streamable HTTP with a standing GET
    stream; a Streamable HTTP server is now "type" => "http". An sse
    descriptor with "headers" or "auth", or with a query string in its URL,
    is refused before anything is dialed (:mcp_sse_credentials_refused,
    :mcp_sse_url_refused): the whole import under the default
    on_failure: :refuse, only that server under on_failure: :drop.
  • When a run's control process ends while the run is still going, the task is
    killed after its registered cancellations are called; its monitor reports
    :killed.
  • A run's owner can receive {:imp_run_event_sink_failed, run_id, details}
    and {:imp_run_event_undelivered, run_id, event}; an owner with a strict
    handle_info/2 needs clauses for them.
  • Imp.Run.start/3, Imp.ACP.start_link/1, Imp.ACP.run/1 and
    Imp.ACP.Local.start_link/1 raise ArgumentError for an option they do
    not know. A transport's own options for Imp.ACP go in
    :transport_options, and :capabilities is spelled :agent_capabilities.
  • Imp.predict/2, Imp.chain_of_thought/2 and Imp.configure/1 raise
    ArgumentError for an option or setting they do not know. Request options
    such as :temperature go under config:; a setting of your own goes
    through Imp.context/2.
  • :max_errors and :retriever are no longer settings, and
    Imp.configure/1 and Imp.context/2 refuse them. Pass :max_errors to
    BootstrapFewShot, RandomSearch or COPRO (10 when not given) and a retriever
    to the program.
  • ReActV2 emits no :final event; :run_finished carries the prediction.
    Imp.Trajectory.to_atif/2's extra.outcome is extra.terminal_event, and
    a tool result's extra.outcome is the recorded Imp.Tool.outcome/1
    instead of "returned" or "error".
  • A ReActV2 or ReAct prediction's fields are its outputs only: history,
    termination_reason, termination_cause, termination_error,
    finished_by_tool, unexecuted_tool_calls and context_projection are in
    prediction.metadata. termination_reason says how the turn ended, and a
    turn without an answer is :incomplete with termination_cause saying why;
    typed extraction is :extracted (no completion_mode), and
    Imp.Predict.ReAct spells :parse_failure as :parse_error and :direct
    as :answered. Use Imp.Prediction.complete?/1 to ask whether a turn
    answered.
  • For a signature with one :string output, ReActV2 offers no submit
    tool, and a step answered in text with no tool call ends the turn.
  • Errors have one shape per tag, with the reason as a term. A failed
    Imp.Clients.ReqLLM request is %Imp.LMError{} (with status,
    retryable and context_window_exceeded; Imp.ContextWindowExceededError
    is gone), and a completion that cannot be parsed is
    %Imp.AdapterParseError{kind: ...}, which Imp.Predict returns
    directly instead of %{reason: {:error, _}, trace: _}. A raise inside a
    client, program, tool, tool policy, retriever, optimizer or ACP callback
    keeps the exception struct where 0.4.0 kept its message.
    {:tool_denied, tool} is {:tool_denied, tool, :tool_policy}, and a run's
    :authorize refusal is {:tool_denied, tool, reason}; Refine and
    Assertions return {:error, reason};
    Imp.optimize! raises Imp.Error for a failed optimization. The CHANGELOG
    lists every tag that changed.
  • Imp.Example and Imp.Prediction keep string keys as strings. Code that
    read a field of data loaded from JSON with map.field or map[:field]
    reads it with `Imp.Example.g...
Read more

Imp v0.4.0

Choose a tag to compare

@deepfates deepfates released this 17 Sep 21:07
ad64af1

Imp v0.4.0

Imp is a framework for typed, optimizable language-model programs on the BEAM.
Declare a task as named inputs and outputs, call it like any other Elixir
program, measure it on examples, compile it with an optimizer, and run the
selected program under OTP.

This release absorbs the protocol adapters that previously lived on main
only, and changes three published shapes. It is 0.4.0 rather than a patch
because a program written against v0.3.2 can need edits.

Install

v0.4.0 is a Git source release from a public repository; no credentials are
required.

{:imp, github: "deepfates/imp", tag: "v0.4.0"}

Imp is not published to Hex. Use a path dependency only while developing
against a local checkout.

ExMCP is declared runtime: false, so an OTP release that uses Imp.ACP or
Imp.MCP must list applications: [ex_mcp: :load] in its release
definition; see protocol runtime in
releases
.
Ordinary Imp startup starts no protocol endpoint.

Headline changes

  • Imp.ACP and Imp.MCP.connect/2 are in the tag. The separate imp_acp
    package is retired with no compatibility shim: a consumer that depended on
    it now depends on imp alone. Imp.MCP.connect/2 also gains OAuth
    credentials for remote HTTP servers (Imp.MCP.OAuth), bearer_env
    descriptor auth, on_failure: :drop with an unavailable list, and a
    per-dial timeout.
  • :reasoning_effort is the one reasoning option on Imp.Clients.ReqLLM.
    :openrouter_reasoning is gone; the wire encoding is the separate
    :openrouter_reasoning_wire.
  • ReActV2 sends the tool roster natively and no longer declares a tools
    input field or writes its instructions into signature.instructions. Loop
    guidance travels to the adapter through :adapter_opts.
  • The :model_response event's metadata.cost is a plain USD float or nil,
    with any provider breakdown under metadata.billing.
  • Structured values in a prompt render complete, the way DSPy renders a dict,
    instead of a truncated inspect/1.

Breaking changes from v0.3.2

  • Replace openrouter_reasoning: ... with reasoning_effort: .... Saved
    programs allowlist :reasoning_effort and :openrouter_reasoning_wire in
    its place, so rebuild artifacts that carried the old key.
  • A caller that passed or read ReActV2's tools input field no longer has
    one; the roster is sent natively.
  • A host that read metadata.cost as a provider billing map reads a number
    now, and finds the map under metadata.billing when the provider sent one.

Upgrade path

  1. Rename the reasoning option and rebuild saved artifacts with 0.4.0.
  2. Drop any tools handling around ReActV2.
  3. Sum spend from metadata.cost as a number.
  4. If you depended on {:imp, github: "deepfates/imp", branch: "main"} for the
    adapters, move to the tag.
  5. Run your held-out evaluation and application smoke test against the tagged
    dependency.

New in this release: Benchmarks
and its results table
carry every number this repository publishes with the command that produces it,
and the ticket-routing rows were re-measured live for this release, a month
after the first run, with both runs recorded.

The CHANGELOG records every user-visible change in this
release. Generated module documentation is the complete API reference. Start
with Imp, Imp.Signature, Imp.Module, Imp.Evaluate, Imp.Optimizer,
Imp.ACP, Imp.MCP, and Imp.Telemetry.

Imp v0.3.1

Choose a tag to compare

@deepfates deepfates released this 23 Aug 16:29

Private source patch release. Clarifies the supported center versus pre-1.0 advanced workflows, documents the exact input/output typing contract and structured field schema, corrects the restricted-interpreter security boundary, and fixes cold-reader prerequisites and cross-references.\n\nPackage SHA-256: d3de6bfb6a4bc04993ba3ec415fa1937ed6a546c260e88c6d619fadf731458fa

Imp v0.3.0

Choose a tag to compare

@deepfates deepfates released this 23 Aug 14:29

Imp v0.3.0 — private source release

Imp is DSPy for the BEAM: declare a language-model task as a typed Elixir
program, then test, measure, improve, and operate it like any other code.
Instead of maintaining prompt strings, you declare signatures — named, typed
inputs and outputs — and programs are ordinary Elixir values you can call,
evaluate against metrics, compile with optimizers, persist as checksummed
artifacts, and run under OTP supervision.

Install

Imp is not published to Hex. Install the private source release from its
immutable tag (GitHub credentials with repository access are required):

{:imp, github: "deepfates/imp", tag: "v0.3.0"}

Use a path dependency only for local development.

Three focused product guides, five Livebooks, a provider-free tutorial, and a
deployment example are bundled for HexDocs. Generated module documentation is
the exhaustive API reference. Research evidence and maintainer procedures
remain in the source repository.

This is a private Git source release tagged v0.3.0, not a Hex release.
Repository visibility remains private.

Breaking changes from 0.2.1

  • Imp.optimize/3, /4, and /5 return {:ok, program} or
    {:error, reason}. Use the corresponding Imp.optimize! arity to retain the
    previous raising behavior.
  • Imp.Adapters.Types and its nested structs moved to Imp.Adapter.Types.
    Replace the Imp.Adapters. prefix with Imp.Adapter..
  • The disconnected Imp.Agent/Imp.Agent.Runtime model was removed. Use the
    ordinary ReActV2 or RLM program path; hosts needing ordered events,
    cancellation, or explicit per-effect approval can use experimental
    Imp.start_run/3 without adopting a second agent abstraction.

What you are getting

Be precise about what kind of thing this release is, in three layers:

Proven here, with receipts you can run. A broad substantive DSPy 3.2.1
surface, realized natively: the signature DSL, program shapes from predict
through ReAct, CodeAct, and a sandboxed recursive controller, evaluation,
optimizers, retrieval, MCP, streaming, and persistence. Fidelity is a checked
per-surface claim, not a blanket promise that everything matches upstream. The
optimizer and adapter families carry executable differential tests that run
pinned DSPy 3.2.1 and compare their declared observations, backed by committed,
content-addressed evidence artifacts. The conformance report records a
differential, an Elixir-native equivalent, or an honest gap for each tracked
surface. The source repository's
evidence guide
and conformance report
retain those research records separately from the packaged manual.
Imp.Experiment.check/5 can repeat noisy selection independently from final
test estimation (repetitions: [selection: n, test: m]), selects on the
selection mean, preserves strict baseline ties, and persists the per-stage runs
and paired deltas. The original integer repetition form remains the uniform
shorthand.

Several scoped effectiveness results remain reviewable with source-repository
artifacts: the latest ticket-routing study
run improves from 30–35% to 95–100% on held-out data across three live repeats,
for about 1.3 cents per repeat. The older retained run recorded 25–30% to 85%;
it remains historical evidence rather than the current tutorial result. On a
separately frozen matched TREC contract, Imp GEPA
improved
its baseline by +0.4000 and cleared the preregistered noninferiority margin
against pinned DSPy GEPA; MIPROv2 improved its own baseline by +0.1458.
The committed compact scored-row inputs let a third party rerun the frozen
scoring and aggregation without publishing the 181 MB private provider traces.
The source repository's TREC case study gives the exact
source-checkout command, input hashes, and limitations. The rows and research
program remain source-repository evidence rather than Hex runtime contents.
Neither result is a general optimizer-effectiveness claim.

Borrowed honestly — with the gap named. The optimization algorithms'
effectiveness evidence comes from their published literature (DSPy, MIPROv2,
SIMBA, GEPA) — which verified those implementations on those tasks. Imp's
differentials verify their named mechanics; they do not establish whole-loop
parity. The matched TREC result answers one outcome question, while most
optimizers still lack matched live outcomes.
That transfer question is tracked as open targets in the ledger, not assumed
away. What we can say from our own committed evidence: demonstration-based
compilation produces real held-out lift (the tutorial's result), and whether
any optimizer improves your task is a question Imp lets you answer in an
afternoon — signature, metric, held-out split, receipt.

Promised, explicitly. Our own matched-control effectiveness science —
multi-seed optimizer studies, paper-scale reproductions, matched-model
Imp-vs-DSPy comparisons — is the open research program, tracked as unasserted
target claims in the same public ledger. We do not assert what we have not
measured, and the ledger is the boundary between the two.

What the BEAM adds

A model call is one more slow, fallible, concurrent effect: bounded supervised
evaluation fan-out, tools in isolated tasks under their own timeouts, scripted
deterministic testing with Imp.LM.Static through the same seams production
uses, compiled programs as checksummed artifacts with no secrets inside,
credentials bound at runtime, and redacted telemetry on every call, retry, and
tool step. The deployment example is a complete OTP
application.

The packaged
provider-free ticket router
is the cold-start product proof: a separate consumer compiles a typed router,
measures a deterministic 25% baseline, attaches four reviewable demonstrations,
and measures 100% after compilation. Its scripted LM proves the package and
program/evaluation/optimizer lifecycle, not real-model effectiveness.

The packaged OTP deployment example is the
release front door: an unpacked consumer runs a two-stage typed program through
disjoint selection and test splits, optimizer selection, checksummed result and
artifact persistence, fresh-process loading, concurrent serving, hot reload,
and contained failure using a provider-free support workflow. The source
repository separately retains real-model Banking77 research, including an
honest negative GEPA selection and successful fresh-service lifecycle; those
research data, runners, results, and artifacts are deliberately excluded from
the packaged example.

Persistent playbooks use the same application-owned deployment posture. Their
optimizer grounds proposals in training failures, requires lift on separate
promotion and audit rows, returns a challenger review with authoritative usage
and exact rollback, and restores a private checksummed checkpoint into freshly
constructed runtime code. It returns immutable state; installing the selected
playbook into a serving process remains an explicit application decision.

Since v0.1.0

  • Packaged for Hex with the product guides, five Livebooks, provider-free
    tutorial, and deployment example; research and internal audit material stay
    in the repository. Publication to hex.pm is still pending — until
    it happens, install from a source checkout.
  • Documentation rebuilt reader-first: the README, one Learning Path, a DSPy
    migration guide, and one operations guide form the reader path; generated
    module docs own exhaustive reference. The evidence ladder remains a
    repository audit surface rather than part of the packaged manual.
  • Evidence campaign: nine new differential artifact families landed; every
    semantic-conformance claim in the ledger is now asserted (was 1 of 4).
    "Asserted" is a maintainer attestation, not a fresh-checkout replay — see
    the source repository's
    evidence reconciliation
    for what a clean clone can verify from committed evidence alone.
  • Streaming promoted to the facade (Imp.stream/3, Imp.collect/3) with
    provider token streaming and an explicit local fallback. Provider mode runs
    the real composed program, can select intermediate fields by named predictor,
    and ends with its typed prediction. It is demand-driven and cleans up on an
    early halt or dead consumer. Programs with no named predictors return a
    terminal unsupported-program error instead of replaying a completed response
    as if it were provider output.
  • Known optimizer identities in Imp.Optimizer.Report retain their atom type
    after a checksummed Artifact write/read/apply cycle; unknown extension
    identifiers remain portable strings.
  • Two silent-failure bugs found and fixed the same day they were exposed by
    the claims census (batch message mangling; swallowed telemetry), plus
    loud-by-default evaluation timeouts and a threadable teacher timeout for
    BootstrapFewShot — that last one found live by our own dogfooding.
  • LabeledFewShot selection semantics pinned deterministic;
    Imp.Datasets.split/2 shuffle now seeded.
  • CI rebuilt: four parallel gates, deterministic dependency cache.

Expectations for 0.x

APIs may change before 1.0. Known flaky tests are ticketed and public in the
repository. The evidence ladder is the contract: if a page claims more than
its receipts support, that is a bug — file it.

The Imp facade and modules marked stable in priv/public_api.json are the
compatibility center. Optimizer and advanced modules are explicitly
experimental before 1.0: they are real implementations with public tests, but
their APIs may still converge with upstream semantics. A callable optimizer is
not thereby proven effective; the task-scoped results above are the evidence
boundary.

This release ...

Read more

Imp v0.1.0

Choose a tag to compare

@deepfates deepfates released this 17 Jul 00:33

Imp v0.1.0 is the first immutable public release of the Elixir/BEAM-native DSPy toolkit.

Highlights:

  • Typed programs, signatures, providers, evaluation, retrieval, tools/agents, RLM, optimizers, persistence, and observability
  • Credential-safe serialization and artifact handling with hostile regression coverage
  • Deterministic package, release, public-API, Livebook, protocol, and documentation gates
  • Source-bound evidence with explicit C0-C5 claim scopes and honest exclusions

Install from the immutable Git tag:

{:imp, github: "deepfates/imp", tag: "v0.1.0"}

Verified source: e7edc10a0dd0b493b3afddcd556869a370a7647e. Hex publication is not part of this release.