llm-driver is a small TypeScript library with one text-generation API and four
switchable targets. Point it at a hosted provider API or at an agent CLI you are
already logged into, and switch between them by changing configuration only —
the generate call site never changes. Models are always explicit; the library
never selects or updates one for you.
| Provider | Flavor | Target | Authentication |
|---|---|---|---|
claude |
api |
Anthropic Messages API | ANTHROPIC_API_KEY |
claude |
cli |
Local claude -p |
Existing Claude CLI login |
openai |
api |
OpenAI Responses API | OPENAI_API_KEY |
openai |
cli |
Local codex exec |
Existing Codex CLI login |
npm install llm-driverRequires Node.js 20.12 or later. Ships dual ESM + CommonJS builds with TypeScript declarations for both.
import { createClient, user } from "llm-driver";
const client = createClient({
provider: "openai",
flavor: "cli",
model: "gpt-5.6-sol",
});
const response = await client.generate({
system: "Answer concisely.",
messages: [user("What is dependency inversion?")],
maxTokens: 1024,
});
console.log(response.text);CommonJS works too:
const { createClient, user } = require("llm-driver");The response is provider-neutral:
interface Response {
id: string; // "" when the target reports none
model: string;
text: string;
usage: Usage; // inputTokens, outputTokens, cachedInputTokens,
// cacheCreationInputTokens, reasoningTokens
completionReason: "stop" | "max_tokens" | "refusal" | "";
provider: "claude" | "openai";
flavor: "api" | "cli";
}Only the config object changes:
// Claude API
createClient({ provider: "claude", flavor: "api", model: "claude-sonnet-4-5" });
// Local Claude CLI
createClient({ provider: "claude", flavor: "cli", model: "claude-sonnet-4-5" });
// OpenAI API
createClient({ provider: "openai", flavor: "api", model: "gpt-5.6-sol" });
// Local Codex CLI
createClient({ provider: "openai", flavor: "cli", model: "gpt-5.6-sol" });The request and the generate call stay identical across all four.
generateStream takes the same request and returns an async iterable of events:
zero or more text deltas, then exactly one final done event carrying the
same Response generate would have returned. Nothing follows done.
for await (const event of client.generateStream({
messages: [user("Explain dependency inversion")],
maxTokens: 1024,
})) {
if (event.type === "text") {
process.stdout.write(event.text);
} else {
console.log("\n", event.response.usage);
}
}type StreamEvent =
| { type: "text"; text: string } // incremental delta, possibly coarse
| { type: "done"; response: Response }; // always last, exactly onceThe concatenated text deltas equal done.response.text. Granularity
is not part of the contract — it is whatever the target reports, and a target
may report nothing until the end:
| Target | Granularity |
|---|---|
claude/api |
Token-level deltas (Messages API SSE) |
openai/api |
Token-level deltas (Responses API SSE) |
claude/cli |
Partial-message chunks (--output-format stream-json --include-partial-messages) |
openai/cli |
One coarse delta: codex exec --json reports completed messages only |
claude/clicaveat.claude -pis an agent, not a completion endpoint. It streams deltas for every assistant message in the turn, but its finalresultevent — the one that becomesdone.response— reports only the last message. The concatenation therefore equalsdone.response.textfor a single-message turn; if the CLI runs tools, the deltas additionally contain the intermediate assistant text spoken before each tool call.done.responseis always exactly whatgeneratewould have returned. Treatdone.response.textas the answer and the deltas as progress output. The other three targets hold the equality unconditionally.
Errors and aborts work exactly as with generate, except that they surface from
the iteration rather than from the call:
- Request validation happens on the first
next(), not whengenerateStreamis called — standard async-generator semantics.const it = client.generateStream(bad)does not throw; thefor awaitthat drives it does. - A failure throws an
LLMDriverErrorfrom the loop; an abort throws the signal's own reason untouched, identically across all four targets. - Stopping early cleans up the transport.
break,return, orthrowinside the loop aborts the HTTP stream, or signals the CLI process group (SIGTERM, then SIGKILL after a short grace period) — the iteration does not wait around to reap the process, it only guarantees the teardown is under way. - That cleanup runs in the generator's
finally, whichfor awaittriggers for you. A manual iterator must call.return()itself (or thefinallynever runs and the transport leaks). Note that.return()cannot preempt a pendingnext(): it is queued behind it, so a stalled stream is only preempted by theAbortSignal.
for await (const event of client.generateStream(request, {
signal: AbortSignal.timeout(30_000),
})) {
if (event.type === "text" && event.text.includes("STOP")) break; // transport torn down
}API flavors use the official provider SDKs. Set ANTHROPIC_API_KEY or
OPENAI_API_KEY in the environment, or pass apiKey explicitly:
const client = createClient({
provider: "claude",
flavor: "api",
model: "claude-sonnet-4-5",
apiKey: process.env.MY_KEY,
baseUrl: "https://proxy.internal/anthropic", // optional
fetch: myInstrumentedFetch, // optional transport override
});baseUrl and fetch are for proxies, custom transports, and tests. Treat
custom endpoints and transports as trusted application configuration: API
credentials and prompts are sent through them.
CLI flavors reuse the installed tool's existing login — no API key is read.
cliPath selects a nonstandard executable and cliArgs appends argv verbatim:
const client = createClient({
provider: "openai",
flavor: "cli",
model: "gpt-5.6-sol",
cliPath: "/opt/homebrew/bin/codex", // optional
cliArgs: ["--config", "model_reasoning_effort=high"], // optional
});Prompts are written to the subprocess's stdin and argv is built directly — no
shell is ever invoked. Treat cliPath and cliArgs as trusted application
configuration, because they control local process execution.
Options are flavor-scoped: passing apiKey, baseUrl, or fetch to a cli
client — or cliPath/cliArgs to an api client — is rejected with an
invalid_config error rather than silently ignored.
Every failure is an LLMDriverError carrying a stable code plus whatever
context the target reported.
import { createClient, LLMDriverError, user } from "llm-driver";
try {
const response = await client.generate({
messages: [user("hello")],
maxTokens: 256,
});
console.log(response.text);
} catch (error) {
if (error instanceof LLMDriverError) {
console.error(error.code, error.provider, error.flavor, error.status);
}
throw error;
}error.code |
Meaning |
|---|---|
invalid_config |
Bad provider/flavor/model, or an option used with the wrong flavor |
invalid_request |
Empty transcript, empty message text, bad role, or non-positive/non-integer maxTokens |
executable_not_found |
The CLI binary could not be launched |
process_failed |
The CLI exited non-zero (status is the exit code) |
parse_failed |
Target output could not be parsed |
api_error |
The provider or CLI reported a failure (status, providerCode when available) |
transport_failed |
Network or transport-level failure |
Other fields: provider, flavor, operation (e.g. "generate"), status
(HTTP status or process exit code), providerCode, and cause.
Abort is deliberately different. Pass an AbortSignal and, when it fires,
generate rejects with the signal's abort reason itself — never a wrapped
LLMDriverError — identically across all four targets:
const response = await client.generate(
{ messages: [user("hello")], maxTokens: 256 },
{ signal: AbortSignal.timeout(30_000) },
);There is no default timeout: without a signal, generate waits as long as the
target takes. Pass AbortSignal.timeout(ms) if you need a deadline.
The shared contract is text-only generation, streaming or not: system text, multi-turn user/assistant messages, normalized text, and usage when reported.
The CLI flavors intentionally wrap agent CLIs. They are not byte-for-byte equivalents of the hosted APIs:
- API flavors map and enforce
maxTokens. CLI flavors validate it as part of the portable request but have no reliable equivalent flag, so it is not enforced. - Claude CLI runs in single-shot print mode with the default permission mode, so a headless run cannot approve tool actions that need approval — but tools that are allowed by default still run (see the security note below). Codex runs with a read-only sandbox. Agent CLI behavior can still differ from a hosted model endpoint.
- System text is passed through Claude's
--append-system-promptflag and Codex's per-invocationdeveloper_instructionsconfig. Conversation text stays on stdin. System text may therefore be visible to local process inspection. - CLI subprocesses inherit the application's working directory and environment
so local authentication works. Do not send untrusted prompts to a CLI
flavor without isolating the host process. Codex's read-only sandbox
prevents writes, not reads.
claude -pruns in the default permission mode, which still lets the agent use its default-allowed read tools (Read,Glob,Grep) without prompting, in the inherited working directory — a prompt-injected run can read local files even though it cannot write them. Run the host process in a sandbox or a directory with nothing sensitive in it, and usecliArgsto tighten the CLI's own limits, e.g.cliArgs: ["--disallowed-tools", "Read,Glob,Grep"]. - Usage fields are populated only when a target reports them; everything else
is
0. - Process-group cleanup on abort is POSIX-only: the subprocess is spawned
detached and aborting signals the whole group (SIGTERM, then SIGKILL after a
grace period) so CLI-spawned helpers die too. On Windows only the direct child
is killed, and the CLI flavors are untested there —
.cmdshims are not resolved, so pass an explicitcliPath. - Being detached also means a terminal Ctrl-C does not reach the CLI child: the
signal goes to your process group, not its own. A live child is killed by an
explicit
AbortSignal, or by the library'sprocess.on("exit")handler when the host shuts down normally. A host killed outright (SIGKILL, or a signal it does not handle) leaves the CLI running until it finishes on its own.
Tool/function calling, images, structured output, retries, automatic fallback, and persisted conversations are outside this library's current scope.
The bundled example takes the target and model from command-line flags:
npm run example -- \
--provider openai \
--flavor cli \
--model gpt-5.6-sol \
--prompt "Explain dependency inversion in one paragraph"Optional flags: --system <text>, --max-tokens <n> (default 1024), and
--stream to print deltas as they arrive followed by the usage line.
Swap in --provider claude --flavor api --model claude-sonnet-4-5 and the
example's generation code is unchanged.
npm install
npm run build # tsup → dist/ (ESM + CJS + .d.ts)
npm test # vitest run
npm run test:coverage
npm run lint # biome check
npm run typecheck # tsc --noEmitThe default test suite is offline: it uses injected fetch/HTTP test servers and a fake process runner, never provider credentials or real agent CLIs.
That also means the CLI flavors' real flags (claude -p --output-format json …
and codex exec --json …) cannot be validated offline. An opt-in smoke test
launches the real binaries so a renamed or removed flag surfaces in the suite
instead of only at runtime. It skips unless the model env vars are set:
LLMWRAPPER_CLAUDE_CLI_MODEL=claude-sonnet-4-5 \
LLMWRAPPER_CODEX_CLI_MODEL=gpt-5.6-sol \
npm testMIT — see LICENSE.