Skip to content

Releases: fire17/apiplan

v0.7.1 — portable release artifact

Choose a tag to compare

@fire17 fire17 released this 31 Aug 07:30

Fixes

  • Makes source-contract checks newline-independent on Windows.
  • Isolates Google credential fingerprint probes in their own process, eliminating cross-file environment races.
  • Gives the client-startup floor enough spawn samples to measure deterministic work rather than scheduler noise.
  • Includes every v0.7 cache/provider upgrade from v0.7.0.

Local verification

  • 272 tests passed, 0 failed, 808 assertions.
  • All 7 performance budgets passed: 22 ms client startup, 3 ms owned dispatch/drain, 59 MB idle daemon.
  • apiplan doctor: all clear after synchronizing 37 global commands.

The attached tarball is the install-tested release artifact.

v0.7.0 — cached API standard and full provider upgrade

Choose a tag to compare

@fire17 fire17 released this 31 Aug 07:16

What changed

  • Makes provider-native cached prompts the permanent localhost API standard: Anthropic cache_control + metadata and OpenAI prompt_cache_key/session routing are preserved end to end.
  • Adds seamless apiplan hotswap upgrade drain/replace on port 8787; the release cutover held 40/40 continuity probes.
  • Folds in the complete provider upgrade: Google Antigravity/Gemini, Ollama, media and video vision, credential single-flight/rotation recovery, tool-call fidelity, evidence-based health, and truncated-stream detection.
  • Addresses every API-capable model in the live Codex catalog, including gpt-reserve and codex-auto-review, while excluding models marked unsupported.

Verified

  • 272 tests pass, 0 fail, 808 assertions after cross-platform probe isolation.
  • All 7 performance budgets pass: 23 ms client startup, 3 ms owned dispatch+drain, 56 MB idle daemon.
  • Live cache receipts: 16,226 Anthropic cache-read tokens and 4,864 OpenAI cached tokens.
  • Live provider matrix: Claude Opus 5, GPT-5.6-Sol, Gemini 3.7 Flash and local heretic all returned the exact requested result.
  • 37 global commands installed; doctor all clear after command sync.

v0.6.1 — voice bye-hangup + first-word latency

Choose a tag to compare

@fire17 fire17 released this 15 Aug 23:02

talk: bye/goodbye ends the call (--bye/--no-bye to tune); reports first word in Nms on first audio vs an LX_T0_MS stamp so launchers can measure call → speech. Powers the live-explain (/lx) skill.

v0.6.0 — point any SDK at localhost

Choose a tag to compare

@fire17 fire17 released this 06 Aug 22:50

apiplan serve runs a local server that speaks OpenAI's and Anthropic's wire shapes exactly. Change one base URL and existing code answers from the subscriptions you already pay for — no API key, no per-token bill. Still zero dependencies.

apiplan serve                 # http://127.0.0.1:8787
from openai import OpenAI
client = OpenAI(api_key="not-needed", base_url="http://127.0.0.1:8787/v1")
client.chat.completions.create(model="opus", messages=[...])   # Claude, in OpenAI's shape
export OPENAI_BASE_URL=http://127.0.0.1:8787/v1
export ANTHROPIC_BASE_URL=http://127.0.0.1:8787

The dialect and the backend are independent

The path decides the response shape. The model field decides who answers. So /v1/chat/completions with model: "opus" returns Claude in OpenAI's format, and /v1/messages with model: "sol" returns GPT in Anthropic's.

That is the point: most tooling speaks exactly one dialect, and this makes every model reachable from all of it.

endpoint shape
POST /v1/chat/completions OpenAI chat, streaming and not
POST /v1/messages Anthropic messages, streaming and not
POST /v1/audio/speech OpenAI speech — instructions steers delivery, like tts --as
POST /v1/images/generations OpenAI images, b64_json
GET /v1/models either shape; the caller's auth header picks which

Errors come back in the caller's own envelope, so both SDKs parse failures correctly.

Verified with the real SDKs

Not curl — the official openai and @anthropic-ai/sdk packages, installed fresh: both dialects, both directions (Claude through OpenAI's SDK, GPT through Anthropic's), streaming on both, plus audio.speech.create and images.generate.

Two bugs this surfaced, both fixed at the root

  • Anthropic replied with an empty string. build() leaves stream to the caller and the CLI sets it at call time; without it Anthropic returns a plain JSON body, which the SSE reader parsed as zero events. Silent, not an error.
  • --max-tokens was broken on every OpenAI model, and had been. The codex backend rejects max_output_tokens outright (400). Rarely hit from the CLI; every API client sets it by default, so the server hit it immediately. The parameter is no longer sent, and the CLI now says the flag is ignored there rather than dropping it quietly.

Safety

Binds 127.0.0.1 only — it hands out your subscription to anything that can reach it. Set APIPLAN_API_KEY to require a key (enforced on Authorization and x-api-key); --port / --host to move it.

134 tests · 7/7 budgets · CI green on ubuntu, macos and windows.

v0.5.0 — direct the performance, not just the words

Choose a tag to compare

@fire17 fire17 released this 02 Aug 15:57

tts can now be told how to say something, not just what. Still on your subscription, still no API key, still zero dependencies.

tts --as "excited, laughing" I cannot believe you actually pulled that off
tts --as "whispering, conspiratorial" --voice cedar keep this between us
tts --as "furious, shouting" the server has been down for three days
tts --as "a grizzled pirate captain, gravelly and theatrical" --voice ash land ho
tts --as-file director-notes.md the long monologue

--style, --emotion and --direction are synonyms. Free text — a character works as readily as an emotion, and so does an arc ("start deadpan, then crack up halfway"). Any language.

Why this was missing

v0.4.0 told the model "Read the following text aloud, verbatim, and say nothing else — no greeting, no comment." gpt-realtime is a conversational speech model, so that instruction was not neutral: it suppressed the expressiveness the model already had. The steering channel was open the whole time; we were holding it shut.

It measurably performs

One identical line, direction the only variable:

--as seconds loudness (RMS)
(none — baseline) 3.70 0.0997
excited, laughing out loud 4.65 0.1364
whispering, conspiratorial 3.75 0.0572
furious, shouting 3.85 0.1354
very slowly, heartbroken, tearful 9.15 0.0509

A whisper is 2.7× quieter than a laugh; a heartbroken read runs 2.5× longer than the baseline. Asked to laugh, it laughs — 3/3 runs put a real [laughs] into the model's own transcript, which is how this was verified objectively: duration, RMS amplitude, and the transcript's own record of non-speech vocalisations.

Whisper repeated 3/3 in a tight band (RMS 0.033–0.042, against 0.096–0.106 for the laughing takes).

Honest limit

Whether inline bracketed cues ([laughs] written mid-sentence) are performed rather than read aloud is consistent with the duration and transcript, but was not verified by ear. DARWIN.md round 20 records the whole method, including that gap.

Without --as, the strict verbatim path is unchanged — the direction is opt-in, and four new tests hold it from leaking into the spoken words.

122 tests · 7/7 budgets · CI green on ubuntu, macos and windows.

v0.4.0 — images and speech, on the subscription

Choose a tag to compare

@fire17 fire17 released this 02 Aug 15:38

Everything new here runs on the Claude Code / ChatGPT logins you already have. No API key is read, and nothing is billed per token. Still zero dependencies.

Draw

imagine a lighthouse keeper reading a letter by lamplight   # draws it, then opens it
imagine --raw a single red triangle on white                # your words, verbatim

Runs as the image_generation tool on the same Codex endpoint as chat.

By default the model rewrites your prompt before drawing — that's what prompt used: reports, and it genuinely helps a terse prompt. --raw pins your text character for character; --enhance asks for the rewrite explicitly.

Speak

tts the lighthouse keeper found a letter in the sand
tts --voice cedar שלום, זה מבחן קצר בעברית
apiplan voices

Ten voices — alloy ash ballad coral echo sage shimmer verse marin cedar — over OpenAI's realtime socket (gpt-realtime), which accepts the ChatGPT login and returns PCM16 that needs a 44-byte wav header and no codec. Any language; Hebrew verified.

The plan was a WebRTC stack. It turned out one retired header (OpenAI-Beta: realtime=v1) was the only thing standing between "needs ICE/DTLS/SRTP/Opus" and "a WebSocket". Dependency count stays at zero.

tts --aloud is a second, older engine: ChatGPT's product read-aloud in the app's own voices (maple juniper orbit fathom breeze ember glimmer vale cove). It can only speak a message already in your history, which is why it's a flag rather than a command.

Also

  • --open on images; imagine opens what it draws
  • apiplan voices shows every voice and where it comes from
  • apiplan install now tops up defaults added since your machine was set up — without this, imagine and tts would never have reached an existing install
  • README rebuilt: banner, architecture diagram, receipts

Honest limits

Speaking fresh text in a ChatGPT product voice isn't possible: it would mean creating a conversation first, and that endpoint sits behind a Cloudflare CAPTCHA (turnstile.required = true). Left uncleared. DARWIN.md rounds 14–19 log every route tried, including the ones that 404'd.

Windows users: v0.4.0 also fixes tests that couldn't locate their own source on Windows — caught by the CI matrix, invisible on macOS.

118 tests · 7/7 budgets · CI green on ubuntu, macos and windows.

apiplan v0.2.3 — verified on macOS, Linux, WSL and Windows

Choose a tag to compare

@fire17 fire17 released this 28 Jul 15:34
curl -fsSL https://raw.githubusercontent.com/fire17/apiplan/main/install.sh | sh

All four target operating systems now run this for real — not unit tests plus optimism.

verified how
macOS developed here: both providers live, images, 25-way parallel calls
Linux the published repo in a container — 88 tests, platform detection, unix-socket IPC, credential fallback to ~/.claude/.credentials.json, and the one-line install on a bare box with only git+curl
WSL a real WSL2 machine with no bun/claude/codex installed — the one-liner fetched bun, osLabel: WSL, 87 tests, aliases, bare sentence through the shim
Windows windows-latest on every push — 88 tests, alias resolution, the commands installed and invoked through the generated .cmd with an unquoted sentence, the .ps1 twin from PowerShell, and the loopback-TCP daemon coming up

What real Windows execution found (and fixed)

  • The repo was uncloneable on Windows. A stray tracked file whose name was a multi-line script fragment — legal on macOS/Linux, impossible on Windows — made git checkout abort before any test ran. Removed, plus a test that now fails if any tracked path contains a character Windows rejects.
  • removeShim() ignored extensionless shims on Windows, so every rename, delete and prune silently orphaned a file there.
  • Git Bash users had no working command at all — .cmd/.ps1 are resolved through PATHEXT, but bash appends only .exe. Windows now installs three files per command (.cmd, .ps1, and a bare sh shim), and PowerShell still resolves the .cmd — checked in CI, not assumed.

None of these were reachable from macOS by any amount of testing.

Also

  • CI badge and a per-platform table in the README stating how each platform was verified.
  • DARWIN.md now logs thirteen measure-fix-verify rounds, including the ones where the fix was to stop measuring the wrong thing.

88 tests · all performance budgets met · apiplan doctor all clear.

apiplan v0.2.2 — test portability fix

Choose a tag to compare

@fire17 fire17 released this 28 Jul 15:22

Patch release. No product changes.

Verifying the v0.2.1 tag from a fresh clone (rather than the tree it was built in) surfaced 2 of 87 tests failing there while passing locally: on macOS /var is a symlink to /private/var, so the installer's pwd and the test's import.meta.dir spell the same directory differently. The assertion now compares resolved paths.

Verified 87/87 in place and from a copy under /tmp, which is where the old assertion broke.

Install: curl -fsSL https://raw.githubusercontent.com/fire17/apiplan/main/install.sh | sh

apiplan v0.2.1 — one-line install, Linux verified

Choose a tag to compare

@fire17 fire17 released this 28 Jul 15:21
curl -fsSL https://raw.githubusercontent.com/fire17/apiplan/main/install.sh | sh

That's now the whole setup on a new machine.

What's new

  • One-line install. install.sh / install.ps1 self-bootstrap: piped from curl/irm with no checkout present they fetch themselves to ~/.apiplan/src and continue. Verified end-to-end on a bare Linux container with only git + curl.
  • apiplan update — pull the install, re-sync every command, refresh model lists. One word to stay current.
  • Linux is now live-verified. The published repo run in a container (aarch64): 80/80 tests, correct platform detection, unix-socket IPC, and the ~/.claude/.credentials.json credential fallback that only ever executes off-macOS.
  • Fixed a bug that silently wired commands to the wrong tree. sh install.sh puts no slash in $0, so the checkout test failed, the installer concluded it was being piped, and it cloned a second copy — commands then ran a stale tree. Source resolution now covers all three invocation styles, the installer prints the tree it wired, and apiplan doctor reports the install root. Three regression tests assert no clone appears.
  • apiplan path --raw for scripts; no PATH suggestion when the directory is already on PATH; an already-wired shell rc is left alone instead of appended to twice.
  • WSL detection is now testable — detectWsl() is a pure function checked against real WSL1/WSL2 kernel strings and the interop environment variables.
  • The perf gate no longer lies. It was failing ~half of identical runs: percentage bands are meaningless at 18 ms, one metric measured the OS Keychain rather than this code, and a small-sample median is the wrong estimator for deterministic work. Fixed and verified over four consecutive clean runs.

Verified

87 tests · all budgets met · apiplan doctor all clear · macOS and Linux live-verified.

WSL and Windows still have not been run — the logic is tested (kernel strings, .cmd/.ps1 shims, loopback-TCP daemon incl. its 403) but no process has executed there. Running it on either and reporting back is the most useful contribution right now.

DARWIN.md logs all ten rounds, including the ones where the fix was to stop measuring the wrong thing.

apiplan v0.2.0 — every frontier model as a shell command

Choose a tag to compare

@fire17 fire17 released this 28 Jul 14:56

Call frontier models from your shell using the Claude Code and Codex subscriptions you already have — no API key, no per-token bill.

opus explain monads in one sentence
cat server.log | sonnet find the root cause | pbcopy
sol -i screenshot.png what is wrong with this layout

Highlights

  • One command per model — opus sonnet fable haiku sol luna terra, plus *-fast twins. No quotes needed around your question.
  • Family names age well. opus means the newest Opus (today claude-opus-5); explicit versions stay reachable (opus48, sonnet46, gpt55, and opus4.8/Opus-4-8 all resolve). Model lists come from each provider's own endpoint, so a new release needs no code change.
  • Two providers, one engine. Anthropic /v1/messages and OpenAI's Codex Responses endpoint behind a single adapter interface. Every alias is verified against the API's own served-model field.
  • Multimodal + pipes. -i accepts a file, URL, data: URI, stdin, or your clipboard; stdin appends to the prompt; output pipes onward.
  • Warm daemon. Caches the credential and holds the provider connection open. Our own overhead on a warm call is 4 ms, measured inside the client — and a cold call never waits for the daemon.
  • apiplan control CLI + TUI. See which providers you're connected to, which models exist and what reaches them, and every command on your PATH — then rename any of them or invent new ones with baked-in flags. Commands are plain JSON, not code.
  • Cross-platform by construction. One file owns every macOS/Linux/WSL/Windows difference: credential location, sh vs .cmd+.ps1 shims, unix socket vs token-protected loopback TCP, clipboard mechanism.

Verified

80 tests · 7 of 7 performance budgets · apiplan doctor all clear · 25 parallel calls 25/25 (p50 1.06 s) · tests pass from a fresh clone of this tag.

Live-verified on macOS. Linux, WSL and Windows are unit-tested (including the Windows transport forced on macOS) but have not yet been run on those systems — see the README's cross-platform section.

Reading

DARWIN.md logs five measure-fix-verify rounds, including the round where the "warm" daemon turned out to be making every call 1.6 s slower, an optimisation that was measured and rejected, and the perf budget that was found to be measuring network weather instead of code.

MIT.