Grok backend, ChatGPT subscription sign-in, cloud voices
Three contributions from @raphaelbarreiros, plus the build fix for current Xcode.
Grok backend (#61)
xAI Grok is the sixth backend, over xAI's public Chat Completions API. Connect with an xAI API key or a SuperGrok subscription sign-in (OAuth, endpoints from auth.x.ai discovery with a trusted-host check). Photos from the glasses, web search, the native tools and conversation history all work, sharing the prompt and tool loop with OpenAI through the new CloudChat. The model picker loads your account's live list; the default is the grok-4.20-non-reasoning alias.
ChatGPT subscription sign-in for OpenAI (#60)
Settings → OpenAI now offers Connect With: API Key or ChatGPT Subscription. Sign in with your ChatGPT Plus/Pro account and use the subscription instead of buying API credits. OAuth with PKCE, tokens in the Keychain (this-device-only), single-flight rotation-safe refresh. Text, photos, web search and tools work; live video still needs an API key (or Gemini).
This uses the Codex CLI's integration surface, not a documented API, so it can change or stop working without notice. Signing in on the phone appears to end a Codex CLI session for the same account, and vice versa.
Cloud voices (#61)
Two speech engines next to Apple and Kokoro: Grok (28 multilingual voices, follows the reply's language; works with the SuperGrok sign-in) and OpenAI (gpt-4o-mini-tts, 13 voices; needs an API key with credits). Sentences are synthesized as they're written, at most three in flight, and played in order and gaplessly; the voice goes into session recordings like Kokoro's. Every neural voice list now has a play button per voice.
Build
mlx-swift-lmpin moved to upstream14414441, so the project builds on Xcode 27 (#59). Still requires the Metal Toolchain and-skipPackagePluginValidation, see SETUP.md.
Fixes
- Ambient watch-loop narration could reset a streaming reply's speaking state and let barge-in fire on the assistant's own voice (Kokoro and cloud engines).
- Completion callbacks from a stopped reply can no longer corrupt the buffer count of the reply that replaced it.
Known limitations
- OpenAI voice is not yet verified on a device (no contributor had API credits). It shares the Grok playback path; a confirmation from anyone with credits is welcome.
- A cloud-voice sentence that fails to synthesize is skipped rather than spoken by the Apple voice (#62).
Thanks @raphaelbarreiros for three well-tested PRs and a very fast review turnaround.
Full changelog: v2.13.0...v2.14.0