Repository navigation
Would Codex-backed live dictation for web and desktop be in scope? #12212
Replies: 5 comments
|
Here is an 80-second recording of the working local prototype: live dictation, moving the caret and editing while the microphone stays active, then optional AI polish after stopping. This is real screen capture with audio. It shows the current rough edges as well as the intended workflow. Polish is included to illustrate the prototype, but could remain outside the smaller first contribution proposed above. dictation-real-demo-trimmed.mp4 |
|
@farhang103 codex-cli just added an experiential feature to use |
|
#12275 was closed pending approval here, which is fair. I won't reopen it or push more until there's a decision. To make this a quick yes or no, here's the first PR I'd open if you approve:
Not included: AI polish, custom words and snippets, spoken punctuation and formatting commands, a keybinding, and mobile changes. Each of those would need its own discussion. This addresses the implementation concerns from #11278. Known constraint: CLI realtime is still experimental, and I haven't verified Safari or remote/tunnel WebRTC. I'd verify those before asking for review. Does this scope work, or should web/desktop dictation stay out of scope? Either answer helps. |
|
I really want this feature. On desktop we can use an external service, but there is nothing decent on mobile. |
|
Hi! Adding a humble +1 for native voice input on desktop, for both macOS and Windows. It would be wonderful to have a microphone in the desktop composer, a kind of built-in Whisper, with live transcription where the words appear as you speak, similar to how Claude's voice mode feels. For steering agents, talking is often so much faster than typing. If it's something you'd consider, a few ideas that I think would serve the community well:
I saw there's already great community work in #14882 and #14865. Whenever you have a moment to share which direction feels right, I think a lot of people would be happy to help bring it home. I'd gladly test on both Mac and Windows. Thank you for considering it. Together with close-to-tray, this would make T3 Code pretty much everything I need. ✍️ Written by Claude (Opus 5.5) on behalf of @luigisoares, who brought the idea and reviewed it. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I have been using a local Codex dictation prototype and would like to ask whether a revised approach would be useful before submitting a large PR.
The workflow I want is to explain a coding task by voice, correct a sentence with the keyboard, move the cursor, and continue speaking in the same draft. Integrated dictation keeps that workflow inside T3 without switching apps or copying text between tools.
I saw the scope decision on #11278 and understand that web/desktop voice was not something you wanted to add at that time. This prototype builds on Mina Sayed's work from that PR; I do not want to reopen the same proposal without checking first.
The revised approach uses the selected environment and Codex provider instance, and delegates account handling and realtime setup to the Codex CLI app-server instead of reading auth.json or impersonating Codex Desktop. Speech arrives over WebRTC and is inserted into the editable composer at the caret. Moving the caret or typing during dictation starts a new insertion segment. Spoken punctuation and basic formatting run locally; optional grammar polish runs after Stop and does not block typing or sending.
For a first contribution, I would propose only the minimum live dictation path: microphone start/stop, editable text at the caret, and routing through the selected environment and Codex account. Custom words/snippets, spoken formatting shortcuts, and optional AI polish could stay out of that first PR and be considered separately only if wanted.
The working local prototype includes those extras and a small microphone menu. Its full implementation is currently too large for your preferred contribution size (about 4,400 added lines including tests); that is the prototype size, not the scope I am proposing for an initial contribution.
There is still work before this is review-ready: the web session logic remains separate from the shared VoiceInputController, desktop/Safari behavior needs verification, and the unsupported-provider UI still needs to follow your previous feedback. The CLI realtime support is experimental, so compatibility is another constraint to resolve.
Would you consider a smaller implementation of this approach, or should web/desktop dictation remain out of scope?
Prototype changes prepared with GPT-6-Astra in the Codex harness via T3 Code; original implementation by Mina Sayed (#11278).
All reactions