Releases: brancusi/voice-tools-releases
Release list
Voice Pipes 1.10.0
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.10.0
- Choose which microphone to use. Setup → Microphone → Input: System follows whatever macOS uses, or pick a mic by name, such as your MacBook's.
- Each track can use its own mic. Open a track's Microphone block and pick one, for example your headset for a track where you want the clearer audio, and the laptop mic for quick dictation. Every mic your tracks use is kept ready, so takes start the moment you press. Bluetooth headsets are the exception: keeping one open switches it to low-quality call audio, so tracks on a headset start a moment after you press.
- If a mic you picked isn't plugged in, its tracks record from the system input until it's back, and Setup says so.
vp inputslists your mics for agents, and config.toml takesinput = "…"under[settings]or in a track's microphone block. - Training a word now records from the same mic you dictate with.
1.9.0
- History keeps everything, for good. Every dictation, answer and read-aloud is kept, not just the last 1,000, in a small database on your Mac. Saving a run no longer rewrites the whole history file, and the app holds only your newest runs in memory: the History window loads older ones as you scroll and searches all of them. A search across 50,000 runs takes a few hundredths of a second.
- Your agents can read all of it.
vp historyreaches back to your first run, whether the app is open or not:--search,--track,--since 3d,--all, andvp history exportgives whole runs (text, what was heard, every step, cost) as JSON lines. - Your existing history moves over by itself on first launch, and the old
history.jsonstays where it was as a backup.
1.8.5
- The version is in the menu bar panel too, on its last line: up to date with Check now, or the new version with Install….
1.8.4
- See which version you're running at a glance. The main window has a bar along the bottom with the version and whether it's up to date. Check now looks for a new version straight away; when there is one, it says so and Install… installs it (a few seconds, then the app relaunches).
1.8.3
- Hotkeys start recording instantly, and your first word is no longer cut off. Before, the microphone switched on only when you pressed, which takes up to half a second on some mics (about 0.45 s on a Studio Display), and anything you said in that gap was lost. Voice Pipes now keeps the microphone ready, so a take starts the moment you press and includes the half second before it. macOS shows its orange mic dot while the mic is open. The audio stays in memory until you press, and is never saved.
- Choose how ready it stays in Setup → Microphone: Always (the default), After use (open for 5 minutes after each take), or Off (the old behaviour). Or set
microphoneunder[settings]in config.toml. Bluetooth headsets are never kept open, because that would switch them into low-quality call audio.
1.8.2
- Every cloud voice now plays, whatever audio it sends back. Some voices, like Google's Gemini TTS (one voice that speaks many languages), only send raw audio, and Voice Pipes stopped with an error instead of speaking. It now plays raw audio as well as MP3, WAV, AAC and the rest. A voice that can't send MP3 is asked once for raw audio, and the app remembers that for next time.
1.8.1
- New defaults: hands on the keyboard, ears on what matters. While something is read aloud, the HUD now takes the keys straight away (Esc stops, Space pauses, j/k and h/l steer; the app you're in stays in front), and agents read aloud anything that needs your attention: questions they're waiting on, finished tasks, problems, and long replies. If you'd kept the old defaults, you now have these; anything you'd changed yourself stays as it was.
- Change it by asking. Tell your agent "don't take my keys", "only read me long stuff" or "stop reading to me" and it updates the setting for you, or do it yourself:
vp reading keys hover|click|never|always,vp agents read-aloud off|long|attention|all, andvp reading key faster periodfor your own keys. Setup → Reading has the same. - Model pickers: free tiers no longer say "free" twice in their name.
1.8.0
- Model pickers now show what a paragraph costs, how fast each model is on your Mac, and a quality rating from Artificial Analysis. Sort by best value, quality, speed or cost; each picker remembers your choice. On-device models stay on top in every sort: they're free, private and work offline. Costs are only shown where the billing unit is known (otherwise "—"); speeds come from your own runs (on-device models show our benchmark until then); quality comes from Artificial Analysis's public leaderboards as of October 2026, and models they don't cover say "not rated". Click a model for what it's good at, then Use it; ↑↓, →, Return and Esc work too.
- The sidebar is a little wider, so track names read in full (hover for the whole name), and ↑↓ move through it again.
- Agents read to me unasked in Setup → Command line and agents: Off, Long replies, When they need me, or Everything, the same setting as
vp agents read-aloud. - The read-along card fades its text at the edges instead of cutting a line in half.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to...
Voice Pipes 1.9.0
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.9.0
- History keeps everything, for good. Every dictation, answer and read-aloud is kept, not just the last 1,000, in a small database on your Mac. Saving a run no longer rewrites the whole history file, and the app holds only your newest runs in memory: the History window loads older ones as you scroll and searches all of them. A search across 50,000 runs takes a few hundredths of a second.
- Your agents can read all of it.
vp historyreaches back to your first run, whether the app is open or not:--search,--track,--since 3d,--all, andvp history exportgives whole runs (text, what was heard, every step, cost) as JSON lines. - Your existing history moves over by itself on first launch, and the old
history.jsonstays where it was as a backup.
1.8.5
- The version is in the menu bar panel too, on its last line: up to date with Check now, or the new version with Install….
1.8.4
- See which version you're running at a glance. The main window has a bar along the bottom with the version and whether it's up to date. Check now looks for a new version straight away; when there is one, it says so and Install… installs it (a few seconds, then the app relaunches).
1.8.3
- Hotkeys start recording instantly, and your first word is no longer cut off. Before, the microphone switched on only when you pressed, which takes up to half a second on some mics (about 0.45 s on a Studio Display), and anything you said in that gap was lost. Voice Pipes now keeps the microphone ready, so a take starts the moment you press and includes the half second before it. macOS shows its orange mic dot while the mic is open. The audio stays in memory until you press, and is never saved.
- Choose how ready it stays in Setup → Microphone: Always (the default), After use (open for 5 minutes after each take), or Off (the old behaviour). Or set
microphoneunder[settings]in config.toml. Bluetooth headsets are never kept open, because that would switch them into low-quality call audio.
1.8.2
- Every cloud voice now plays, whatever audio it sends back. Some voices, like Google's Gemini TTS (one voice that speaks many languages), only send raw audio, and Voice Pipes stopped with an error instead of speaking. It now plays raw audio as well as MP3, WAV, AAC and the rest. A voice that can't send MP3 is asked once for raw audio, and the app remembers that for next time.
1.8.1
- New defaults: hands on the keyboard, ears on what matters. While something is read aloud, the HUD now takes the keys straight away (Esc stops, Space pauses, j/k and h/l steer; the app you're in stays in front), and agents read aloud anything that needs your attention: questions they're waiting on, finished tasks, problems, and long replies. If you'd kept the old defaults, you now have these; anything you'd changed yourself stays as it was.
- Change it by asking. Tell your agent "don't take my keys", "only read me long stuff" or "stop reading to me" and it updates the setting for you, or do it yourself:
vp reading keys hover|click|never|always,vp agents read-aloud off|long|attention|all, andvp reading key faster periodfor your own keys. Setup → Reading has the same. - Model pickers: free tiers no longer say "free" twice in their name.
1.8.0
- Model pickers now show what a paragraph costs, how fast each model is on your Mac, and a quality rating from Artificial Analysis. Sort by best value, quality, speed or cost; each picker remembers your choice. On-device models stay on top in every sort: they're free, private and work offline. Costs are only shown where the billing unit is known (otherwise "—"); speeds come from your own runs (on-device models show our benchmark until then); quality comes from Artificial Analysis's public leaderboards as of October 2026, and models they don't cover say "not rated". Click a model for what it's good at, then Use it; ↑↓, →, Return and Esc work too.
- The sidebar is a little wider, so track names read in full (hover for the whole name), and ↑↓ move through it again.
- Agents read to me unasked in Setup → Command line and agents: Off, Long replies, When they need me, or Everything, the same setting as
vp agents read-aloud. - The read-along card fades its text at the edges instead of cutting a line in half.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another age...
Voice Pipes 1.8.5
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.8.5
- The version is in the menu bar panel too, on its last line: up to date with Check now, or the new version with Install….
1.8.4
- See which version you're running at a glance. The main window has a bar along the bottom with the version and whether it's up to date. Check now looks for a new version straight away; when there is one, it says so and Install… installs it (a few seconds, then the app relaunches).
1.8.3
- Hotkeys start recording instantly, and your first word is no longer cut off. Before, the microphone switched on only when you pressed, which takes up to half a second on some mics (about 0.45 s on a Studio Display), and anything you said in that gap was lost. Voice Pipes now keeps the microphone ready, so a take starts the moment you press and includes the half second before it. macOS shows its orange mic dot while the mic is open. The audio stays in memory until you press, and is never saved.
- Choose how ready it stays in Setup → Microphone: Always (the default), After use (open for 5 minutes after each take), or Off (the old behaviour). Or set
microphoneunder[settings]in config.toml. Bluetooth headsets are never kept open, because that would switch them into low-quality call audio.
1.8.2
- Every cloud voice now plays, whatever audio it sends back. Some voices, like Google's Gemini TTS (one voice that speaks many languages), only send raw audio, and Voice Pipes stopped with an error instead of speaking. It now plays raw audio as well as MP3, WAV, AAC and the rest. A voice that can't send MP3 is asked once for raw audio, and the app remembers that for next time.
1.8.1
- New defaults: hands on the keyboard, ears on what matters. While something is read aloud, the HUD now takes the keys straight away (Esc stops, Space pauses, j/k and h/l steer; the app you're in stays in front), and agents read aloud anything that needs your attention: questions they're waiting on, finished tasks, problems, and long replies. If you'd kept the old defaults, you now have these; anything you'd changed yourself stays as it was.
- Change it by asking. Tell your agent "don't take my keys", "only read me long stuff" or "stop reading to me" and it updates the setting for you, or do it yourself:
vp reading keys hover|click|never|always,vp agents read-aloud off|long|attention|all, andvp reading key faster periodfor your own keys. Setup → Reading has the same. - Model pickers: free tiers no longer say "free" twice in their name.
1.8.0
- Model pickers now show what a paragraph costs, how fast each model is on your Mac, and a quality rating from Artificial Analysis. Sort by best value, quality, speed or cost; each picker remembers your choice. On-device models stay on top in every sort: they're free, private and work offline. Costs are only shown where the billing unit is known (otherwise "—"); speeds come from your own runs (on-device models show our benchmark until then); quality comes from Artificial Analysis's public leaderboards as of October 2026, and models they don't cover say "not rated". Click a model for what it's good at, then Use it; ↑↓, →, Return and Esc work too.
- The sidebar is a little wider, so track names read in full (hover for the whole name), and ↑↓ move through it again.
- Agents read to me unasked in Setup → Command line and agents: Off, Long replies, When they need me, or Everything, the same setting as
vp agents read-aloud. - The read-along card fades its text at the edges instead of cutting a line in half.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another agent to read you its summary: the skill now tells it to write for listening, open the follow-along card and read it with Voice Pipes, then wait while you listen and steer.
1.6.4
- Follow along with anything read aloud. The HUD has a new button while it's reading: it opens a card above it with the whole text, one sentence per line, the one being read lit up and the card scrolling with the voice. Hover over it to look ahead; it picks up the voice again when you move away. Handy for long answers you didn't select yourself.
- Speed it up or slow it down as it reads. − and + on the card change the speed on the spot, from 0.6× to 2×, with every kind of voice. The card stays open for the next reading until you close it.
vp speed 1.4changes the speed from a t...
Voice Pipes 1.8.4
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.8.4
- See which version you're running at a glance. The main window has a bar along the bottom with the version and whether it's up to date. Check now looks for a new version straight away; when there is one, it says so and Install… installs it (a few seconds, then the app relaunches).
1.8.3
- Hotkeys start recording instantly, and your first word is no longer cut off. Before, the microphone switched on only when you pressed, which takes up to half a second on some mics (about 0.45 s on a Studio Display), and anything you said in that gap was lost. Voice Pipes now keeps the microphone ready, so a take starts the moment you press and includes the half second before it. macOS shows its orange mic dot while the mic is open. The audio stays in memory until you press, and is never saved.
- Choose how ready it stays in Setup → Microphone: Always (the default), After use (open for 5 minutes after each take), or Off (the old behaviour). Or set
microphoneunder[settings]in config.toml. Bluetooth headsets are never kept open, because that would switch them into low-quality call audio.
1.8.2
- Every cloud voice now plays, whatever audio it sends back. Some voices, like Google's Gemini TTS (one voice that speaks many languages), only send raw audio, and Voice Pipes stopped with an error instead of speaking. It now plays raw audio as well as MP3, WAV, AAC and the rest. A voice that can't send MP3 is asked once for raw audio, and the app remembers that for next time.
1.8.1
- New defaults: hands on the keyboard, ears on what matters. While something is read aloud, the HUD now takes the keys straight away (Esc stops, Space pauses, j/k and h/l steer; the app you're in stays in front), and agents read aloud anything that needs your attention: questions they're waiting on, finished tasks, problems, and long replies. If you'd kept the old defaults, you now have these; anything you'd changed yourself stays as it was.
- Change it by asking. Tell your agent "don't take my keys", "only read me long stuff" or "stop reading to me" and it updates the setting for you, or do it yourself:
vp reading keys hover|click|never|always,vp agents read-aloud off|long|attention|all, andvp reading key faster periodfor your own keys. Setup → Reading has the same. - Model pickers: free tiers no longer say "free" twice in their name.
1.8.0
- Model pickers now show what a paragraph costs, how fast each model is on your Mac, and a quality rating from Artificial Analysis. Sort by best value, quality, speed or cost; each picker remembers your choice. On-device models stay on top in every sort: they're free, private and work offline. Costs are only shown where the billing unit is known (otherwise "—"); speeds come from your own runs (on-device models show our benchmark until then); quality comes from Artificial Analysis's public leaderboards as of October 2026, and models they don't cover say "not rated". Click a model for what it's good at, then Use it; ↑↓, →, Return and Esc work too.
- The sidebar is a little wider, so track names read in full (hover for the whole name), and ↑↓ move through it again.
- Agents read to me unasked in Setup → Command line and agents: Off, Long replies, When they need me, or Everything, the same setting as
vp agents read-aloud. - The read-along card fades its text at the edges instead of cutting a line in half.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another agent to read you its summary: the skill now tells it to write for listening, open the follow-along card and read it with Voice Pipes, then wait while you listen and steer.
1.6.4
- Follow along with anything read aloud. The HUD has a new button while it's reading: it opens a card above it with the whole text, one sentence per line, the one being read lit up and the card scrolling with the voice. Hover over it to look ahead; it picks up the voice again when you move away. Handy for long answers you didn't select yourself.
- Speed it up or slow it down as it reads. − and + on the card change the speed on the spot, from 0.6× to 2×, with every kind of voice. The card stays open for the next reading until you close it.
vp speed 1.4changes the speed from a terminal;vp open reading/vp close readingshow or hide the text.vp uino longer reports a focused field after you've left the page it w...
Voice Pipes 1.8.3
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.8.3
- Hotkeys start recording instantly, and your first word is no longer cut off. Before, the microphone switched on only when you pressed, which takes up to half a second on some mics (about 0.45 s on a Studio Display), and anything you said in that gap was lost. Voice Pipes now keeps the microphone ready, so a take starts the moment you press and includes the half second before it. macOS shows its orange mic dot while the mic is open. The audio stays in memory until you press, and is never saved.
- Choose how ready it stays in Setup → Microphone: Always (the default), After use (open for 5 minutes after each take), or Off (the old behaviour). Or set
microphoneunder[settings]in config.toml. Bluetooth headsets are never kept open, because that would switch them into low-quality call audio.
1.8.2
- Every cloud voice now plays, whatever audio it sends back. Some voices, like Google's Gemini TTS (one voice that speaks many languages), only send raw audio, and Voice Pipes stopped with an error instead of speaking. It now plays raw audio as well as MP3, WAV, AAC and the rest. A voice that can't send MP3 is asked once for raw audio, and the app remembers that for next time.
1.8.1
- New defaults: hands on the keyboard, ears on what matters. While something is read aloud, the HUD now takes the keys straight away (Esc stops, Space pauses, j/k and h/l steer; the app you're in stays in front), and agents read aloud anything that needs your attention: questions they're waiting on, finished tasks, problems, and long replies. If you'd kept the old defaults, you now have these; anything you'd changed yourself stays as it was.
- Change it by asking. Tell your agent "don't take my keys", "only read me long stuff" or "stop reading to me" and it updates the setting for you, or do it yourself:
vp reading keys hover|click|never|always,vp agents read-aloud off|long|attention|all, andvp reading key faster periodfor your own keys. Setup → Reading has the same. - Model pickers: free tiers no longer say "free" twice in their name.
1.8.0
- Model pickers now show what a paragraph costs, how fast each model is on your Mac, and a quality rating from Artificial Analysis. Sort by best value, quality, speed or cost; each picker remembers your choice. On-device models stay on top in every sort: they're free, private and work offline. Costs are only shown where the billing unit is known (otherwise "—"); speeds come from your own runs (on-device models show our benchmark until then); quality comes from Artificial Analysis's public leaderboards as of October 2026, and models they don't cover say "not rated". Click a model for what it's good at, then Use it; ↑↓, →, Return and Esc work too.
- The sidebar is a little wider, so track names read in full (hover for the whole name), and ↑↓ move through it again.
- Agents read to me unasked in Setup → Command line and agents: Off, Long replies, When they need me, or Everything, the same setting as
vp agents read-aloud. - The read-along card fades its text at the edges instead of cutting a line in half.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another agent to read you its summary: the skill now tells it to write for listening, open the follow-along card and read it with Voice Pipes, then wait while you listen and steer.
1.6.4
- Follow along with anything read aloud. The HUD has a new button while it's reading: it opens a card above it with the whole text, one sentence per line, the one being read lit up and the card scrolling with the voice. Hover over it to look ahead; it picks up the voice again when you move away. Handy for long answers you didn't select yourself.
- Speed it up or slow it down as it reads. − and + on the card change the speed on the spot, from 0.6× to 2×, with every kind of voice. The card stays open for the next reading until you close it.
vp speed 1.4changes the speed from a terminal;vp open reading/vp close readingshow or hide the text.vp uino longer reports a focused field after you've left the page it was on.
1.6.3
- The command line keeps up by itself.
vpalways runs the app's own version, so updating from the menu updates it too. Now the rest follows as well: the agent skill is refreshed at every launch, an agent you install later (say Codex) gets it automatically, and the Claude Code sessi...
Voice Pipes 1.8.2
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.8.2
- Every cloud voice now plays, whatever audio it sends back. Some voices, like Google's Gemini TTS (one voice that speaks many languages), only send raw audio, and Voice Pipes stopped with an error instead of speaking. It now plays raw audio as well as MP3, WAV, AAC and the rest. A voice that can't send MP3 is asked once for raw audio, and the app remembers that for next time.
1.8.1
- New defaults: hands on the keyboard, ears on what matters. While something is read aloud, the HUD now takes the keys straight away (Esc stops, Space pauses, j/k and h/l steer; the app you're in stays in front), and agents read aloud anything that needs your attention: questions they're waiting on, finished tasks, problems, and long replies. If you'd kept the old defaults, you now have these; anything you'd changed yourself stays as it was.
- Change it by asking. Tell your agent "don't take my keys", "only read me long stuff" or "stop reading to me" and it updates the setting for you, or do it yourself:
vp reading keys hover|click|never|always,vp agents read-aloud off|long|attention|all, andvp reading key faster periodfor your own keys. Setup → Reading has the same. - Model pickers: free tiers no longer say "free" twice in their name.
1.8.0
- Model pickers now show what a paragraph costs, how fast each model is on your Mac, and a quality rating from Artificial Analysis. Sort by best value, quality, speed or cost; each picker remembers your choice. On-device models stay on top in every sort: they're free, private and work offline. Costs are only shown where the billing unit is known (otherwise "—"); speeds come from your own runs (on-device models show our benchmark until then); quality comes from Artificial Analysis's public leaderboards as of October 2026, and models they don't cover say "not rated". Click a model for what it's good at, then Use it; ↑↓, →, Return and Esc work too.
- The sidebar is a little wider, so track names read in full (hover for the whole name), and ↑↓ move through it again.
- Agents read to me unasked in Setup → Command line and agents: Off, Long replies, When they need me, or Everything, the same setting as
vp agents read-aloud. - The read-along card fades its text at the edges instead of cutting a line in half.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another agent to read you its summary: the skill now tells it to write for listening, open the follow-along card and read it with Voice Pipes, then wait while you listen and steer.
1.6.4
- Follow along with anything read aloud. The HUD has a new button while it's reading: it opens a card above it with the whole text, one sentence per line, the one being read lit up and the card scrolling with the voice. Hover over it to look ahead; it picks up the voice again when you move away. Handy for long answers you didn't select yourself.
- Speed it up or slow it down as it reads. − and + on the card change the speed on the spot, from 0.6× to 2×, with every kind of voice. The card stays open for the next reading until you close it.
vp speed 1.4changes the speed from a terminal;vp open reading/vp close readingshow or hide the text.vp uino longer reports a focused field after you've left the page it was on.
1.6.3
- The command line keeps up by itself.
vpalways runs the app's own version, so updating from the menu updates it too. Now the rest follows as well: the agent skill is refreshed at every launch, an agent you install later (say Codex) gets it automatically, and the Claude Code session hook follows the app. If you move Voice Pipes to another folder,vpis repointed at it; if that needs your password, Setup → Checks shows "vp points at a moved app" with a Fix… button.
1.6.2
- Install everything from a terminal. One line installs the app, the
vpcommand and the agent skill, and starts Voice Pipes:curl -fsSL https://github.com/brancusi/voice-tools-releases/releases/latest/download/install.sh | bash. It asks nothing, so an agent can run it too, and it refuses any download that isn't signed by Voice Pipes' developer and notarized by Apple. Run it again to repair or reinstall;--uninstallremoves the app, itsvplinks and the skill, and keeps your config, history and keys.
1.6.1
- Agents can show you, not just tell you.
vp opennow reaches anything in the app: a track ...
Voice Pipes 1.8.1
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.8.1
- New defaults: hands on the keyboard, ears on what matters. While something is read aloud, the HUD now takes the keys straight away (Esc stops, Space pauses, j/k and h/l steer; the app you're in stays in front), and agents read aloud anything that needs your attention: questions they're waiting on, finished tasks, problems, and long replies. If you'd kept the old defaults, you now have these; anything you'd changed yourself stays as it was.
- Change it by asking. Tell your agent "don't take my keys", "only read me long stuff" or "stop reading to me" and it updates the setting for you, or do it yourself:
vp reading keys hover|click|never|always,vp agents read-aloud off|long|attention|all, andvp reading key faster periodfor your own keys. Setup → Reading has the same. - Model pickers: free tiers no longer say "free" twice in their name.
1.8.0
- Model pickers now show what a paragraph costs, how fast each model is on your Mac, and a quality rating from Artificial Analysis. Sort by best value, quality, speed or cost; each picker remembers your choice. On-device models stay on top in every sort: they're free, private and work offline. Costs are only shown where the billing unit is known (otherwise "—"); speeds come from your own runs (on-device models show our benchmark until then); quality comes from Artificial Analysis's public leaderboards as of October 2026, and models they don't cover say "not rated". Click a model for what it's good at, then Use it; ↑↓, →, Return and Esc work too.
- The sidebar is a little wider, so track names read in full (hover for the whole name), and ↑↓ move through it again.
- Agents read to me unasked in Setup → Command line and agents: Off, Long replies, When they need me, or Everything, the same setting as
vp agents read-aloud. - The read-along card fades its text at the edges instead of cutting a line in half.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another agent to read you its summary: the skill now tells it to write for listening, open the follow-along card and read it with Voice Pipes, then wait while you listen and steer.
1.6.4
- Follow along with anything read aloud. The HUD has a new button while it's reading: it opens a card above it with the whole text, one sentence per line, the one being read lit up and the card scrolling with the voice. Hover over it to look ahead; it picks up the voice again when you move away. Handy for long answers you didn't select yourself.
- Speed it up or slow it down as it reads. − and + on the card change the speed on the spot, from 0.6× to 2×, with every kind of voice. The card stays open for the next reading until you close it.
vp speed 1.4changes the speed from a terminal;vp open reading/vp close readingshow or hide the text.vp uino longer reports a focused field after you've left the page it was on.
1.6.3
- The command line keeps up by itself.
vpalways runs the app's own version, so updating from the menu updates it too. Now the rest follows as well: the agent skill is refreshed at every launch, an agent you install later (say Codex) gets it automatically, and the Claude Code session hook follows the app. If you move Voice Pipes to another folder,vpis repointed at it; if that needs your password, Setup → Checks shows "vp points at a moved app" with a Fix… button.
1.6.2
- Install everything from a terminal. One line installs the app, the
vpcommand and the agent skill, and starts Voice Pipes:curl -fsSL https://github.com/brancusi/voice-tools-releases/releases/latest/download/install.sh | bash. It asks nothing, so an agent can run it too, and it refuses any download that isn't signed by Voice Pipes' developer and notarized by Apple. Run it again to repair or reinstall;--uninstallremoves the app, itsvplinks and the skill, and keeps your config, history and keys.
1.6.1
- Agents can show you, not just tell you.
vp opennow reaches anything in the app: a track with one block's settings open (vp open track clean-dictation --step 4), a route card, the cursor in a field (--field prompt), History filtered to a track or a search, a word in Vocabulary, a section of Setup, a step of the setup window, or the menu bar panel (vp open menu). What it points at flashes, and it answers with what's on screen.vp uiprints that on its own;vp closeclose...
Voice Pipes 1.8.0
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.8.0
- Model pickers now show what a paragraph costs, how fast each model is on your Mac, and a quality rating from Artificial Analysis. Sort by best value, quality, speed or cost; each picker remembers your choice. On-device models stay on top in every sort: they're free, private and work offline. Costs are only shown where the billing unit is known (otherwise "—"); speeds come from your own runs (on-device models show our benchmark until then); quality comes from Artificial Analysis's public leaderboards as of October 2026, and models they don't cover say "not rated". Click a model for what it's good at, then Use it; ↑↓, →, Return and Esc work too.
- The sidebar is a little wider, so track names read in full (hover for the whole name), and ↑↓ move through it again.
- Agents read to me unasked in Setup → Command line and agents: Off, Long replies, When they need me, or Everything, the same setting as
vp agents read-aloud. - The read-along card fades its text at the edges instead of cutting a line in half.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another agent to read you its summary: the skill now tells it to write for listening, open the follow-along card and read it with Voice Pipes, then wait while you listen and steer.
1.6.4
- Follow along with anything read aloud. The HUD has a new button while it's reading: it opens a card above it with the whole text, one sentence per line, the one being read lit up and the card scrolling with the voice. Hover over it to look ahead; it picks up the voice again when you move away. Handy for long answers you didn't select yourself.
- Speed it up or slow it down as it reads. − and + on the card change the speed on the spot, from 0.6× to 2×, with every kind of voice. The card stays open for the next reading until you close it.
vp speed 1.4changes the speed from a terminal;vp open reading/vp close readingshow or hide the text.vp uino longer reports a focused field after you've left the page it was on.
1.6.3
- The command line keeps up by itself.
vpalways runs the app's own version, so updating from the menu updates it too. Now the rest follows as well: the agent skill is refreshed at every launch, an agent you install later (say Codex) gets it automatically, and the Claude Code session hook follows the app. If you move Voice Pipes to another folder,vpis repointed at it; if that needs your password, Setup → Checks shows "vp points at a moved app" with a Fix… button.
1.6.2
- Install everything from a terminal. One line installs the app, the
vpcommand and the agent skill, and starts Voice Pipes:curl -fsSL https://github.com/brancusi/voice-tools-releases/releases/latest/download/install.sh | bash. It asks nothing, so an agent can run it too, and it refuses any download that isn't signed by Voice Pipes' developer and notarized by Apple. Run it again to repair or reinstall;--uninstallremoves the app, itsvplinks and the skill, and keeps your config, history and keys.
1.6.1
- Agents can show you, not just tell you.
vp opennow reaches anything in the app: a track with one block's settings open (vp open track clean-dictation --step 4), a route card, the cursor in a field (--field prompt), History filtered to a track or a search, a word in Vocabulary, a section of Setup, a step of the setup window, or the menu bar panel (vp open menu). What it points at flashes, and it answers with what's on screen.vp uiprints that on its own;vp closecloses windows, the panel or a sheet. No screen recording or Accessibility access involved.--backgroundshows a window without taking your keyboard. - Watch a track being built. When an agent (or you, in an editor) changes config.toml, the open track flashes what changed and opens a single new or changed block. The editor also stays where it was: a block you had open no longer closes on every save.
1.6.0
- Your setup is now a file. Tracks, hotkeys and settings live in
~/.config/voice-pipes/config.toml, and the vocabulary invocabulary.tomlbeside it: readable, commented TOML with a full block reference at the end and a schema your editor can check. Save and it applies within a second. If an edit doesn't check out, the last good version keeps running and Setup → Checks says what's wrong, on which line, with a "did you mean". Every change keeps a back...
Voice Pipes 1.7.4
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.7.4
- Ask an agent to build or change a pipeline in plain words. "Make me a track that cleans up my dictation and posts it to my notes", "read hard text with a cloud voice": the agent skill now walks Claude Code, Codex and others through it: they edit config.toml (which holds everything the app does, one to one), check it, open the track so you can watch it change, try it, and tell you what they built. The skill carries the full block and settings reference, generated from the same source as config.toml's, so it always matches.
- Agents read aloud what you want them to. Tell an agent "read me anything that needs my attention" and it's saved for every agent and every session:
off(only when asked),long(summaries, reports, explanations),attention(that, plus questions, finished tasks and problems) orall. Set it withvp agents read-aloud attention, orread_aloudin config.toml under [settings.agents]. - Setup → Reading, tidied: each key is one chip with its ×, "Anywhere while reading" lists only the shortcuts you've set plus + Add shortcut, and Reset to defaults is clearly off when there's nothing to reset.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another agent to read you its summary: the skill now tells it to write for listening, open the follow-along card and read it with Voice Pipes, then wait while you listen and steer.
1.6.4
- Follow along with anything read aloud. The HUD has a new button while it's reading: it opens a card above it with the whole text, one sentence per line, the one being read lit up and the card scrolling with the voice. Hover over it to look ahead; it picks up the voice again when you move away. Handy for long answers you didn't select yourself.
- Speed it up or slow it down as it reads. − and + on the card change the speed on the spot, from 0.6× to 2×, with every kind of voice. The card stays open for the next reading until you close it.
vp speed 1.4changes the speed from a terminal;vp open reading/vp close readingshow or hide the text.vp uino longer reports a focused field after you've left the page it was on.
1.6.3
- The command line keeps up by itself.
vpalways runs the app's own version, so updating from the menu updates it too. Now the rest follows as well: the agent skill is refreshed at every launch, an agent you install later (say Codex) gets it automatically, and the Claude Code session hook follows the app. If you move Voice Pipes to another folder,vpis repointed at it; if that needs your password, Setup → Checks shows "vp points at a moved app" with a Fix… button.
1.6.2
- Install everything from a terminal. One line installs the app, the
vpcommand and the agent skill, and starts Voice Pipes:curl -fsSL https://github.com/brancusi/voice-tools-releases/releases/latest/download/install.sh | bash. It asks nothing, so an agent can run it too, and it refuses any download that isn't signed by Voice Pipes' developer and notarized by Apple. Run it again to repair or reinstall;--uninstallremoves the app, itsvplinks and the skill, and keeps your config, history and keys.
1.6.1
- Agents can show you, not just tell you.
vp opennow reaches anything in the app: a track with one block's settings open (vp open track clean-dictation --step 4), a route card, the cursor in a field (--field prompt), History filtered to a track or a search, a word in Vocabulary, a section of Setup, a step of the setup window, or the menu bar panel (vp open menu). What it points at flashes, and it answers with what's on screen.vp uiprints that on its own;vp closecloses windows, the panel or a sheet. No screen recording or Accessibility access involved.--backgroundshows a window without taking your keyboard. - Watch a track being built. When an agent (or you, in an editor) changes config.toml, the open track flashes what changed and opens a single new or changed block. The editor also stays where it was: a block you had open no longer closes on every save.
1.6.0
- Your setup is now a file. Tracks, hotkeys and settings live in
~/.config/voice-pipes/config.toml, and the vocabulary invocabulary.tomlbeside it: readable, commented TOML with a full block reference at the end and a schema your editor can check. Save and it applies within a second. If an edit doesn't check out, the last good version keeps running and Setup → Checks says what's wrong, on which line, with a "did you mean". Every change keeps a backup, including edits from outside the app, and a config.toml symlinked from your dotfiles stays linked. Your current tracks and words move over by themselves. vp, the command line. Setup → Install command-line tool (also a new step in the setup window) putsvpon your PATH. Run a track with text (vp run clean-dictation --text "…", or pipe text in), speak (vp say), ask a question out loud and get the spoken answer back (vp ask), listen, transcribe a file on this Mac, browse history, edit vocabulary, check the config, and watch runs as they happen. It's built for agents: compact output, next-step hints, and no prompts.- Sign in to OpenRouter from the command line.
vp auth login openrouteropens the browser and saves the key to your Keychain;vp auth set typesafetakes a Jev key. Keys are never printed. - Secrets for your own endpoints.
vp secret set notes, then${secret:notes}in an HTTP block's URL, headers or body. The value stays in the Keychain, not in the file. - **Agents know Voice Pipes...
Voice Pipes 1.7.3
Voice Pipes (formerly Voice Tools): a menu bar app that runs tracks, hotkey-triggered pipelines for dictation and read-aloud.
1.7.3
- Steer a reading from the keyboard. While something is read aloud: Esc stops, Space pauses, j/k (or ↓/↑) move a sentence, h/l (or −/=) change the speed, g/G jump to the start or end. Point at the HUD or click it and its keys work; move away or click elsewhere and your typing goes straight back to the app you were in. Voice Pipes never comes to the front. A small "KEYS ON" in the HUD shows when it's listening, with your keys.
- Make it yours, in Setup → Reading. Choose when the HUD takes the keyboard: Always (the moment something starts reading, great for agent read-backs), Point or click, Click, or Never; whether clicking away keeps reading or stops; and every key, so faster and slower can be whatever you like. Add shortcuts that work in any app but only while something is reading (say ⌃⌥→ for faster); none are set until you choose. It's all in config.toml too, under [settings.reading].
- Agents can steer too:
vp next,vp prev, alongsidevp speed.
1.7.2
- History now shows every step of a run: what each one took, cost and gave, including which branch Jev picked and what happened when a step failed. Open a run with the ▸ beside it (or → on the keyboard): every step in order with its time and output, the path Jev took through a Branch or Route and the ones it didn't take, and, if an LLM failed and passed its text on, what the next step got instead. Copy log copies it as plain text.
- What it cost. Each step shows its tokens and cost (exact, as OpenRouter reports them; cloud speech is estimated from its price, marked ≈), on-device steps say "on this Mac", and each run shows its total. The top of History adds up today and the last 7 days. From a terminal:
vp history show <n>for a run's log,vp history usagefor totals. - History shows each run as its own card, and fits a narrow window: search on its own line, track filters that wrap. Runs from before this version have no log.
vp --versionnow shows the real version (it said "dev" when run from your PATH).
1.7.1
- A new Add step picker. "+ Add step" now opens a proper picker in the Voice Pipes look instead of a plain menu: every block with its icon and a one-line description, what fits after the previous step first, and the rest in a dimmed "Doesn't fit" section that says why ("needs audio", "starts a pipeline"). It shows which blocks need a Jev or OpenRouter key you haven't added yet. Type to filter (it searches descriptions too: "clipboard" finds Copy), ↑↓ to move, Return to add, Esc to close. Inside a branch it says which branch you're adding to.
1.7.0
- Branches: one track, several paths. The new Branch · Jev block asks Jev a question about the text, like "how hard is this to read aloud?" or "what is this about?", and runs the matching branch's own steps (any blocks, even another branch) before the track carries on. You decide what each branch does, right in the editor; Jev picks in about a third of a second.
- Read aloud handles tricky text. It now starts with a branch: plain prose is read as it is; text with numbers, prices or dates is first spelled out by a fast model (about 0.6 s), so "$4.2M" is read as "four point two million dollars"; code, file paths, URLs and tables are rewritten into sentences you can follow by ear. Want a different voice for the hard stuff? Add a Speak block to that branch. (New installs get this; to switch your existing Read aloud, add a Branch block, or ask your agent to.)
- Download progress you can trust. The model downloads in setup and in Setup → On this Mac now show a steady bar with "212 / 483 MB · 11 MB/s", then "preparing for this Mac…" while the model compiles. No more flickering, and the sizes shown are the real ones.
1.6.5
- Click a sentence to jump there. In the HUD's follow-along card, click any sentence and reading carries on from it, with any voice. Handy when an agent reads you a long summary: skip ahead to the part you care about, or go back over something you missed.
- A running word cursor. A thin underline moves along under the word being spoken, and the part of the sentence already read is a touch brighter. It's exact with macOS voices and closely estimated with the others.
- Agents know they can read to you. Ask Claude Code or another agent to read you its summary: the skill now tells it to write for listening, open the follow-along card and read it with Voice Pipes, then wait while you listen and steer.
1.6.4
- Follow along with anything read aloud. The HUD has a new button while it's reading: it opens a card above it with the whole text, one sentence per line, the one being read lit up and the card scrolling with the voice. Hover over it to look ahead; it picks up the voice again when you move away. Handy for long answers you didn't select yourself.
- Speed it up or slow it down as it reads. − and + on the card change the speed on the spot, from 0.6× to 2×, with every kind of voice. The card stays open for the next reading until you close it.
vp speed 1.4changes the speed from a terminal;vp open reading/vp close readingshow or hide the text.vp uino longer reports a focused field after you've left the page it was on.
1.6.3
- The command line keeps up by itself.
vpalways runs the app's own version, so updating from the menu updates it too. Now the rest follows as well: the agent skill is refreshed at every launch, an agent you install later (say Codex) gets it automatically, and the Claude Code session hook follows the app. If you move Voice Pipes to another folder,vpis repointed at it; if that needs your password, Setup → Checks shows "vp points at a moved app" with a Fix… button.
1.6.2
- Install everything from a terminal. One line installs the app, the
vpcommand and the agent skill, and starts Voice Pipes:curl -fsSL https://github.com/brancusi/voice-tools-releases/releases/latest/download/install.sh | bash. It asks nothing, so an agent can run it too, and it refuses any download that isn't signed by Voice Pipes' developer and notarized by Apple. Run it again to repair or reinstall;--uninstallremoves the app, itsvplinks and the skill, and keeps your config, history and keys.
1.6.1
- Agents can show you, not just tell you.
vp opennow reaches anything in the app: a track with one block's settings open (vp open track clean-dictation --step 4), a route card, the cursor in a field (--field prompt), History filtered to a track or a search, a word in Vocabulary, a section of Setup, a step of the setup window, or the menu bar panel (vp open menu). What it points at flashes, and it answers with what's on screen.vp uiprints that on its own;vp closecloses windows, the panel or a sheet. No screen recording or Accessibility access involved.--backgroundshows a window without taking your keyboard. - Watch a track being built. When an agent (or you, in an editor) changes config.toml, the open track flashes what changed and opens a single new or changed block. The editor also stays where it was: a block you had open no longer closes on every save.
1.6.0
- Your setup is now a file. Tracks, hotkeys and settings live in
~/.config/voice-pipes/config.toml, and the vocabulary invocabulary.tomlbeside it: readable, commented TOML with a full block reference at the end and a schema your editor can check. Save and it applies within a second. If an edit doesn't check out, the last good version keeps running and Setup → Checks says what's wrong, on which line, with a "did you mean". Every change keeps a backup, including edits from outside the app, and a config.toml symlinked from your dotfiles stays linked. Your current tracks and words move over by themselves. vp, the command line. Setup → Install command-line tool (also a new step in the setup window) putsvpon your PATH. Run a track with text (vp run clean-dictation --text "…", or pipe text in), speak (vp say), ask a question out loud and get the spoken answer back (vp ask), listen, transcribe a file on this Mac, browse history, edit vocabulary, check the config, and watch runs as they happen. It's built for agents: compact output, next-step hints, and no prompts.- Sign in to OpenRouter from the command line.
vp auth login openrouteropens the browser and saves the key to your Keychain;vp auth set typesafetakes a Jev key. Keys are never printed. - Secrets for your own endpoints.
vp secret set notes, then${secret:notes}in an HTTP block's URL, headers or body. The value stays in the Keychain, not in the file. - Agents know Voice Pipes. Installing the command-line tool also installs a skill for Claude Code, Codex and other agents, so they can tell you things out loud, ask you questions by voice, run your tracks and safely edit your config.
1.5.1
- Welcome to Voice Pipes, redrawn. The first-run window has a step bar, plain titles, and live status: Continue waits until both permissions are on (or Skip), the models step shows Parakeet's download in MB and offers the optional Supertonic voices, the keys step links to openrouter.ai and lets you Skip, stay local, and Try it shows your first run's timings before the Wrangler tips his hat.
- Empty pages with somewhere to go. An empty Vocabulary offers Train a word… (type the spelling, then say it a few times); with no tracks left, Restore the starter tracks brings back Fast dictation, Clean dictation and Read aloud, leaving off any hotkey another track already uses.
- The setup window now opens centred, and a failed model download no longer leaves a stuck percentage.
1.5.0
- A guided first launch. New installs open Set up Voice Pipes: it explains and asks for Microphone and Accessibility one ...