-
Notifications
You must be signed in to change notification settings - Fork 0
AI Enhancement
Artificial intelligence (AI) Enhancement uses a local language model to revise transcribed text. It can smooth phrasing, remove distracting repetitions, improve readability, and follow custom instructions. The built-in models run on your computer; text is not sent to a cloud service.
Open the AI Enhancement tab and enable AI enhancement. Select a model that fits the available system memory and graphics memory, then download it. Voice2Win shows whether the model is available locally, loading, active, or in an error state.
Smaller and more heavily quantized models load faster and use less memory. Larger models and higher-precision variants can produce better results but need more memory and take longer. If Voice2Win warns that the selected speech and AI models are close to or beyond graphics-memory capacity, choose a smaller model or assign one workload to the main processor.
The Standard Instruction (Default Prompt) runs automatically after normal speech recognition when it is enabled. Its default behavior is to revise the transcription without changing its meaning and output only the final text.
If the standard instruction is disabled, normal dictation is inserted without AI enhancement.
Create additional prompts for specialized workflows. Each prompt can have:
- A title.
- An enabled or disabled state.
- Its own global keyboard shortcut.
- Instruction text sent with the transcription.
- An option to copy currently selected text into the instruction.
- Separate instruction text before and after the selected content.
Pressing a custom prompt shortcut records speech like normal dictation, then applies that specific prompt instead of the default one. Examples include summarizing, changing tone, translating, or formatting a dictated note.
This workflow edits text directly in another application:
- Select editable text in the target application.
- Press the configured AI Processing and Replace shortcut.
- Speak an instruction such as “Translate into English” or “Correct the spelling.”
- Voice2Win copies the selection, sends the selected text and spoken instruction to the configured AI runtime, and replaces the selection with the result.
The target application must support normal copy and paste operations. Do not move the selection or cursor while the operation is running. If no selection can be copied, the instruction is too short, or the model returns no output, Voice2Win leaves the original text unchanged and shows an indicator message.
The defaults suit most users:
| Setting | Default | Effect |
|---|---|---|
| Context Size | 4096 | Amount of text the model can consider at once. Higher values need more memory. |
| Max Tokens | 512 | Maximum output length. Higher values permit longer results. |
| Temperature | 0.2 | Lower values are more restrained and consistent; higher values are more creative. |
| Minimum transcription length | 20 characters | Shorter recordings skip AI enhancement. |
| System Prompt | Built-in role and rules | Sent before ordinary enhancement prompts. An empty value sends no system prompt. |
The system prompt does not apply to AI Processing and Replace templates because those already contain their complete instruction.
AI models normally use the model storage area. You can define an alternative AI-model folder, preferably on another drive, and copy existing downloads there.
AI Enhancement can run on the main processor, a compatible graphics card, Ollama, LM Studio, or a shared Voice2Win instance. See AI Models and Hardware.
Languages
English
Deutsch
Español
Français
Português (Brasil)
Italiano
Nederlands
Polski
Svenska
简体中文
日本語
한국어
हिन्दी
繁體中文
Bahasa Indonesia
Tiếng Việt
- Chuyển Giọng Nói Thành Văn Bản
- Chuyển Văn Bản Thời Gian Thực
- Tăng cường AI
- Lệnh Giọng Nói
- Ghi âm Cuộc trò chuyện
- Điều Khiển Từ Xa Thiết Bị Di Động