Repository navigation
Releases: Dreaminko/VRCS
Release list
VRCS 0.1.11
This release introduces native live translation and improves how real-time source and translated subtitles stay aligned.
New Features
- Native live translation: Added Gemini Live Translate (Preview) and OpenAI Realtime Translation, allowing each audio source's first target language to be translated directly from the recognition stream.
- Language-learning AI preferences: The explanation language can now follow the app language or use a selected language, and the preference is preserved across sessions.
Fixes & Improvements
- More reliable sentence alignment: Source and translation streams are segmented independently, late translations stay attached to their source subtitles, and ambiguous multi-sentence results are shown as a shared translation instead of being assigned to the wrong lines.
- Improved streaming stability: Gemini streaming responses and WebSocket events now handle partial updates, completion, and interrupted sessions more reliably.
Notes
- OpenAI Realtime Translation is still being optimized and is not recommended for regular use yet. Service delays and differences in source and target sentence boundaries may still affect subtitle grouping and alignment.
- Live Translate limitations: Live Translate applies only to the first target language for each audio source and requires automatic translation. Text prompts, glossaries, and thinking settings do not apply. The service may also generate and bill translated audio even though VRCS does not play it.
本次更新加入原生流式翻译,并改进实时字幕中原文与译文的对应关系。
新功能
- 原生流式翻译: 新增 Gemini Live Translate(Preview)和 OpenAI Realtime Translation,可直接通过识别流将每路音源的第一目标语言进行翻译。
- 语言学习 AI 偏好: 解释语言现在可以跟随软件语言或使用指定语言,相关偏好会在不同会话间保留。
修复与改进
- 更可靠的字幕对齐: 原文与译文现在会独立断句,延迟到达的译文会继续关联原字幕;无法可靠逐句对应的多句结果会显示为合并译文,减少译文错位。
- 更稳定的流式处理: 改进 Gemini 流式响应和 WebSocket 事件处理,提高临时结果、完成事件及连接中断场景下的稳定性。
使用说明
- OpenAI Realtime Translation 仍在优化中,现阶段不太建议作为日常翻译方案使用。 服务延迟,以及原文与译文断句数量不一致,仍可能影响字幕的分组和对齐。
- Live Translate 使用限制: Live Translate 仅用于每路音源的第一目标语言,并且需要开启自动翻译。文本 Prompt、术语表和思考设置不适用。服务还可能生成译后音频并计费,但 VRCS 不会播放这些音频。
Full Changelog: 0.1.10...0.1.11
VRCS 0.1.10
- Added Gemini real-time speech recognition with
gemini-3.5-transcribe-live, including interim transcripts, language selection, and custom vocabulary. - Added
gpt-live-transcribeto the OpenAI speech recognition model list, contributed by @blufish1234 in #3. - Improved audio capture when following Windows default devices, with automatic switching and recovery after repeated device disconnections.
- Fixed audio device listing failures caused by individual unavailable devices, so other available devices remain selectable.
- Fixed disconnected devices blocking audio settings changes, allowing each channel's device selection and voice activation threshold to be updated independently.
- Fixed microphone tests stopping unexpectedly during app status updates, including during setup and calibration.
- Added a Traditional Chinese README, contributed by @blufish1234 in #4.
- 新增 Gemini 实时语音识别,支持
gemini-3.5-transcribe-live,提供实时中间字幕、语言选择和自定义词汇。 - 在 OpenAI 语音识别模型列表中新增
gpt-live-transcribe,由 @blufish1234 贡献,见 #3。 - 改进跟随 Windows 默认设备时的音频采集,支持自动切换设备,并在设备多次断连后自动恢复。
- 修复个别音频设备不可用时导致设备列表加载失败的问题,其他可用设备仍可正常选择。
- 修复设备断连导致音频设置无法保存的问题,支持独立调整各通道的设备选择和语音触发阈值。
- 修复应用状态更新时麦克风测试意外停止的问题,改善首次设置和校准体验。
- 新增繁体中文 README,由 @blufish1234 贡献,见 #4。
Contributors / 贡献者
VRCS 0.1.9
VRCS 0.1.7
New Features
- Semantic endpointing (experimental): Added optional Smart Turn endpointing to check whether speech is complete after a pause, producing more natural segment boundaries for recognition services controlled by VRCS.
- Resizable compact subtitles: Compact mode can now be resized vertically to show up to four recent subtitles and their translations.
Fixes & Improvements
- More natural speech segmentation: Improved voice activity stability and made target segment length prefer nearby pauses, reducing abrupt splits during continuous speech.
- Custom recognition models: Recognition settings now accept custom provider model names instead of limiting input to built-in suggestions.
- Reliable live subtitle scrolling: When you scroll back through live subtitles, new entries no longer pull the view to the bottom until you return there.
- More accurate language matching: Equivalent language codes, including Simplified and Traditional Chinese variants, are now matched correctly to avoid unnecessary translations and configuration errors.
新功能
- 语义断句(实验性): 新增可选的 Smart Turn 语义断句,在停顿后判断话语是否完整,让由 VRCS 本地控制的识别服务在更自然的位置分段。
- 可调整的紧凑字幕: 紧凑模式现在支持垂直调整窗口大小,最多可同时显示最近四条字幕及其译文。
修复与改进
- 更自然的语音分段: 提高了语音活动检测的稳定性;达到目标片段长度时会优先寻找附近的停顿,减少连续讲话被生硬切断的情况。
- 自定义识别模型: 识别设置现在支持填写服务商的自定义模型名称,不再局限于内置建议。
- 更可靠的实时字幕滚动: 手动向上查看历史字幕后,新字幕不会再将页面拉回底部;回到底部后会恢复自动跟随。
- 更准确的语言匹配: 现在能正确识别等价的语言代码,包括简体与繁体中文的不同代码,避免同语种内容被重复翻译或因未配置翻译服务而报错。
VRCS 0.1.6
New Features
- Traditional Chinese UI: Added Traditional Chinese as an interface language.
- Alibaba Cloud Token Plan: Added native text-generation and real-time speech recognition support, including
qwen-audio-3.0-realtime-plus, with automatic migration for existing Token Plan recognition profiles. - Qwen Audio streaming ASR: Added
qwen-audio-3.0-asr-flash-streamingto the Alibaba Cloud Qwen Audio / Fun-ASR realtime service.
Fixes & Improvements
- In-app updates: Fixed download progress events so update status is reported correctly.
- Qwen model handling: Added thinking controls for
qwen3.8-maxand improved automatic translation model selection for Token Plan.
新功能
- 繁体中文界面: 新增繁体中文作为界面语言。
- 阿里云 Token Plan:新增原生文本生成与实时语音识别支持,包含
qwen-audio-3.0-realtime-plus,并会自动迁移已有的 Token Plan 语音识别配置。 - Qwen Audio 流式语音识别:阿里云 Qwen Audio 和 Fun-ASR 实时服务现已支持
qwen-audio-3.0-asr-flash-streaming。
修复与改进
- 应用内更新:修复下载进度事件,确保应用能正确显示更新状态。
- Qwen 模型处理:新增
qwen3.8-max的思考模式控制,并改进 Token Plan 的翻译模型自动选择。
VRCS 0.1.5
New Features
- Search across all subtitles. Search originals and translations from the conversation sidebar, highlight matching text, and jump directly to the relevant subtitle with its surrounding context; press
Ctrl+Fto start searching.
Fixes & Improvements
- More resilient audio capture. When Windows resources or device formats block startup, VRCS now tries compatible capture modes automatically, improving reliability for system audio, microphones, and VRChat process audio.
- Fixed OpenAI Realtime Transcription connections. VRCS now opens the session with the correct realtime model while continuing to use the selected transcription model.
- API profiles refresh immediately. Creating, editing, or deleting a profile or credential now updates related selectors without leaving stale settings behind.
- More reliable SteamVR subtitle refreshes. Headset and wrist overlays now keep shared textures active between frames, reducing repeated submissions and improving continuous updates.
- Faster first navigation. The Learning and Settings pages now preload while VRCS is idle, reducing the wait when they are first opened.
- Improved sidebar accessibility. The collapsed conversation sidebar now preserves its semantic container, making its role clearer to screen readers.
新功能
- 搜索全部字幕。 可在会话侧边栏中搜索所有对话的原文和译文,结果会高亮关键词,并可直接跳到对应字幕及上下文;按
Ctrl+F即可开始搜索。
修复与改进
- 音频采集更稳定。 当 Windows 资源或设备格式导致采集无法启动时,VRCS 会自动尝试兼容模式,提高系统音频、麦克风和 VRChat 进程音频的采集成功率。
- 修复 OpenAI 实时转写连接。 VRCS 现在会通过正确的实时会话模型建立连接,同时继续使用所选转写模型。
- API 配置即时同步。 新建、修改或删除 API 配置及凭据后,相关选择器会立即刷新,不再显示旧配置。
- SteamVR 字幕刷新更可靠。 头显和手腕字幕会在帧间持续复用共享纹理,减少重复提交,并提高连续更新的稳定性。
- 首次切换页面更快。 学习与设置页面会在 VRCS 空闲时提前加载,首次打开时等待更少。
- 改进侧边栏辅助功能。 折叠后的会话侧边栏现在会保留完整的语义容器,让屏幕阅读器更准确地识别其作用。
VRCS 0.1.3
VRCS 0.1.3
This release adds subtitle search across conversation history and multilingual translation display for SteamVR overlays, while improving navigation responsiveness, overlay rendering, and settings reliability.
Highlights
- Added subtitle search across all conversations, covering both original text and translations, with highlighted matches and direct navigation to the surrounding context.
- Added a VR Overlay translation display option to show only the preferred language or all translated languages.
Improvements
- Learning and Settings pages now preload during idle time and on navigation intent, reducing delays when switching workspaces.
- VR Overlay updates now reuse shared DXGI textures, reducing texture update overhead in SteamVR.
- Simplified translation route settings with cleaner labels and numeric ordering.
Fixes
- Prevented stale settings responses from restoring deleted API profiles or overwriting newer configuration changes.
- Kept the collapsed conversation sidebar inside its semantic container for more consistent layout and accessibility behavior.
本次更新新增了对话历史字幕搜索和 SteamVR 覆盖层多语言译文显示,并改进了页面切换速度、覆盖层渲染和设置可靠性。
主要更新
- 新增全部对话字幕搜索,覆盖原文和译文,支持关键词高亮,并可直接跳转到所在对话的上下文。
- VR Overlay 新增译文显示选项,可只显示首选语言,也可同时显示全部目标语言。
改进
- 学习与设置页面会在空闲时或用户准备切换时预加载,减少页面切换等待。
- VR Overlay 更新现在会复用 DXGI 共享纹理,降低 SteamVR 中的纹理更新开销。
- 精简翻译路线设置,采用数字排序并删去重复说明。
修复
- 防止过期的设置响应恢复已删除的 API 配置,或覆盖较新的配置更改。
- 修复对话侧栏折叠后脱离语义容器的问题,使布局和辅助功能行为保持一致。
VRCS 0.1.2
VRCS 0.1.2
This release expands multilingual translation workflows and subtitle interaction, while improving recognition, OSC output, audio capture, SteamVR integration, and release reliability.
Highlights
- Added Ask AI for selected subtitle text. Ask a custom question or use quick prompts for meaning, grammar, and phrasing.
- Added up to three ordered translation routes for both microphone and speaker audio, with independent target languages, API profiles, and models.
- Added reusable language presets for recognition, translation routes, and OSC output preferences.
- Added a dedicated Glossaries settings section shared by LLM translation and supported ASR services.
- Added a glossary table editor with search, filtering, inline editing, multi-selection, and bulk actions.
- Added three OSC multilingual output strategies:
- Preferred language only
- Round robin
- All languages in one message
- Added an option to send translated text without including the original subtitle.
Improvements
- Recognition model lists can now load compatible versioned Qwen Realtime and FunASR models dynamically.
- OpenAI Realtime transcription now uses the model selected in settings.
- Speaker and microphone capture pipelines now start in parallel for faster initialization.
- VR overlay reconnection now waits until SteamVR is running.
- Explicit manual Chatbox messages can be sent while mute synchronization is blocking automatic output.
- Existing translation routes and glossary settings are migrated to the new configuration format automatically.
- Improved release validation, updater configuration, and CUDA architecture checks.
Fixes
- Suppressed repeated VRCX world-name echoes in partial and final recognition results.
- Prevented duplicate final ASR events for the same utterance.
- Improved OSC revision handling to discard stale automatic output after configuration or mute-state changes.
- Fixed unbalanced WASAPI COM initialization during audio device discovery and capture.
- Improved rollback behavior when one of the audio capture pipelines fails to start.
Compatibility note
The CUDA edition now requires the CUDA 13.x Runtime, cuBLAS, and a compatible NVIDIA driver. The standard edition does not require CUDA.
本次更新扩展了多语言翻译流程与字幕交互能力,并改进了语音识别、OSC 输出、音频采集、SteamVR 集成和发布可靠性。
主要更新
- 新增字幕划词 “问 AI” 功能,可自由提问,也可快速查询词义、语法和表达方式。
- 麦克风与扬声器翻译现在分别支持最多三条有序翻译路线,每条路线可独立设置目标语言、API 配置和模型。
- 新增语言预设,可保存识别语言、翻译路线和 OSC 输出偏好。
- 新增独立的 “术语表” 设置页面,术语可同时供 LLM 翻译和支持上下文的 ASR 服务使用。
- 新增术语表格编辑器,支持搜索、筛选、行内编辑、多选和批量操作。
- 新增三种 OSC 多语言输出策略:
- 仅输出首选语言
- 按语言轮换输出
- 在同一条消息中输出所有语言
- 新增仅发送译文、不附带原文的选项。
改进
- Qwen Realtime 与 FunASR 现在可以动态加载兼容的版本化识别模型。
- OpenAI Realtime 转写现在会使用设置中选择的模型。
- 扬声器与麦克风采集管线改为并行启动,缩短初始化时间。
- VR 覆盖层会等待 SteamVR 启动后再尝试连接。
- 静音同步阻止自动输出时,用户主动发送的 Chatbox 消息仍可正常发送。
- 旧版翻译和术语表配置会自动迁移到新的配置格式。
- 改进发布验证、自动更新配置和 CUDA 架构检查。
修复
- 过滤 VRCX 上下文造成的世界名称重复识别结果。
- 防止同一段语音产生重复的最终识别事件。
- 改进 OSC 修订状态处理,配置或静音状态变化后不再发送过期的自动输出。
- 修复音频设备查询和采集过程中的 WASAPI COM 初始化不平衡问题。
- 改进音频管线启动失败时的回滚处理。
VRCS 0.1.1
English
VRCS 0.1.1 focuses on API configuration, cloud speech recognition, real-time VR overlay output, and Windows audio compatibility. It also introduces in-app software updates.
Features and improvements
-
Redesigned API management
- Unified cloud providers, local services, and custom protocols.
- API profiles can now enable speech-to-text, text generation, and text translation independently.
- Connection fields, authentication methods, models, and available features are adapted to each provider.
- Improved profile creation, editing, validation, credential management, and model selection.
-
Expanded cloud speech recognition
- Unified management for Qwen Realtime, Fun-ASR Realtime, OpenAI Realtime, and Groq Transcription.
- Added segmented-upload transcription with Groq Whisper models.
- Recognition settings now distinguish real-time streaming from segmented upload services.
- Model and context settings are stored separately for each recognition service.
-
Improved translation services
- Improved translation model discovery, refresh, and automatic selection.
- Reasoning controls now adapt to the selected provider and model.
- Internal reasoning content is removed from the final translation.
- Available languages and translation options now follow provider capabilities.
-
Added in-app updates
- Check, download, and install updates from GitHub Releases.
- Optional automatic update checks.
- Separate update targets for standard and CUDA builds.
-
Real-time VR overlay output
- Partial recognition results can now be shown in headset and wrist overlays.
- Improved transitions between live recognition, finalized subtitles, and translation results.
Fixes
- Migrated OpenAI Realtime transcription to the generally available API.
- Improved Windows WASAPI format negotiation with automatic conversion fallback.
- Added retries for transient audio initialization failures.
- Added actionable errors for busy devices, denied permissions, stopped audio services, unsupported formats, and other capture failures.
- Fixed translation prompt editing becoming unresponsive or reverting during autosave.
- Clarified the successful ASR test status in the onboarding wizard.
- Updated README links, including references to VRCX-0.
Upgrade notes
- Existing API profiles and recognition settings are migrated to the new service and capability model.
- Active transcription must be stopped before installing an application update.
Join our Discord or QQ group to share feedback and join the discussion!
中文
VRCS 0.1.1 重点改进了 API 配置、云端语音识别、VR 覆盖层实时显示和 Windows 音频兼容性,同时加入应用内更新功能。
新功能与改进
-
重新设计 API 管理
- 统一管理云服务、本地服务和自定义协议。
- API 配置现在可以分别启用语音转文字、文本生成和文本翻译能力。
- 根据服务提供商自动显示适用的连接参数、认证方式、模型和功能。
- 改进 API 配置的创建、编辑、验证、凭据管理和模型选择体验。
-
扩展云端语音识别
- 统一管理 Qwen Realtime、Fun-ASR Realtime、OpenAI Realtime 和 Groq Transcription。
- 新增基于 Groq Whisper 模型的分段上传识别。
- 在设置中明确区分实时流式识别与分段上传识别。
- 支持按识别服务分别保存模型和上下文设置。
-
改进翻译服务
- 优化翻译模型的发现、刷新和自动选择。
- 根据提供商和模型显示正确的推理模式选项。
- 自动隐藏模型的推理过程,避免其混入最终译文。
- 根据服务能力限制可选语言和翻译设置。
-
新增应用内更新
- 支持从 GitHub Releases 检查、下载并安装新版本。
- 支持自动检查更新。
- 分别匹配标准版和 CUDA 版更新包。
-
VR 覆盖层实时输出
- 可选择在头显和手腕覆盖层中显示语音识别的临时结果。
- 优化实时字幕、最终字幕及翻译结果之间的状态切换和显示。
修复
- 将 OpenAI Realtime 语音转写迁移到正式版 API。
- 改进 Windows WASAPI 音频格式协商,并在必要时启用自动格式转换。
- 对暂时性的音频初始化失败增加自动重试。
- 为设备占用、权限不足、音频服务未运行、格式不支持等问题提供更明确的错误提示。
- 修复翻译提示词在自动保存期间输入卡顿或内容回退的问题。
- 修复初始设置向导中 ASR 测试成功状态容易被误认为失败的问题。
- 更新 README 中的项目与 VRCX-0 相关链接。
升级说明
- 旧版 API 配置和语音识别设置会迁移到新的服务与能力模型。
- 安装应用更新前需要先停止正在进行的语音转写。
欢迎加入我们的 Discord 频道 或 QQ 群 参与讨论与交流!
Full Changelog: 0.1.0...0.1.1
VRCS 0.1.0
这是 VRCS 的首个正式版本。我投入了大量时间来打磨和完善这款软件,现在终于可以与大家分享了。
由于软件仍处于早期阶段,可能还存在不少需要解决的 Bug 和问题。
欢迎加入我们的 Discord 频道 或 QQ 群 参与讨论与交流!
This is the initial release of VRCS. I have spent a lot of time polishing and refining the software, and I think it's finally time to share it.
Since the software is still in its early stages, there may still be bugs and issues that need to be addressed.
Join our Discord or QQ group to share feedback and join the discussion!
Full Changelog: https://github.com/Dreaminko/VRCS/commits/0.1.0