LLM-based two-way translator for VRChat
🇺🇸 English · 🇰🇷 한국어 · 🇯🇵 日本語 · 🇨🇳 简体中文 · 🇷🇺 Русский
pph_clip_1_github.mp4
If you want to see more of actual communication with other foreign friends through PuriPuly:
You've been there.
Wanting to comfort a friend,
but only managing: "Are you okay?"
You already know a 'translator'
can't carry what's truly in your heart.
So I built one that can.
- LLM-Powered Localization — Slang, colloquialisms, and casual/formal speech, all rendered naturally.
- Context Memory — Keeps the conversation flowing naturally with awareness of prior context.
- Two-way Voice Translation — Translates the other person's voice too, with VR subtitle overlay support.
- Start via Discord — Get going right away without a complex setup process.
-
How good is the translation quality? → When both you and the other person use this translator, you can have even the deepest kinds of conversations. Quantitatively, with Gemma 4 it scored 6× better than DeepL. See the 'Translation Comparison' section below for details.
-
How long does it take from speaking to getting a translation? → With Gemma 4 and a cloud STT service, latency is typically in the mid-to-late 1-second range.
-
Does it cost money to use? → Yes, but only later. New users get a free usage allowance, and even after that the pricing is very cheap; you can use it thousands of times for $1.
-
Do I need to get an API key? → Yes, but again, only later. At first, just install and authenticate via Discord to start using it.
-
How polished is the feature for translating the other person's voice? → It works best for one-on-one conversations in low-noise environments. Up to three people may be okay, but usability is not guaranteed. When using it in VRChat, use Earmuff to control the environment.
-
Voice recognition is poor / slow. → If you're using local Qwen ASR, we recommend switching to a cloud STT service. If you're on Intel, configure PuriPuly so it's pinned to P-cores only.
-
How are voice and conversation contents handled? → Voice and conversation contents are stored locally and are not sent to Puripuly servers. Other people's voices, transcripts, and translation results are never recorded. That said, the STT service and translation provider may process the data.
- We ran the experiment using Microsoft's Gemba MQM framework.
- It was set up as a multi-turn environment to better resemble real conversation.
- For the full results, see here.
| LLM \ ASR | Qwen ASR (Local) | Qwen ASR (Cloud) | Soniox | Deepgram |
|---|---|---|---|---|
| Gemma 4 26B A4B + 31B | 14,380 | 2,920 | 3,710 | 1,180 |
| DeepSeek V4 Flash | 19,410 | 3,080 | 3,980 | 1,210 |
| LLM \ ASR | Qwen ASR (Local) | Qwen ASR (Cloud) | Soniox | Deepgram |
|---|---|---|---|---|
| Gemma 4 26B A4B | 14,380 | 2,920 | 3,710 | 1,180 |
| Gemma 4 31B (OpenRouter) | 13,700 | 2,780 | 3,530 | 1,120 |
| Gemma 4 31B (Cerebras) | 920 | 730 | 770 | 540 |
| Gemini 3 Flash | 1,710 | 1,170 | 1,280 | 740 |
| Gemini 3.1 Flash-Lite | 3,430 | 1,770 | 2,030 | 940 |
| Qwen 3.5 Plus | 7,460 | 2,460 | — | — |
| Local LLMs | Unlimited | 3,660 | 5,000 | 1,290 |
| LLM \ ASR | Qwen ASR (Local) | Qwen ASR (Cloud) | Soniox | Deepgram |
|---|---|---|---|---|
| Gemma 4 26B A4B + 31B | ~$0.00007 | ~$0.0003 | ~$0.0003 | ~$0.0008 |
| DeepSeek V4 Flash | ~$0.00005 | ~$0.0003 | ~$0.0003 | ~$0.0008 |
| LLM \ ASR | Qwen ASR (Local) | Qwen ASR (Cloud) | Soniox | Deepgram |
|---|---|---|---|---|
| Gemma 4 26B A4B | ~$0.00007 | ~$0.0003 | ~$0.0003 | ~$0.0008 |
| Gemma 4 31B (OpenRouter) | ~$0.00007 | ~$0.0003 | ~$0.0003 | ~$0.0009 |
| Gemma 4 31B (Cerebras) | ~$0.0011 | ~$0.0014 | ~$0.0013 | ~$0.0019 |
| Gemini 3 Flash | ~$0.0006 | ~$0.0009 | ~$0.0008 | ~$0.0014 |
| Gemini 3.1 Flash-Lite | ~$0.0003 | ~$0.0006 | ~$0.0005 | ~$0.0011 |
| Qwen 3.5 Plus | ~$0.0001 | ~$0.0004 | — | — |
| Local LLMs | $0 | ~$0.0003 | ~$0.0002 | ~$0.0008 |
- Based on (Input 900 tokens + Output 12 tokens) × 1.2 avg LLM calls per utterance.
- Uses per Dollar is derived from the un-rounded values in the Cost per Utterance table.
- All costs and usage counts are approximate.
- DeepSeek assumes a 70% cache hit rate.
- Qwen API costs are based on the Beijing region.
- Pricing as of May 25, 2026 / Fast Response mode active.
| Service | Free Credit | Duration | Note |
|---|---|---|---|
| Deepgram | $200 | None | - |
| Google AI Studio | $10 | 1 year | Monthly for Gemini subscribers |
| Alibaba Cloud | 1M tokens per model | 90 days | Singapore region |
| Alibaba Cloud | ¥300 | 1 year | Students in China |
| Cerebras | 1M tokens daily | None | 5 calls per minute limit |
If you run into problems or anything feels unclear, feel free to DM me on Twitter/X.
-
Download the latest version from the Download page.
-
Install PuriPuly.
-
Click the TALK button.
-
Click the TRANS button, then authenticate via Discord.
-
Click the CAPTIONS button to turn on VR subtitles.
-
(Optional) Click the LISTEN button to enable translation of the other person's voice.
Peer voice translation needs a low-noise space to work properly. When using it in VRChat, use Earmuff to control the environment.
-
Enable OSC in VRChat: Action menu → Settings → OSC → Enable.
For bidirectional control setup and the stable parameter ABI, see VRChat OSC controls.
If audio capture does not work, open Settings > General and follow these steps.
- Change Audio Host API to Auto or MME.
- Select the correct microphone.
- Restart the app.
If Soniox/Gemini/Deepgram are blocked in your region, please use the following combination:
-
STT: Qwen ASR
-
LLM: DeepSeek V4 Flash
You can authenticate through QQ instead of Discord.
Follow the guide that matches the service you want to use.
For the translation LLM, we recommend using the Gemma 4 model through OpenRouter.
By the way, while you're setting things up, why not configure ASR too? PuriPuly delivers the best experience when paired with a cloud STT. For instance, even with the same Qwen ASR, local and cloud voice-recognition performance differ noticeably.
We recommend starting with Deepgram. Just signing up gets you $200 in free credits.
-
Set the options inside the red circle as shown in the screenshot.

-
Click the button inside the red circle to exit the payment screen.

If you clicked Authorize but you're still not authenticated, retry, or directly issue an API key as below and paste it in.
-
Set the options inside the red circle as shown in the screenshot.

-
Go to the DeepSeek official homepage and click the Access API button.

-
Click the button to copy the API key, then paste it into the API tab of the translator.

-
Login to the Deepgram Console.

-
Select STT (Speech-to-Text) on the service selection screen.

-
Go to Google AI Studio and click the Get API key button.

-
(Recommended) Click the yellow Set Up Billing button to upgrade to the paid tier. The tier transition may take a moment.

-
Go to Google Developer Program and join the program.

-
Login to Soniox Console.

-
Soniox requires prepaid credits. Once added, go to the API Keys menu.

-
Go to Cerebras and click Get started.

-
Choose the plan you want. We recommend starting with the free tier.

See ARCHITECTURE.md.
Upcoming work is tracked publicly on the PuriPuly project board.
| Surface | Recommended environment | Documentation |
|---|---|---|
| Python desktop application | Windows | This section |
| Broker service | Linux | broker/README.md |
| Native VR overlay | Windows | native/overlay/README.md |
The Python application requires Python 3.12 or 3.13.
Create and activate the Windows environment:
python -m venv .venv
.venv\Scripts\Activate.ps1Install the application and development dependencies:
python -m pip install --upgrade pip
pip install -e ".[dev]"uv may be used instead:
uv sync --devInstall the repository hooks:
pre-commit installFor Linux or WSL work, use .venv-wsl when it is available.
UV_PROJECT_ENVIRONMENT=.venv-wsl uv sync --devRepositories configured with direnv may run commands through:
direnv exec . <command>Run the Flet desktop application:
python -m puripuly_heart.main run-guiThe equivalent uv command is:
uv run python -m puripuly_heart.main run-guiDeveloper preview controls for hidden UI states are enabled with:
python -m puripuly_heart.main run-gui --debug-ui-previewFormat the Python sources and tests:
black src testsCheck formatting without modifying files:
black --check src testsRun lint checks:
ruff check src testsRun the complete Python test suite:
python -m pytestRun a focused test file or directory during development:
python -m pytest tests/path/to/test_file.pyBroker documentation is maintained in broker/README.md.
Native VR overlay documentation is maintained in native/overlay/README.md.
Custom HTTP API extension documentation is maintained in docs/http-extensions.md. For the JSON Schema required for connection, see docs/http-extension.schema.json.
VRChat OSC controls are documented in docs/vrchat-osc.md.
RICHARDwuxiaofei fzcfweasdferttgg-png
SUI_32C, Nagikokoro, motoka96, _Ykol魚, kascr_, Just Monika V, FLUVIA, Han โชเล่ย์, EA_PE, Ephedrine, ~ eri ~, fzcfweasdferttgg-png, Welcius, nunu299
Third-party licenses and notices: src/puripuly_heart/data/THIRD_PARTY_NOTICES.txt




































