Full songs with an editable score, generated on your own GPU. One executable — no Python, no Node.js, no launcher.
YuE2 Studio is a desktop studio for YuE2, the open song model from M-A-P that writes a score before it sings. Describe a style, write the lyrics, and the model composes a melody with chords as sheet music, then performs it as a full song with vocals. The score comes back with the track: read it, edit it, and render the same composition again with a different sound. Windows installer with auto-update or a portable folder, runs offline on an NVIDIA card with 6 GB of VRAM or more; AMD and Intel cards through Vulkan are experimental.
It is built on yue2.cpp, the native C++/CUDA port of YuE2. The studio around it is Rust and React in a Tauri window — nothing in the runtime path is Python.
Given this repository, an agent can set everything up and drive the studio by itself:
- Install the studio from the latest release and start it.
- Connect to its MCP server at
http://127.0.0.1:8791/mcp:claude mcp add --transport http yue2-studio http://127.0.0.1:8791/mcp - Read the skill it serves (resource
studio://skill, promptstudio), the same text as docs/mcp-skill.md, and start with the toolstudio_status.
llms.txt says the same for tools that look for it. To keep the skill in Claude
Code, save docs/mcp-skill.md as ~/.claude/skills/yue2-studio/SKILL.md.
- Full songs from a style and lyrics — up to six minutes, in the languages the model sings. On an RTX 4090 with the Q8_0 set a 3:38 song renders in about 46 seconds.
- Read and edit the score — the model writes its composition in ABC notation first; the studio engraves it as sheet music. Edit the notes, tempo or key and create again: the composition stays, the performance changes. Or switch to melody-only, or no score at all.
- Compose the score first — write only the score from the style and lyrics, without
singing it (the studio's take on yue2.cpp's
yue-plan), read and fix it, then create. - Covers — SheetSage2 listens to a recording and writes its melody as a score; YuE2 then sings that melody in your style. The words have to fit the melody: the original lyrics work best, new ones need the same syllables line by line with the stresses on the same notes, or the singing drifts off the tune — lines stretch, pause and slide into the wrong section. Transcription mistakes carry into the cover too. A track from your library can be transcribed from its menu.
- Exact replay — every track keeps its request and its audio codes, so it can be re-rendered bit for bit, or re-rendered with other steps, a new sound seed, several variations, or another output format, without composing again.
- 110 ready examples — the style/lyrics/score sets that ship with yue2.cpp, covers included, one click to load.
- Every engine setting — all seven sampling knobs (temperature, top-p, top-k,
repetition penalty and its window, minimum and maximum tokens) for the score and for the
audio codes, flow-matching steps, guidance, both seeds, peak normalisation, MP3 or
16/24/32-bit WAV. Prompts open and save as JSON or YAML in the engine's own request
format, so they move freely between the studio, the yue2.cpp WebUI and
yue-synth. - A writing assistant — a local Gemma model or OpenRouter writes the style and lyrics from an idea and edits the score on request; or pick your connected agent (MCP) and it writes instead. Lyrics you wrote yourself get their section tags with one button, the words exactly as you wrote them.
- Word-level karaoke — enhanced LRC with a timestamp on every word, aligned by Parakeet or Whisper. Your lyrics are kept; only the timing is borrowed.
- Six stems on the GPU — drums, bass, other, vocals, guitar and piano with HT-Demucs.
- Any track to MIDI — a song, a stem or a processed take becomes multi-instrument MIDI (34 instrument groups and drums) with MuScriptor on the GPU, through HOT-Step's native port. A piano roll fills in while it listens; play it against the original, mute or solo an instrument, save the .mid. Downloaded the first time it is used.
- Every result is a track — stems, a processed take, a re-render and a cover land in the library as tracks of their own, each marked with the one it was made from and keeping the settings it was made with.
- LoRA — LoRA and LoKr files for either half of the model, the composition or the sound, each with its own strength, picked in the create form. A catalogue of ready ones with their authors credited, and a search on Hugging Face that downloads what you pick.
- Train your own LoRA — an optional tab on the LoRA page: songs of one artist or
style become a LoRA on your own card with HOT-Step's trainer and its tuned recipe (LoKr,
Prodigy, lyric timing from a forced aligner that also reads Cyrillic). Every setting of
the recipe is editable. Datasets move between this studio and MiniMax Music3 Studio as a
folder, and each saved checkpoint goes into the LoRA library in one click.
- A three-step wizard — drop a folder of songs, check them, train. Albums with a cue sheet are cut into songs; titles and artists come from the tags, the file name and the folders.
- Preparation on its own — lyrics come from the lyric databases players use (LRCLIB, QQ Music, Kugou), word for word; only a song none of them knows has its vocals separated and is heard by Whisper, with its usual hallucinations filtered out. The assistant lays the lyrics out in sections. MOSS-Music listens to every song and the assistant writes its style, with the tempo measured by Beat This! on the card.
- Every song shows where it is — the lyrics and styles appear as each song is done, one model on the card at a time, each loaded once for the whole batch. A song that failed or was stopped has a button that finishes just that song.
- Picks up after a restart — each song keeps what it has; a preparation cut off by a crash or a restart carries on by itself and redoes nothing.
- Stop by drift or by epochs — stop when the composition half drifts too far from the base, or after a set number of passes over the songs.
- A trigger word from the start — every dataset gets a rare word made from its name, which you can change.
- Train further — not there yet at 750 steps? Set 1000 and the run goes on from its latest checkpoint with the same recipe and songs; the loss chart and the checkpoints continue instead of starting over.
- Audio processing — noise reduction, the Spectral Lifter, a vocal naturaliser, your own VST3 plugins in a chain, and mastering to a reference track. Compare before and after while it plays, then keep the result as a version of the track or throw it away.
- MP3 made by the studio — the engine renders 32-bit float and the studio encodes the MP3 with LAME, so nothing is lost before the encoder.
- A library of plain files — search, playlists, cover art from prompt templates, MP3s exported with title, lyrics and cover in their ID3 tags. Interface in English, Russian, Chinese, Japanese and Korean.
The same screens in the language you read: Русский, 中文, 日本語, 한국어 — on the project page, or in docs/screenshots.
Made in the released build on an RTX 4090 with the recommended BF16 set, nothing edited afterwards. Listen in the browser on the project page or download the MP3s from docs/samples.
| Song | How it was made | Style |
|---|---|---|
| Houseplant Party | English funk disco | funk disco, slap bass, wah guitar, brass, 118 BPM |
| My Truck Talks Back | Country | country, male baritone, pedal steel, fiddle, 96 BPM |
| La Siesta | Latin pop in Spanish | latin pop, nylon guitar, congas, piano montuno, 102 BPM |
| 外卖小哥 | Mandarin city pop | city pop, electric piano, funky bass, 108 BPM |
| 月曜日のサムライ | Japanese rock, anime-opening style | j-rock, female voice, distorted guitars, 168 BPM |
| Дача, лето, комары | A full song from a style and lyrics | folk pop, warm male voice, accordion, 104 BPM |
| Дача, лето, комары — джаз | Cover: SheetSage2 wrote down the melody of the song above, YuE2 sang it as jazz | vintage jazz swing, sultry female crooner, muted trumpet |
| Понедельник | Blues rock, raspy voice | blues rock, overdriven guitar, hammond organ, 96 BPM |
| Бабушка на дискотеке | Re-rendered from its saved audio codes with 64 solver steps | disco pop, funky bass, string section, 120 BPM |
| Борщ на орбите | Ska punk | ska punk, brass section, offbeat guitar, 150 BPM |
| Кот-программист | Synthwave pop, the second take | synthwave pop, analog synths, drum machine, 112 BPM |
| Ночной вор Барсик | Style and lyrics written by the built-in assistant from a one-line idea | pop-punk, distorted guitars, 185 BPM |
- Windows 10/11 x64.
- A GPU with 6 GB of VRAM or more:
- NVIDIA, from the GTX 900 series on, runs on CUDA — the fastest path. The studio ships two CUDA builds of the engine and picks the one your card and driver run: CUDA 13 for Turing and newer (GTX 16, RTX 20–50, Tesla T4, A100, RTX A-series, L4/L40, H100) with driver 580 or newer, CUDA 12 for Maxwell, Pascal and Volta (GTX 900/1000, Titan X/Xp/V, Tesla M40, P40, P100, V100) and for any card on a driver from 525 to 579. Both carry compiled code for every one of those architectures, so nothing is left for the driver to compile.
- AMD or Intel — experimental. The engine runs on Vulkan through the card's own driver. The Vulkan path itself is verified on NVIDIA (the same words heard as on CUDA), but on AMD Radeon integrated graphics the song came out with unintelligible vocals, and discrete AMD and Intel cards are untested. Reports from owners are welcome.
- Without a GPU the engine falls back to the processor, which works but is many times slower.
- 4–10 GB of disk for one model set.
- Training a LoRA (optional): an NVIDIA RTX 30-series card or newer with 11 GB of VRAM and about 8 GB more disk for the trainer and its weights, downloaded only when you open training. Describing songs by ear needs about 12 GB of VRAM and 10 GB more disk for MOSS-Music; without it the styles are written by hand.
- Install — run
YuE2.Studio_x.y.z_x64-setup.exefrom the latest release, or unzip the portable archive anywhere and runYuE2-Studio.exe. - Choose a model set — the first screen preselects the set your card can run. Press download; it fetches only what is missing and resumes if interrupted.
- Create — write a style and lyrics, or load one of the examples, and press Create. The engine starts by itself and the song lands in your library with its score.
Everything the studio owns stays in its own folder: models, songs, settings, logs,
temporary files and the WebView2 profile, beside YuE2-Studio.exe. That holds for the
portable archive and for an installation into any folder the studio can write to;
deleting the folder removes the studio. Only an installation into a read-only location
such as Program Files falls back to %LOCALAPPDATA%\YuE2 Studio. The installed version
updates itself: a new release is offered inside the studio and installed in place.
While the studio is open it serves MCP at http://127.0.0.1:8791/mcp: an agent such as
Claude Code, Claude Desktop or Cursor does everything the page does, through the same code -
songs and scores, the library, covers, stems, MIDI, karaoke, processing, video clips, the player,
LoRA, and a LoRA from a folder of songs end to end - and sees and works the window itself:
a screenshot, its controls, clicks and typing. 162 tools, grouped by area. The model's
writing rules and official examples come with the server, so the agent writes the styles,
lyrics and lyric layouts itself instead of the studio's small assistant.
With Agent (MCP) chosen as the writing assistant (Settings, Models), the connected agent also answers the studio's own write buttons and dataset preparation. Settings, Agent (MCP) shows whether an agent is connected and what to paste into the client.
claude mcp add --transport http yue2-studio http://127.0.0.1:8791/mcpdocs/mcp-skill.md is the skill an agent reads: every tool, what the model expects, and step-by-step recipes.
A runnable YuE2 installation is a backbone (the 3B model that writes the score and the audio codes) and the VAE (turns them into 48 kHz stereo). SheetSage2 is optional: without it the studio generates, but cannot transcribe recordings for covers.
| Your GPU | Set | Download |
|---|---|---|
| 12 GB VRAM and above | Full native — BF16 backbone, original weights | 9.7 GB |
| 8 GB and above | Quality — Q8_0 backbone, near lossless | 4.9 GB |
| 7 GB and above | Balanced — Q6_K backbone | 4.0 GB |
| 5.5 GB and above | Light — Q5_K_M backbone | 3.6 GB |
Sizes include SheetSage2 at the matching quantisation. The studio detects your card and preselects the set, but the download is always your decision; the model manager also builds a custom mix role by role. Q5_K_M is the lightest quantisation published for YuE2.
All files come from Serveurperso/YuE2-GGUF,
pinned to revision 64b030e, and are checked by size and SHA-256. They are written to, and
can be dropped into by hand at:
<the studio's folder>\data\models\yue2-cpp\— the portable folder, or the folder you installed into%LOCALAPPDATA%\YuE2 Studio\models\yue2-cpp\— only for an installation into a read-only location
A file placed by hand with the exact catalogue name is recognised and never downloaded again.
Every file, with direct download links
| File | Role | Size |
|---|---|---|
YuE2-3B-BF16.gguf |
backbone | 6.67 GB |
YuE2-3B-Q8_0.gguf |
backbone | 3.55 GB |
YuE2-3B-Q6_K.gguf |
backbone | 2.74 GB |
YuE2-3B-Q5_K_M.gguf |
backbone | 2.44 GB |
YuE2-Vae-F32.gguf |
VAE, every set | 506 MB |
SheetSage2-F32.gguf |
transcriber | 2.52 GB |
SheetSage2-Q8_0.gguf |
transcriber | 913 MB |
SheetSage2-Q6_K.gguf |
transcriber | 776 MB |
SheetSage2-Q5_K_M.gguf |
transcriber | 703 MB |
Each part is downloaded when you first use it, and each file can be downloaded by hand
and put into the studio's data folder — <the studio's folder>\data\, or
%LOCALAPPDATA%\YuE2 Studio\ for an installation into a read-only location. A file with
the exact name in the listed folder is recognised and not downloaded again.
Training, listening, lyrics, stems and the assistant — direct links
Training a LoRA — scragnog/YuE2-GGUF, pinned to eb7de09
| File | What for | Size | Put in data\ |
|---|---|---|---|
yue2_3b_int8_convrot.safetensors |
the model the LoRA is trained on | 3.69 GB | training\models\ |
yue2-vae-standard-f32.gguf |
audio to latents | 506 MB | training\models\ |
yue2-tok-f16.gguf |
audio to codes | 1.13 GB | training\models\ |
sheetsage2-f16.gguf |
the songs' scores | 1.27 GB | training\models\ |
mms-fa-f32.gguf |
lyric timing | 1.18 GB | training\models\ |
Describing songs by ear (optional)
| File | What for | Size | Put in data\ |
|---|---|---|---|
moss-aud-f16.gguf |
MOSS-Music-8B, the ear | 1.61 GB | training\models\moss\ |
moss-lm-q8_0.gguf |
MOSS-Music-8B, the words | 8.11 GB | training\models\moss\ |
beat_this.onnx |
Beat This!, the tempo | 79 MB | training\models\audio-facts\ |
Lyrics by ear — only for songs no lyric database knows; pick one recogniser
| Files | What for | Size | Put in data\ |
|---|---|---|---|
faster-whisper-large-v3: config.json, model.bin, preprocessor_config.json, tokenizer.json, vocabulary.json |
Whisper large-v3 | 2.9 GB | karaoke\models\whisper\faster-whisper-large-v3\ |
parakeet-tdt-0.6b-v3-onnx: config.json, encoder-model.int8.onnx, decoder_joint-model.int8.onnx, nemo128.onnx, vocab.txt |
Parakeet, European languages | 0.7 GB | karaoke\models\parakeet\ |
Stems and vocals for recognition
| File | What for | Size | Put in data\ |
|---|---|---|---|
htdemucs_6s_fp16weights.onnx |
HT-Demucs, six stems | 130 MB | separation\models\htdemucs\htdemucs_6s_fp16.onnx (this name) |
MIDI from audio — the transcriber and one of the models
| File | What for | Size | Put in data\ |
|---|---|---|---|
music-midi-cuda-windows-x64.zip |
HOT-Step's ace-midi, unpacked |
126 MB | midi\runtime\music-midi\ |
muscriptor-small: config.json, model.safetensors |
fast, 103M | 0.4 GB | midi\models\muscriptor-small\ |
muscriptor-medium: config.json, model.safetensors |
balanced, 307M | 1.2 GB | midi\models\muscriptor-medium\ |
muscriptor-large: config.json, model.safetensors |
best, 1.4B | 5.5 GB | midi\models\muscriptor-large\ |
MuScriptor by Kyutai & Mirelo, weights CC BY-NC 4.0 (non-commercial); the mirror carries the official files byte for byte, without the Hugging Face sign-in.
The writing assistant — one of
| File | Size | Put in data\ |
|---|---|---|
gemma-4-E4B_q4_0-it.gguf |
4.80 GB | assistant\models\ |
gemma-4-12b-it-qat-q4_0.gguf |
6.50 GB | assistant\models\ |
yue-server.exe → ggml.dll, ggml-base.dll shipped inside the app
loads at run time, whichever the machine can use:
├─ ggml-cuda.dll → cublas64_13.dll, cublasLt64_13.dll (downloaded once), nvcuda.dll (NVIDIA driver)
├─ ggml-vulkan.dll → vulkan-1.dll (every AMD, Intel and NVIDIA driver)
└─ ggml-cpu-*.dll nine builds, from SSE4.2 to AVX-512; the best one for the processor is picked
+ msvcp140, vcruntime140, vcruntime140_1, vcomp140 Visual C++ runtime, shipped app-local
Shipped inside the app. yue-server.exe, every ggml*.dll, built from the pinned
yue2.cpp commit with all backends, and the Visual C++ runtime it needs live in
resources\yue2-cpp\ beside the main executable. Nothing is installed into Windows. Settings → Local engine chooses the compute device: Auto (CUDA on NVIDIA,
Vulkan on AMD and Intel, the processor without a GPU), or CUDA, Vulkan or the processor
explicitly. On Vulkan the studio turns on the engine's FP16 clamp: without it AMD Radeon
integrated graphics rendered pure silence and the engine crashed encoding it.
Downloaded once, on the first engine start — on NVIDIA only.
| File(s) | Where from | Size | Why |
|---|---|---|---|
cublas64_13.dll, cublasLt64_13.dll |
NVIDIA's redistributable libcublas-windows-x86_64-13.5.1.27-archive.zip |
391 MB (zip) | The CUDA linear algebra ggml-cuda.dll is linked against; too large, and under NVIDIA's licence, to bundle. |
A machine with the CUDA 13 toolkit already has cuBLAS on its PATH and downloads nothing.
Behind a proxy, take the two DLLs from the archive's bin\ folder and drop them next to
yue-server.exe; the studio finds and uses them.
React UI ─┐
├─ YuE2-Studio.exe (Tauri window + Rust/Axum service on 127.0.0.1:8791)
Rust axum ┘ │
├─ yue2.cpp `yue-server` (C++/CUDA, GGUF, 127.0.0.1:18087)
├─ music-train.exe (HOT-Step ace-train, LoRA training, optional)
└─ vst-host.exe (HOT-Step VST3 host, a process of its own)
The service is compiled into the desktop binary. It supervises the engine process, restarts it when you switch model sets, and imports finished songs itself, so a result is never lost if the window was reloaded or closed mid-generation. Progress is read from the engine's own log: score, audio codes, acoustic rendering, decoding.
npm --prefix app install
npm --prefix desktop install
cargo test --workspace
npm --prefix app testDeveloping the UI against a running service:
cargo run -p music-server # service on 127.0.0.1:8791
npm --prefix app run dev # UI on 127.0.0.1:3791The engine: scripts/build-yue-runtime.ps1 -RuntimeBackend all builds the pinned yue2.cpp
commit (engines/yue2-cpp-source.json) with runtime-loaded CUDA, Vulkan and CPU backends,
using CUDA 13, the Vulkan SDK, MSVC and Ninja. scripts/build-release.ps1 -Version X.Y.Z produces the NSIS installer, the
portable archive and the signed latest.json for the updater; it reads the signing key from
TAURI_SIGNING_PRIVATE_KEY or %USERPROFILE%\.tauri\yue2-studio.key. Model weights are
never part of a release.
yue2.cpp also builds for Linux and macOS (Metal); the studio's release pipeline ships the Windows build with CUDA, Vulkan and CPU backends for now.
Other Projects by @timoncool
| Project | Description |
|---|---|
| MiniMax Music3 Studio | The same studio on MiniMax Music3 — the one this grew out of |
| ACE-Step Studio | AI music studio — songs, vocals, covers, videos |
| Foundation Music Lab | Music generation + timeline editor |
| VibeVoice ASR | Portable speech recognition |
| Qwen3-TTS | Portable text-to-speech with voice cloning |
| telegram-api-mcp | Full Telegram Bot API as an MCP server |
- Nerual Dreming — Telegram | neuro-cartel.com | ArtGeneration.me
- Нейро-Софт — Telegram | portable neural networks
- M-A-P for YuE2-3B, the YuE2 VAE and SheetSage2.
- Serveurperso for yue2.cpp, its examples and the GGUF conversions.
- scragnog for HOT-Step-CPP: the LoRA trainer the studio runs (its native joint AR/NAR training for YuE2), the training weights in scragnog/YuE2-GGUF, the VST3 host, and the noise reduction, Spectral Lifter and mastering designs the studio's audio processing is ported from.
- sergree for matchering, the reference mastering algorithm, and jeankassio for the vocal naturalizer in ComfyUI_MusicTools.
- The authors of the LoRA in the catalogue, each credited and linked on its card: Mothersuperior, monsterovich, atomtanstudio, HaileyStorm and ntc-ai.
- The LAME project for the MP3 encoder.
I build open-source software and do AI research. Most of what I create is free and available to everyone. Your donations help me keep creating without worrying about where the next meal comes from =)
All donation methods · Русский · 中文 · 日本語 · 한국어 | dalink.to/nerual_dreming | boosty.to/neuro_art
- BTC:
1E7dHL22RpyhJGVpcvKdbyZgksSYkYeEBC - ETH (ERC20):
0xb5db65adf478983186d4897ba92fe2c25c594a0c - USDT (TRC20):
TQST9Lp2TjK6FiVkn4fwfGUee7NmkxEE7C
The studio is MIT, and so is yue2.cpp. The models are not: YuE2-3B, the YuE2 VAE and SheetSage2 are released under CC BY-NC 4.0 — songs you make with them are for non-commercial use unless you obtain other terms from their authors.
What changed and when is in CHANGELOG.md.










