YuE2 generates full songs by first writing a readable score in text-based notation, then filling in semantic and audio detail, all inside one model. Symbolic planning wins expert preference 49.3% vs 34.6% over skipping the score.
Modern song generators come in two flavors that don’t talk to each other. Symbolic models like Music Transformer output notation you can read and edit, but stop before producing a finished recording. Audio models like MusicLM produce polished recordings but never expose the composition. If you want to change a chord in the chorus or move the song to a new key, you’re stuck: the audio model has no notion of “chord” you can grab, and the symbolic model can’t render a convincing vocal performance.
The practical scenario: imagine you generate a song and love the arrangement but hate one melodic phrase. Today with a Suno-style tool, you either regenerate the whole thing and hope for the best, or you accept it. YuE2 argues the fix is to make the model commit to melody, harmony, key, meter, and form as readable notation first, then perform that notation as audio. The score becomes both a quality-improving planning step and an editing surface.
The generation runs in three stages inside one 28-layer backbone. First, the model writes a symbolic score s in ABC notation, serialized with byte-pair encoding. Second, it produces c, a stream of discrete semantic music tokens at 25 Hz that carry musical detail a lead sheet can’t express (timbre hints, expressive nuance). Third, it produces z, continuous acoustic latents at 25 Hz that a separate 48-kHz stereo Variational Autoencoder decodes to waveform.
The backbone is an AR-NAR Mixture-of-Transformers. Stages s and c are autoregressive (causal next-token prediction over discrete tokens). Stage z is non-autoregressive and trained with Flow matching, so every acoustic frame updates in parallel and can attend bidirectionally to the full plan. Inside each layer, AR and NAR streams have separate normalization, projections, and MLPs but share one attention operation with a hybrid mask: discrete positions can’t peek at acoustic targets, acoustic positions can read everything.
Getting training targets is the hard part, because ordinary recordings don’t come with aligned scores. The paper introduces two supervisors. SheetSage2 transcribes recordings into readable lead sheets (melody, chords, key, beats, sections). MERT2 is a music representation encoder whose quantized branch produces the 25 Hz semantic tokens at 375 bits/s. Both are trained first, then frozen and used to build training pairs from raw audio.
# YuE2 generation, sketched
s = ar_decode(text_and_lyrics) # write ABC score
c = ar_decode(text_and_lyrics, s) # semantic tokens, 25 Hz
z = nar_flow_match(text_and_lyrics, s, c) # acoustic latents
waveform = vae_decoder(z) # 48 kHz stereo
# Editing: user supplies s_tilde, model regenerates c and z
# Covering: SheetSage2 extracts s_hat from a reference recording
One checkpoint handles creation, editing, and covers because training mixes four tasks: full pipeline, no score, no semantic tokens, and direct audio. The same model can therefore be asked to plan or skip planning, which is exactly what the ablation exploits.
Symbolic planning helps, tested cleanly. Same checkpoint, same prompts, same decoder, same candidate budget. With planning, 49.3% of expert judgments prefer the planned song for overall quality vs 34.6% for the unplanned version (rest are ties). Musicality shows a similar gap (45.0% vs 29.4%). Both are significant with clustered standard errors.
Unified beats separated. With planning held constant on both sides, experts prefer YuE2’s single Mixture-of-Transformers to a separate language model plus diffusion Transformer setup (the LM+DiT design used by several recent song generators). Overall quality: 53.4% vs 35.6%. Audio quality: 48.5% vs 29.6%.
Frontier full-song quality. On WildSongBench, YuE2 (best-of-8) reaches 6.96 on SongBench Global Avg, the highest observed among all evaluated systems including proprietary ones. Against Suno v4.5, experts prefer YuE2 (best-of-8) 57.3% to 30.5% for overall quality; against Suno v5 preferences are nearly balanced at 40.4% vs 39.9%. Suno v6 still receives more overall preferences than YuE2. Audio quality is a particular strength: mean tie-adjusted preference against six proprietary systems is 58.9%.
The score genuinely drives the audio. When YuE2 generates a score then a recording, transcribing the recording back and comparing to the score yields melody similarity of 0.9464 vs 0.2160 for a mismatched score from the same prompt. The audio actually realizes the planned notes and chords.
Editing is local. Changing chords in the first chorus achieves 79.5% target-chord agreement while preserving 93.4% of the vocal melody elsewhere. Melody edits hit 84.2% pitch accuracy with 93-94% preservation of unedited content.
Zero-shot covers work. Given a score transcribed by SheetSage2 from an unseen song, YuE2 renders it in a new style. On 948 works from SHS100K, YuE2 with full scores beats both evaluated cover systems on all eight work-identity retrieval measures, without any cover-specific training.
The supervisors are also strong. MERT2 surpasses previous best results on 14 of 15 Marble benchmark metrics. SheetSage2-AR leads 12 of 15 pairs in the paper’s full-song transcription comparison.
If you’re building or evaluating music generation, the mechanism-level takeaway is that committing to a readable plan before rendering audio measurably improves the finished result, even holding everything else fixed. This is a general pattern worth testing in other generation domains where the output has natural hierarchical structure (video with storyboards, code with type signatures, etc.).
If you want an editable song generator with an inspectable intermediate, the 3B YuE2 checkpoint and WildSongBench are released. The score interface is the editing hook: modify ABC notation, feed it as a prefix, regenerate. This is worth trying if your workflow involves iterating on musical content, not just re-rolling seeds.
For cover generation specifically, the finding that removing chords from the source score improves target-style alignment but hurts work-identity retrieval gives you a knob. If fidelity to the original composition matters more, keep the full score; if stylistic adaptation matters more, drop chords or let the language-model agent revise the score.
Caution on generality: the expert listening panel evaluated 192 WSB prompts with paid conservatory-trained annotators, not consumer preference at scale. The “unified beats separated” claim compares to one specific LM+DiT baseline matched in data volume, not to every possible two-model design. And best-of-8 selection uses SongBench Musicality plus Q3O plus PER as the ranking, so improvements at best-of-8 reflect both the model and the availability of a reasonable automatic ranker.
Experts still prefer Suno v6 overall, so “frontier” here means competitive with, not beating, the strongest proprietary system on every axis. HeartMuLa reports post-training with SongEval and AudioBox, which are among the automatic metrics reported, so its scores on those metrics should be read with that in mind (the paper flags this). The 346,000 hours of training music are described as primarily CC0 and synthetic, with details limited; the paper doesn’t fully specify the corpus composition. The zero-shot cover claim was verified with two cover-similarity models to check work-level overlap with training, which is a real check but not the same as guaranteed non-contamination. Finally, the editing evaluation regenerates the whole song from the edited score prefix; the paper measures preservation of unedited content statistically but this is not the same as surgical local re-synthesis.