Hum-to-Song on YuE2: how it works
Hum a melody for 10 to 30 seconds, add a style line and lyrics, get a finished song that keeps your melody, builds a structure around it, and continues long after your hum stops. This document describes the two mechanisms behind it, how they were trained, and how to run them with the scripts here or inside ComfyUI.
The problem, and why it splits in two
YuE2 generates a song in two stages. The AR planner first writes a symbolic ABC score (its chain-of-thought, cot="melody" or "full"), then writes 25 Hz semantic tokens conditioned on that score. The NAR decoder then turns tokens into VAE latents by flow matching, and the VAE decodes 48 kHz audio.
That split decides where a hum has to enter. Melody is decided by the planner, not by the decoder: once the semantic tokens exist, the decoder can shape timing, articulation and timbre, but it cannot change which notes are sung. So a hum module needs two parts:
- Score continuation puts the hum's melody into the planner, where the melody is chosen. This is the part that makes the song be about your melody. It needs no training.
- A prosody adapter puts the hum's actual pitch contour and timing into the decoder, so the sung line follows how you hummed it, not just what you hummed. This is the trained part.
flowchart LR
H[hum audio] --> T[melody transcriber -> ABC]
T --> P["open score prefix<br/>[ABC_START] + hum ABC (no end token)"]
S[style + lyrics] --> P
P --> AR1[AR: continue the score]
AR1 --> AR2[AR: semantic tokens from the full score]
H --> C[pitch tracker -> sine carrier -> VAE encode]
AR2 --> NAR[NAR decoder + prosody adapter]
C -->|added at layers 0,7,14,21| NAR
NAR --> VAE[VAE decode] --> A[song]
Stage 1: score continuation (no training)
Transcribe the hum. A melody transcriber (SheetSage2, melody-only mode) turns the hum into ABC in YuE2's own two-voice score format: a Vocal melody line and an Ins line, key, meter, tempo. A 20 s hum becomes roughly 200 to 250 score tokens. Trailing bars that are only rests are stripped, so the score does not end where the hum ended.
Leave the score open. YuE2's prompt for a provided score is [EOD] text(instructions, style, lyrics) [ABC_START] score [ABC_END] [MUSIC_START]. We build the same prompt but stop after the hum's score tokens, with no ABC_END. The planner is then simply asked to keep writing. Because it was pretrained to write complete scores after [ABC_START], it treats the hum's bars as the opening of a song and writes the rest: new sections, transitions, an ending, all in the same key, meter and range. On our tests the hummed phrase recurred verbatim two to four times later in the continued score, typically as the hook.
Then the normal path. The full score (hum bars plus continuation) is closed with [ABC_END] [MUSIC_START], the planner writes semantic tokens for the whole song, and the decoder renders them. Song length comes from the score the planner wrote, so 20 s of hum yields a 2 to 3 minute song with an ending.
Three melody modes fall out of this: continue from hum (the above), hum only (close the score right after the hum, so the song is exactly your melody), and ignore hum (let the planner write its own score; the adapter below can still shape phrasing).
Stage 2: the prosody adapter (trained)
The conditioning signal
The decoder is not shown the hum itself. It is shown a deliberately reduced version that carries melody and timing and nothing else:
- pitch is tracked with pYIN over 65 to 1000 Hz (hop 512 at 48 kHz), gaps forward-filled, interpolated to the sample grid;
- a sine wave is synthesized at that pitch;
- its amplitude follows the rectified hum, low-passed with a 4th-order 30 Hz and a 2nd-order 80 Hz Butterworth filter, which smears syllables away;
- the carrier is encoded by YuE2's own VAE into 25 Hz latents,
[T, 64].
Because the decoder only ever sees sine tones, it cannot depend on the singer's timbre or words. That is what lets a training set built from real vocal stems generalize to a person humming into a laptop microphone.
Where it enters the decoder
Four zero-initialized Linear(64 → 2048) projections take the (edge-padded) carrier latents and add them to the decoder's hidden state: one at the input (alongside vae2llm(x_t), the time embedding and positional embedding) and one each before layers 7, 14 and 21 of the 28 layers. Frames with no hum carry zeros, which the projections map to a learned bias, so "no hum here" is an explicit state rather than an absence.
Trainable parameters: the four projections, a rank-64 LoRA on the decoder's attention and MLP projections (nar_self_attn.{q,k,v,o}_proj, nar_mlp.{gate,up,down}_proj), and the two full vae2llm / llm2vae layers. The adapter was trained on top of our real-audio decoder LoRA (folded into the base first); the combined release merges both into one rank-96 LoRA so it applies to the stock model in one step.
Training pairs
Pairs are self-supervised from ordinary songs with separated vocals: the condition is the sine carrier of the vocal stem, the target is the full mix's VAE latents, and the semantic tokens in context are the mix's own tokens from our audio-to-token head (the same kind of tokens the planner produces at inference). 3,163 training pairs, 188 held out. Loss is the standard flow-matching objective on 30 s windows: MSE(v_pred, noise − x1) with x_t = t·noise + (1 − t)·x1, the whole song's tokens in the AR cache.
The partial-hum rule
Real use is "hum a bit, get a whole song", so half of the training draws see only part of the carrier: for eligible songs the hum is kept only from 2 s before the vocal onset for a random 12 to 28 s and zeroed everywhere else. Onset is detected from the pitch tracker's voiced flag combined with energy (sustained 0.4 s). Training windows are biased 60/40 toward the hum's edges so the model often sees the handoff from "follow the hum" to "continue on your own". Vocal onset was found for 3,304 of the 3,351 clips; median onset 7 s, so trimming the dead air mattered.
10% of draws zero the whole carrier, which enables classifier-free guidance on the hum channel at inference: v = v_no_hum + g · (v_hum − v_no_hum), exposed as hum influence.
Numbers
12,000 steps, held-out clips:
| step 0 | step 12,000 | |
|---|---|---|
| with hum | 0.960 | 0.906 |
| no hum (carrier zeroed) | 0.960 | 0.913 |
| partial hum (fixed 20 s from onset) | 0.901 | 0.871 |
The with-hum advantage appears late (crossover near step 5,000) and stays modest. That is expected: the semantic tokens in context already fix the melody, so the carrier can only add timing and pitch nuance. This is exactly why melody had to go into the planner (stage 1) and why the adapter is the finishing layer, not the melody carrier.
Inference outside ComfyUI
# stage 1: hum -> ABC (melody-only transcription), then open-score continuation + song
python scripts/hum_continue.py hum_score.abc style.txt lyrics.txt out_tag <seed> [nar_lora]
# stage 2: render the same song's tokens through the prosody adapter with the hum's carrier
VOICE_CFG=1.0 HUM_OFFSET_S=0 python scripts/infer_hum.py hum_adapter_v1.pt hum.wav out_tag_score.abc out_tag_tokens.npy style.txt lyrics.txt out.flac <seed>
hum_continue.py builds the open prefix with token_prefixes(request, tok) + hum_abc_ids, calls the pipeline's ABC generator on it, then runs the stock semantic and decoder path. infer_hum.py folds the adapter into the decoder, places the carrier latents at HUM_OFFSET_S, and integrates the flow ODE with hum-channel guidance. hum_prep.py and train_hum.py reproduce the dataset and training.
Inference in ComfyUI
The ComfyUI-HumSong node pack (ComfyUI-HumSong/) implements both stages natively with four nodes and no stock nodes: HumSong Loader (checkpoint, hum adapter, melody transcriber), Hum Input (file or in-node microphone recording), Hum to Song (style, lyrics, seed, length cap, hum influence, melody mode), HumSong Output (FLAC, player, score view with the hummed bars highlighted). Internally the loader applies the adapter's decoder LoRA through ComfyUI's own LoRA machinery (humsong_yue2_adapter_v1_comfy.safetensors, fused-projection layout) and keeps the four injection projections; the song node wraps the decoder's forward to add the carrier at the four depths and to apply hum-channel guidance, and uses ComfyUI's built-in SheetSage2 for transcription.
Weights
| file | use |
|---|---|
hum_adapter_v1.safetensors |
adapter as trained; fold nar_lora_joint_v4 into the base first |
hum_adapter_v1_combined.safetensors |
self-contained against stock YuE2-3B (rank-96 LoRA + projections) |
humsong_yue2_adapter_v1_comfy.safetensors |
ComfyUI layout for the node pack |
hum_adapter_v1.pt |
training checkpoint |
minted_regularizer_pack_v2.pt |
regularizer for training planner LoRAs (FS_Audio Suite trainer): 12,247 YuE2 self-generated songs as semantic tokens + style + lyrics + ABC score text, no audio. Coin-flipped against your artist so the LoRA keeps the base model's range. |
Limits worth knowing
The transcriber needs something voice-like; it does not transcribe pure sine tones, so a real hum works and a synthesized tone does not. The hum sets the melody and the opening, not the total length; the planner decides where the song ends. The adapter's effect is subtle by design and best judged in the first hummed section. Weights derive from YuE2-3B (CC BY-NC 4.0): non-commercial use only.
Model tree for Mothersuperior/YuE2-hum-to-song
Base model
m-a-p/YuE2-3B