YouTube Speech to Song vs Sume Music Router: prompt in, audio out

YouTube's Speech to Song turns a Short's speech into a song with Lyria 2. Sume's Music Router takes a text prompt (and an optional image), with no audio input.

4 min readSume
All posts

Speech to Song is a YouTube Shorts feature that uses Lyria 2 to turn speech from an existing video into a song, with attribution to the original creator. Sume's Music Router does not do that: its request takes a text prompt of 1 to 5000 characters (plus an optional image URL) and rejects duration fields, with no audio input. You can generate a track from a prompt, then lay it under a clip with a Timeline soundtrack.

What each side does

YouTube's announcement post, read 2026-10-04, says Speech to Song uses Lyria 2, that the original creators are attributed, and that SynthID and labels apply to outputs. The Music Router docs list model ids sume/music-auto, lyria-3.5 and lyria-3-pro; the Lyria generation there is newer than the one YouTube names.

Speech to Song against Sume Music Router (YouTube blog read 2026-10-04; Sume from docs)
QuestionYouTube Speech to SongSume Music Router
InputSpeech from a ShortText prompt, 1-5000 characters, optional image_url
Audio inYes, inside YouTubeNo
Model namedLyria 2sume/music-auto, lyria-3.5, lyria-3-pro
Where it runsShorts in the appAPI

The workflow Sume supports

Generate a track from a prompt, then mix it in Timeline 1.0 with the soundtrack object: url, gain_db, loop, fade_out_seconds up to 10 and duck_db 0 to 20 so speech stays clear. Speech to Song's remix of someone else's voice is not offered; do not clone a voice you have no right to.

Sources

Related posts

More in Models

All Models posts

Written by Sume