StepAudio 3 Music from a dry vocal or reference audio vs Sume

StepAudio 3 Music accepts lyrics, vocals or reference audio, even scoring a dry vocal. Sume Music takes a text prompt and one optional image, no audio.

4 min readSume
All posts

If your query is "generate an accompaniment for my a cappella vocal" or "make a song from a reference track", the StepFun page for StepAudio 3 Music lists those inputs: natural language, lyrics, vocals or reference audio, with "intelligent scoring from dry vocals" as a named capability. Sume Music does not take audio as input. Its request is a text prompt (1 to 5000 characters) plus an optional public HTTPS image_url, per the Music 1.0 docs and the Music Router.

What can StepAudio 3 Music take as input?

The model page lists lyrics to song, instrumental generation, scoring from dry vocals, natural-language control and reference-audio-driven generation. It marks ABC notation input as coming soon. For voice cloning, covers and reference-audio generation it asks you to obtain lawful authorization from the voice subject and the rights holders of the lyrics, composition and recording, and to complete copyright and content-safety review before public or commercial use.

Inputs listed on the StepAudio 3 Music page vs Sume docs, read 2026-10-02.
InputStepAudio 3 Music pageSume Music Router
Text descriptionYesYes, prompt up to 5000 characters
Lyrics with section tagsYesPut lyrics and markers inside prompt
Dry vocal or reference audioYesNo audio input
ImageNot listedOptional image_url, public HTTPS
Negative promptNot statedNon-empty value returns 400 negative_prompt_unsupported

What do I do on Sume instead?

Describe the arrangement in words: genre, tempo as a number, key, two to four instruments with texture, and one named arc moment, as the Music 1.0 docs suggest. The result is a new track, not a track built around your vocal, so line it up with the vocal yourself. If you already have a finished vocal and a generated bed, a Timeline render can mix them (the soundtrack bed sits under an audio spine); that is a mix, not a composition that follows your melody.

Is an image a reasonable substitute for a reference track?

Only for mood. Sume's docs say image_url conditions the music visually; use the accepted scene still when appropriate. It will not carry a melody or a timbre. Read Image to music AI for how that input behaves.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume