StepAudio 3 Music from a dry vocal or reference audio vs Sume
StepAudio 3 Music accepts lyrics, vocals or reference audio, even scoring a dry vocal. Sume Music takes a text prompt and one optional image, no audio.

If your query is "generate an accompaniment for my a cappella vocal" or "make a song from a reference track", the StepFun page for StepAudio 3 Music lists those inputs: natural language, lyrics, vocals or reference audio, with "intelligent scoring from dry vocals" as a named capability. Sume Music does not take audio as input. Its request is a text prompt (1 to 5000 characters) plus an optional public HTTPS image_url, per the Music 1.0 docs and the Music Router.
What can StepAudio 3 Music take as input?
The model page lists lyrics to song, instrumental generation, scoring from dry vocals, natural-language control and reference-audio-driven generation. It marks ABC notation input as coming soon. For voice cloning, covers and reference-audio generation it asks you to obtain lawful authorization from the voice subject and the rights holders of the lyrics, composition and recording, and to complete copyright and content-safety review before public or commercial use.
| Input | StepAudio 3 Music page | Sume Music Router |
|---|---|---|
| Text description | Yes | Yes, prompt up to 5000 characters |
| Lyrics with section tags | Yes | Put lyrics and markers inside prompt |
| Dry vocal or reference audio | Yes | No audio input |
| Image | Not listed | Optional image_url, public HTTPS |
| Negative prompt | Not stated | Non-empty value returns 400 negative_prompt_unsupported |
What do I do on Sume instead?
Describe the arrangement in words: genre, tempo as a number, key, two to four instruments with texture, and one named arc moment, as the Music 1.0 docs suggest. The result is a new track, not a track built around your vocal, so line it up with the vocal yourself. If you already have a finished vocal and a generated bed, a Timeline render can mix them (the soundtrack bed sits under an audio spine); that is a mix, not a composition that follows your melody.
Is an image a reasonable substitute for a reference track?
Only for mood. Sume's docs say image_url conditions the music visually; use the accepted scene still when appropriate. It will not carry a melody or a timbre. Read Image to music AI for how that input behaves.
Sources
Related posts
More in Comparisons
- Synthesia burned-in captions on dubbed videos, and Sume captions
Synthesia added one-click burned-in captions to dubbed videos on 9/30/2026. Here is what that means, and how to burn captions onto a finished video URL on Sume.
- Which Synthesia plan includes API access? Pro is limited
Synthesia lists no API on Basic or Starter, a limited API with 360 minutes a year on Pro, and full access on Enterprise. What a Sume API key gives you instead.
- Synthesia video quizzes are Enterprise-only; what Sume offers instead
Synthesia adds scored quiz questions and a pass threshold to videos on Enterprise only. Sume has no quiz component, only avatar clips you assemble yourself.
- Synthesia voice clone: consent passcode and 1-5 minute upload vs Sume
Synthesia clones a voice from a recording or a 1-5 minute upload and a spoken consent passcode. Sume has no voice-clone route; here is the practical gap.
Written by Sume