StepAudio 3 Gen: voice, SFX and music in one clip vs Sume jobs
StepFun's stepaudio-3-gen-preview makes voice, effects, ambience and music in one audio output. Sume uses separate music, speech and Timeline mix steps.

StepFun's StepAudio 3 Gen turns one natural-language description into a single audio clip that can hold voices, sound effects, ambient sound, background music and singing, and the description can say where and in what order each element appears. Sume does not have a one-call equivalent. On Sume you make the pieces as separate jobs (music from the Music Router, speech from the speech surface) and combine them in a Timeline render, which has a soundtrack bed with gain, fade and duck. StepFun's side is from its StepAudio 3 Gen page.
What does StepAudio 3 Gen say it does?
The page lists model id stepaudio-3-gen-preview, inputs of text, audio, and role and script descriptions, and a unified audio output. Billing is "Free (limited time)"; the page says the preview version is retired when the free trial ends and a paid version is added. The endpoint is POST /v1/audio/generate.
| Element | StepAudio 3 Gen page | On Sume |
|---|---|---|
| Voices and dialogue | Multi-role lines, timbre, emotion | Separate speech job |
| Sound effects | Generated in the same clip | No sound-effect endpoint listed |
| Background music | Generated in the same clip | Music Router job, 1 to 5000 character prompt |
| Placement and order | Specified in the description | Timeline soundtrack, gain_db, duck_db |
| Output | One unified audio file | One audio artifact per job |
What do I give up by splitting the steps?
One-pass coordination: with separate jobs you decide the levels and timing yourself, and the music does not react to the dialogue. With Sume's Timeline duck_db the bed drops while the spine speaks, a fixed-shape duck, which is a mix control rather than a composition choice. What you gain is that each piece is its own artifact: if one voice line is wrong you redo that job, not the whole clip.
How would I rebuild the clip on Sume?
Generate the voice line, then a music bed with a prompt that says what the bed should do (the Music 1.0 docs suggest naming tempo, instruments and one arc moment, and closing with "Instrumental, no vocals"). Put the voice on the Timeline audio spine and the music in soundtrack, with duck_db above 0 so speech stays clear. Jobs are polled the same way (Jobs and results). Effects have no Sume endpoint; see Sume's sound effects page.
Should I build on a preview model id?
The page says the preview id is retired at the end of the trial, so treat stepaudio-3-gen-preview as temporary and keep the model id in config, not in code.
Sources
Related posts
More in Comparisons
- StepAudio 3 Music from a dry vocal or reference audio vs Sume
StepAudio 3 Music accepts lyrics, vocals or reference audio, even scoring a dry vocal. Sume Music takes a text prompt and one optional image, no audio.
- Synthesia burned-in captions on dubbed videos, and Sume captions
Synthesia added one-click burned-in captions to dubbed videos on 9/30/2026. Here is what that means, and how to burn captions onto a finished video URL on Sume.
- Which Synthesia plan includes API access? Pro is limited
Synthesia lists no API on Basic or Starter, a limited API with 360 minutes a year on Pro, and full access on Enterprise. What a Sume API key gives you instead.
- Synthesia video quizzes are Enterprise-only; what Sume offers instead
Synthesia adds scored quiz questions and a pass threshold to videos on Enterprise only. Sume has no quiz component, only avatar clips you assemble yourself.
Written by Sume