StepAudio 3 Gen: voice, SFX and music in one clip vs Sume jobs

StepFun's stepaudio-3-gen-preview makes voice, effects, ambience and music in one audio output. Sume uses separate music, speech and Timeline mix steps.

5 min readSume
All posts

StepFun's StepAudio 3 Gen turns one natural-language description into a single audio clip that can hold voices, sound effects, ambient sound, background music and singing, and the description can say where and in what order each element appears. Sume does not have a one-call equivalent. On Sume you make the pieces as separate jobs (music from the Music Router, speech from the speech surface) and combine them in a Timeline render, which has a soundtrack bed with gain, fade and duck. StepFun's side is from its StepAudio 3 Gen page.

What does StepAudio 3 Gen say it does?

The page lists model id stepaudio-3-gen-preview, inputs of text, audio, and role and script descriptions, and a unified audio output. Billing is "Free (limited time)"; the page says the preview version is retired when the free trial ends and a paid version is added. The endpoint is POST /v1/audio/generate.

StepAudio 3 Gen page facts and the Sume equivalent, read 2026-10-02.
ElementStepAudio 3 Gen pageOn Sume
Voices and dialogueMulti-role lines, timbre, emotionSeparate speech job
Sound effectsGenerated in the same clipNo sound-effect endpoint listed
Background musicGenerated in the same clipMusic Router job, 1 to 5000 character prompt
Placement and orderSpecified in the descriptionTimeline soundtrack, gain_db, duck_db
OutputOne unified audio fileOne audio artifact per job

What do I give up by splitting the steps?

One-pass coordination: with separate jobs you decide the levels and timing yourself, and the music does not react to the dialogue. With Sume's Timeline duck_db the bed drops while the spine speaks, a fixed-shape duck, which is a mix control rather than a composition choice. What you gain is that each piece is its own artifact: if one voice line is wrong you redo that job, not the whole clip.

How would I rebuild the clip on Sume?

Generate the voice line, then a music bed with a prompt that says what the bed should do (the Music 1.0 docs suggest naming tempo, instruments and one arc moment, and closing with "Instrumental, no vocals"). Put the voice on the Timeline audio spine and the music in soundtrack, with duck_db above 0 so speech stays clear. Jobs are polled the same way (Jobs and results). Effects have no Sume endpoint; see Sume's sound effects page.

Should I build on a preview model id?

The page says the preview id is retired at the end of the trial, so treat stepaudio-3-gen-preview as temporary and keep the model id in config, not in code.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume