Suno Speech makes voice and music as one track. Sume keeps them apart

Suno Speech beta renders voice and music together. On Sume you generate speech, a music bed and a ducked mix as separate jobs, so each can be redone alone.

4 min readSume
All posts

Suno Speech is a beta model that generates a spoken voice and its background music as one finished track. Sume does not list a model that does this. Instead, you make the voice with TTS, the bed with the Music Router, and mix them in Timeline 1.0, so you can redo any one part without touching the others.

Suno announced Speech on October 1, 2026 and says it is opening the beta to everyone (Suno blog, read 2026-10-10). The decision is not which is better. It is whether you need to edit the voice and the music separately after the first render.

What Suno says Speech does

Suno describes Speech as the first audio model that generates voice and music together as one cohesive track. You type an idea, a poem or something you wrote, then describe the voice and the musical style. The output is spoken audio set to original background music.

The post does not state length limits, pricing, credits or commercial rights, so this article does not either. It also admits the beta is rough: Suno writes that British accents can wander off to Australia and back, and that dramatic pauses may be very dramatic.

Suno Speech beta versus the Sume path. Suno facts from its Oct 1 2026 post (read 2026-10-10); Sume facts from docs.sume.com (read 2026-10-10).
QuestionSuno Speech betaSume
OutputOne track: voice plus original background musicSeparate files: a TTS voice file, a music file, and a rendered mix
How you describe itText plus a description of voice and musical styleTranscript to TTS; a separate music prompt (up to 5000 characters) to the Music Router
Change only the musicNot stated in the postRegenerate the bed; the voice file stays as it is
Change only a spoken lineNot stated in the postRegenerate that line and join with timeline audio concat
Music level under voiceNot stated in the postTimeline soundtrack field duck_db, 0 to 20, plus gain_db
Price and lengthNot stated in the postMusic: fixed $0.125 per generation; Timeline: $0.10 per output minute

The Sume path, step by step

Each step is its own job with its own result, so a bad take costs you one step and not the whole track.

Generate the bed first. The Music Router takes a text prompt and picks the engine with sume/music-auto. Sume rejects duration fields, so write the length into the prompt, and put exclusions in the positive prompt such as Instrumental, no vocals, because a non-empty negative_prompt is refused.

  • Voice: POST /v1/tts-1.0/generate with a transcript and an avatar_id or voice.id. Set language for any non-English text. Ask for wav when the file feeds a join.
  • Music: POST /v1/music-router/generate with a prompt. Read the audio artifact from result.artifacts[] where type is audio.
  • Import both files so they live on media.sume.com. Timeline accepts only this workspace's hosted media.
  • Mix: POST /v1/timeline-1.0/render with the voice as the audio spine and the music as soundtrack with duck_db.

The mix request

The audio spine is the voice. The soundtrack field is the bed, and duck_db lowers it while the spine speaks. Duck needs a real spine, not silence, so a ducked bed always sits under actual speech. This request follows the Timeline 1.0 docs; replace the demo URLs with your own imported artifacts.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: speech-mix-001" \
  -d '{
    "audio": {
      "url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
      "duration_seconds": 24
    },
    "soundtrack": {
      "url": "https://media.sume.com/artifacts/artf_demo/bed.mp3",
      "gain_db": -6,
      "duck_db": 12,
      "fade_out_seconds": 3
    },
    "video": [
      {
        "source_url": "https://media.sume.com/artifacts/artf_demo/scene.mp4",
        "start": 0,
        "duration": 24
      }
    ]
  }'

When one track is enough, and when to split

A single generated track suits a poem, a greeting or a quick spoken piece where the words and the music are one idea and you are happy to re-roll the whole thing. Sume has no equivalent single-prompt model, so that use case is outside what Sume ships today.

Split stems suit anything with a script that must stay exact, a brand voice that must stay the same across versions, or a bed you want to swap per channel. Long scripts can be sliced into takes and rejoined gaplessly with timeline audio concat (1 to 20 parts, $0.01 per job), and the returned segments[] offsets tell you where each line starts.

  • Need the same voice with a different bed: regenerate music only.
  • Need one corrected sentence: regenerate that take and concat again.
  • Need the mix on a picture: the render already takes the clip and the audio together.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume