Suno Speech beta: voice and music in one pass, or separate tracks?

Suno's Speech beta makes voice and music in one track. Its blog lists wandering accents and long pauses. When to prefer separate TTS, music and a timeline mix.

4 min readSume
All posts

Use one-pass voice and music when you want a quick, complete clip and can accept whatever balance comes out. Use separate speech, music and a mix when you need to fix a single word, re-time a line or change the music without touching the voice. Suno's October 1 post introduces Speech beta, which combines voice and music in one track from text, and lists known quirks.

What Suno says about Speech beta

The post describes a beta that produces spoken content with music in a single generation. It names two known quirks: accents can wander during a read, and long pauses can appear. Those are vendor-stated limits of a beta, which is the right level of caution to carry into a plan.

Suno Speech beta, as stated by Suno (read 2026-10-03)
ItemWhat the blog says
OutputVoice and music together in one track
InputText
StageBeta
Known quirkWandering accents
Known quirkLong pauses

Where one pass hurts

When voice and music are one file, a fix to either means regenerating both. A wandering accent in line three costs you the music you liked. A long pause cannot be trimmed without cutting the music underneath. You also cannot lower the music under the voice later, because they are no longer separate.

The separate route on Sume

Split the work into three jobs. Text-to-speech makes each line. The Music Router makes the bed through POST /v1/music-router/generate, where sume/music-auto picks the engine (Lyria 3.5 today) and job.request.routed_model names the one that ran. Then Timeline 1.0 assembles them: its soundtrack field takes the bed and a duck_db value lowers it under dialogue.

The Music Router rejects duration and duration_seconds, so the length is steered in the prompt, for example a 30-second track. Music is charged a fixed price per generation. Unbilled /plan calls on Timeline let you check the assembly before you pay for it.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: bed-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Calm acoustic bed, a 30-second track, instrumental, no vocals."
  }'

A quick decision guide

  • Social teaser, one take, low stakes: one pass is faster.
  • Brand voice, legal read, numbers or names: separate tracks, so each line can be checked and redone alone.
  • Same narration over several music options: separate tracks, one voice job and several beds.
  • Localised versions: separate tracks, since only the voice changes.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume