Two-voice dialogue audio on Sume: TTS lines joined with audio concat

Build a role-play or interview track by generating one TTS line per turn in each speaker's voice, then joining up to 20 turns into one gapless file for $0.01.

5 min readSume
All posts

The recipe

To make a two-voice dialogue on Sume, generate each turn as its own text-to-speech job with the right voice, then pass the audio files in order to one timeline-audio concat job. You get a single gapless file and a segments[] list that tells you where each turn starts. Microsoft's launch material lists distinct speakers for tutoring, role-play and simulation as a use case; this is the Sume way to build that kind of track from separate voices.

Source note: the use case comes from the Unite.AI article on the MAI launch (read 2026-10-07). The steps below use only Sume's documented jobs.

Step by step

Each turn is a plain TTS request with the speaker's voice. Ask for the same output format on every turn, WAV for anything you will join again, because the concat job needs the same channel layout across parts and fails with audio_parts_channel_mismatch if they differ. Then make the concat request from the finished files. Parts must already be media.sume.com audio from your workspace.

The script below builds the concat request body from a list of turns. It uses placeholder URLs and prints the JSON, so it runs as-is. Replace the URLs with the audio_url values from your finished TTS jobs.

import json

turns = [
    ("customer", "https://media.sume.com/artifacts/artf_demo/turn01.wav"),
    ("agent", "https://media.sume.com/artifacts/artf_demo/turn02.wav"),
    ("customer", "https://media.sume.com/artifacts/artf_demo/turn03.wav"),
]
assert 1 <= len(turns) <= 20, "concat takes 1-20 parts"

body = {
    "operation": "concat",
    "parts": [{"url": url} for _, url in turns],
}
print(json.dumps(body, indent=2))
# POST this to https://api.sume.com/v1/timeline-1.0/audio
# with Authorization: Bearer $SUME_API_KEY and an Idempotency-Key header.

What it costs

Short lines are billed in whole cents per job, so cost scales with the number of turns more than with their length. A turn of 140 characters is 140 characters at $47.50 per million, 0.665 cents, which rounds up to 1 cent. Twelve turns of that size cost 12 cents, plus $0.01 for the join, 13 cents in total. The concat job is flat at $0.01 and uses no provider inference.

Using the segment offsets

The concat result returns one audio_url, the total duration_seconds and segments[] with index, start and duration_seconds for each part. Use start to cue subtitles per speaker, or to re-base video[].start if you lay a talking-head or still image per turn on a timeline.

Limits to plan for

  • The limit is 20 parts per join. Longer dialogues need two joins, then a join of the joins; see the post on the 1800-second and 20-part limits.
  • The output is capped at 1800 seconds.
  • The docs describe the join as gapless, with no silence at the seams, so any pause between speakers has to be in the audio you join.
  • Keep each speaker on the same voice id across turns, and check the language of the voice against the language field before a long run.

Pacing between turns

A concat job joins parts with no silence at the seams, so a dialogue made this way has no gap between turns unless the audio itself carries one. If a pause is needed between speakers, shorten or lengthen it on the player side or add a short silent file as one of the parts. A silent file counts as a part, and a job takes up to 20 parts.

Check the language and voice pairing first. Sume's TTS returns a 409 tts_voice_language_mismatch when a voice does not match the language of a line, and you can send confirm_language_mismatch if you mean it, so a line in another language needs that decision on purpose.

For a long dialogue with more than 20 turns, join in groups of 20 and then join the groups, as the output of a concat job can be a part of a later one.

  • Silence is a part, not a setting.
  • Match voice language to line language.
  • More than 20 turns: join in groups, then join the groups.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume