Two-voice dialogue audio on Sume: TTS lines joined with audio concat
Build a role-play or interview track by generating one TTS line per turn in each speaker's voice, then joining up to 20 turns into one gapless file for $0.01.

The recipe
To make a two-voice dialogue on Sume, generate each turn as its own text-to-speech job with the right voice, then pass the audio files in order to one timeline-audio concat job. You get a single gapless file and a segments[] list that tells you where each turn starts. Microsoft's launch material lists distinct speakers for tutoring, role-play and simulation as a use case; this is the Sume way to build that kind of track from separate voices.
Source note: the use case comes from the Unite.AI article on the MAI launch (read 2026-10-07). The steps below use only Sume's documented jobs.
Step by step
Each turn is a plain TTS request with the speaker's voice. Ask for the same output format on every turn, WAV for anything you will join again, because the concat job needs the same channel layout across parts and fails with audio_parts_channel_mismatch if they differ. Then make the concat request from the finished files. Parts must already be media.sume.com audio from your workspace.
The script below builds the concat request body from a list of turns. It uses placeholder URLs and prints the JSON, so it runs as-is. Replace the URLs with the audio_url values from your finished TTS jobs.
import json
turns = [
("customer", "https://media.sume.com/artifacts/artf_demo/turn01.wav"),
("agent", "https://media.sume.com/artifacts/artf_demo/turn02.wav"),
("customer", "https://media.sume.com/artifacts/artf_demo/turn03.wav"),
]
assert 1 <= len(turns) <= 20, "concat takes 1-20 parts"
body = {
"operation": "concat",
"parts": [{"url": url} for _, url in turns],
}
print(json.dumps(body, indent=2))
# POST this to https://api.sume.com/v1/timeline-1.0/audio
# with Authorization: Bearer $SUME_API_KEY and an Idempotency-Key header.What it costs
Short lines are billed in whole cents per job, so cost scales with the number of turns more than with their length. A turn of 140 characters is 140 characters at $47.50 per million, 0.665 cents, which rounds up to 1 cent. Twelve turns of that size cost 12 cents, plus $0.01 for the join, 13 cents in total. The concat job is flat at $0.01 and uses no provider inference.
Using the segment offsets
The concat result returns one audio_url, the total duration_seconds and segments[] with index, start and duration_seconds for each part. Use start to cue subtitles per speaker, or to re-base video[].start if you lay a talking-head or still image per turn on a timeline.
Limits to plan for
- The limit is 20 parts per join. Longer dialogues need two joins, then a join of the joins; see the post on the 1800-second and 20-part limits.
- The output is capped at 1800 seconds.
- The docs describe the join as gapless, with no silence at the seams, so any pause between speakers has to be in the audio you join.
- Keep each speaker on the same voice id across turns, and check the language of the voice against the
languagefield before a long run.
Pacing between turns
A concat job joins parts with no silence at the seams, so a dialogue made this way has no gap between turns unless the audio itself carries one. If a pause is needed between speakers, shorten or lengthen it on the player side or add a short silent file as one of the parts. A silent file counts as a part, and a job takes up to 20 parts.
Check the language and voice pairing first. Sume's TTS returns a 409 tts_voice_language_mismatch when a voice does not match the language of a line, and you can send confirm_language_mismatch if you mean it, so a line in another language needs that decision on purpose.
For a long dialogue with more than 20 turns, join in groups of 20 and then join the groups, as the output of a concat job can be a part of a later one.
- Silence is a part, not a setting.
- Match voice language to line language.
- More than 20 turns: join in groups, then join the groups.
Sources
Related posts
More in Media tools
- One lip-sync model per video: Fabric runs at 25 fps, measure H3 Max
Mixing Fabric and MiniMax H3 Max lip-sync clips in one Sume video risks a frame-rate mismatch. Pick one per run and check fps with ffprobe on the first clip.
- video-filter dim amount 0 or 1.2 is refused: the (0, 1] range
Dim amount takes values above 0 up to 1. Zero, negatives and 1.2 return video_filter_amount_out_of_range. What it does, and how to lift a dark clip.
- video-filter invalid_filtergraph: every reason and the fix for each
A video-filter filtergraph is refused for eight reasons, from a [0:v] label to a quote character. What each one means and how to rewrite the graph so it passes.
- video-filter unsupported_filter_op_field: crop and dim keys allowed
A crop op takes four fractions and dim takes one amount; any extra key returns unsupported_filter_op_field. The exact key lists and where the extra work goes.
Written by Sume