Two voices, one conversation: Sume TTS jobs joined by concat
Make a two-person dialogue file with Sume: one TTS job per line using two avatar voices, then one Timeline audio concat. Cost, code and the gap.

To make a two-voice dialogue with Sume, render each line as its own TTS job with the voice of the speaker, then join the files in order with one POST /v1/timeline-1.0/audio concat. There is no multi-speaker field on the TTS request: one job is one voice. The concat takes up to 20 parts, so a dialogue of up to 20 lines is one join, costs a flat $0.01 (plus the TTS jobs) and returns one file with segments[] that tell you where each line starts. The script below runs it for four lines.
Interest in voice that behaves like a conversation is rising. The October tracker lists Tavus Griffin-Lite, an invite-only research preview of a full-duplex video-to-video conversation model, and MAI-Voice-2.1 Flash at a claimed 150 ms (read 2026-10-06). Those are live. What follows is the recorded kind: a scripted exchange, rendered once, played back as a file.
One job per line, then one concat
import json, os, time, urllib.request as u
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}
API = "https://api.sume.com/v1/"
def call(url, body=None, key=None):
h = dict(H, **({"Idempotency-Key": key} if key else {}))
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
def finish(job):
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(job["status_url"])
return call(job["result_url"])["result"]
LINES = [("HOST", "Welcome back. Today: why the audio matters."),
("GUEST", "Because people forgive grain, not bad sound."),
("HOST", "Say more."), ("GUEST", "Fix the voice first, then polish the picture.")]
HANDLES = {"HOST": os.environ["HOST_HANDLE"], "GUEST": os.environ["GUEST_HANDLE"]}
jobs = [call(API + "tts-router/generate", {"model": "sonic-3.6", "language": "en",
"avatar_handle": HANDLES[who], "transcript": text,
"output_format": {"container": "wav"}}, f"dialogue-v1-{i}")
for i, (who, text) in enumerate(LINES)]
parts = [{"url": next(a["url"] for a in finish(j)["artifacts"] if a["type"] == "audio")}
for j in jobs]
joined = finish(call(API + "timeline-1.0/audio",
{"operation": "concat", "parts": parts}, "dialogue-join-v1"))
print(joined["audio_url"])
for seg in joined["segments"]:
print(LINES[seg["index"]][0], f"starts at {seg['start']:.2f}s")How the script works
All submits go out first and the jobs run in parallel, so four lines take roughly as long as the slowest one. Each job has its own Idempotency-Key, so a rerun reuses finished takes. The join requires every part on media.sume.com, which TTS artifacts are, and every part with the same channel layout (audio_parts_channel_mismatch otherwise). Two voices from the same engine and the same output format will match. The result is also usable as audio.url on a Timeline 1.0 render, and the produced audio is capped at 1800 seconds. The default output is wav, which is sample-exact; mp3 is smaller but adds priming padding at each edge, so keep wav if you will join the file again. The concat result carries segments[] with index, start and duration_seconds, one per part, which is what the last loop prints.
What it cannot do
Be honest about the seam: the concat is gapless and has no pause parameter. A line that ends with little trailing silence runs straight into the next one, which can feel rushed in a back-and-forth. Listen to a short test before you build a long script. If the pacing is wrong, adjust each line, with speed from 0.6 to 1.5 and punctuation that ends sentences cleanly, and measure again. Tone also differs between takes, because each line is a separate request with no memory of the one before it. For a scripted explainer that is fine. For a scene that needs overlap, interruption and breath, a live conversation model is the right class of tool.
| Step | Price | Note |
|---|---|---|
| Four TTS jobs | $0.04 | About $0.002 each at $0.0475 per 1,000 characters, rounded up to the cent per job |
| Timeline audio concat | $0.01 | Flat per job |
| Total | $0.05 | Per dialogue file |
Picking voices and next steps
Voice picks come from the avatars in your workspace: list them with GET /v1/avatar-1.0/avatars and use handles whose voice is ready, as shown in pick a voice with only an API key. A pair of voices that differ in pitch and pace makes the speakers easy to tell apart without labels. For the picture side, two-speaker avatar video covers the video path. Prices are on the pricing page, the join is described in Timeline audio, and the job envelope in Jobs and results.
Sources
Related posts
More in Use cases
- UGC-style ad: test five hooks on one body with Timeline plans
Join five 3-second hooks to one 12-second body clip with Timeline 1.0. Plan each cut unbilled, then render the winners at $0.10 a minute.
- UGC-style ad: a music bed that ducks under the voice
Add a looped music bed to a UGC-style ad and lower it under the voice with soundtrack.duck_db in a Timeline 1.0 render. Needs a real voice spine.
- Video ad end card: add a fade-in and fade-out with Timeline
Close an ad with a 1-second fade-out and open it with a 0.5-second fade-in using output.fade_in_seconds and fade_out_seconds in a Timeline 1.0 render.
- Which Sume video model makes a 15-second 9:16 ad? A catalog filter
List the models in GET /v1/videos/models that accept 15 seconds at 9:16 with a short Python filter, instead of trusting a stale table.
Written by Sume