Two voices, one conversation: Sume TTS jobs joined by concat

Make a two-person dialogue file with Sume: one TTS job per line using two avatar voices, then one Timeline audio concat. Cost, code and the gap.

5 min readSume
All posts

To make a two-voice dialogue with Sume, render each line as its own TTS job with the voice of the speaker, then join the files in order with one POST /v1/timeline-1.0/audio concat. There is no multi-speaker field on the TTS request: one job is one voice. The concat takes up to 20 parts, so a dialogue of up to 20 lines is one join, costs a flat $0.01 (plus the TTS jobs) and returns one file with segments[] that tell you where each line starts. The script below runs it for four lines.

Interest in voice that behaves like a conversation is rising. The October tracker lists Tavus Griffin-Lite, an invite-only research preview of a full-duplex video-to-video conversation model, and MAI-Voice-2.1 Flash at a claimed 150 ms (read 2026-10-06). Those are live. What follows is the recorded kind: a scripted exchange, rendered once, played back as a file.

One job per line, then one concat

import json, os, time, urllib.request as u
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}
API = "https://api.sume.com/v1/"

def call(url, body=None, key=None):
    h = dict(H, **({"Idempotency-Key": key} if key else {}))
    data = json.dumps(body).encode() if body else None
    return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]

def finish(job):
    while not job["terminal"]:
        time.sleep(job.get("next_poll_after_seconds") or 2)
        job = call(job["status_url"])
    return call(job["result_url"])["result"]

LINES = [("HOST", "Welcome back. Today: why the audio matters."),
         ("GUEST", "Because people forgive grain, not bad sound."),
         ("HOST", "Say more."), ("GUEST", "Fix the voice first, then polish the picture.")]
HANDLES = {"HOST": os.environ["HOST_HANDLE"], "GUEST": os.environ["GUEST_HANDLE"]}
jobs = [call(API + "tts-router/generate", {"model": "sonic-3.6", "language": "en",
         "avatar_handle": HANDLES[who], "transcript": text,
         "output_format": {"container": "wav"}}, f"dialogue-v1-{i}")
        for i, (who, text) in enumerate(LINES)]
parts = [{"url": next(a["url"] for a in finish(j)["artifacts"] if a["type"] == "audio")}
         for j in jobs]
joined = finish(call(API + "timeline-1.0/audio",
                     {"operation": "concat", "parts": parts}, "dialogue-join-v1"))
print(joined["audio_url"])
for seg in joined["segments"]:
    print(LINES[seg["index"]][0], f"starts at {seg['start']:.2f}s")

How the script works

All submits go out first and the jobs run in parallel, so four lines take roughly as long as the slowest one. Each job has its own Idempotency-Key, so a rerun reuses finished takes. The join requires every part on media.sume.com, which TTS artifacts are, and every part with the same channel layout (audio_parts_channel_mismatch otherwise). Two voices from the same engine and the same output format will match. The result is also usable as audio.url on a Timeline 1.0 render, and the produced audio is capped at 1800 seconds. The default output is wav, which is sample-exact; mp3 is smaller but adds priming padding at each edge, so keep wav if you will join the file again. The concat result carries segments[] with index, start and duration_seconds, one per part, which is what the last loop prints.

What it cannot do

Be honest about the seam: the concat is gapless and has no pause parameter. A line that ends with little trailing silence runs straight into the next one, which can feel rushed in a back-and-forth. Listen to a short test before you build a long script. If the pacing is wrong, adjust each line, with speed from 0.6 to 1.5 and punctuation that ends sentences cleanly, and measure again. Tone also differs between takes, because each line is a separate request with no memory of the one before it. For a scripted explainer that is fine. For a scene that needs overlap, interruption and breath, a live conversation model is the right class of tool.

Cost of a four-line dialogue, about 170 characters (read 2026-10-06)
StepPriceNote
Four TTS jobs$0.04About $0.002 each at $0.0475 per 1,000 characters, rounded up to the cent per job
Timeline audio concat$0.01Flat per job
Total$0.05Per dialogue file

Picking voices and next steps

Voice picks come from the avatars in your workspace: list them with GET /v1/avatar-1.0/avatars and use handles whose voice is ready, as shown in pick a voice with only an API key. A pair of voices that differ in pitch and pace makes the speakers easy to tell apart without labels. For the picture side, two-speaker avatar video covers the video path. Prices are on the pricing page, the join is described in Timeline audio, and the job envelope in Jobs and results.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume