Two-voice dialogue: Gemini TTS vs ElevenLabs Text to Dialogue
Gemini TTS handles up to 2 speakers per request; ElevenLabs Text to Dialogue keeps continuity with surrounding text and request IDs. In Sume you stitch jobs.

For a two-person conversation, Gemini TTS can render both voices in one request (up to 2 speakers with prebuilt voices), while ElevenLabs' Text to Dialogue handles continuity across requests by letting you pass the text before and after a line, or earlier request IDs. In Sume, every TTS job has one voice, so a dialogue is a set of jobs you order and join yourself.
What the vendors document
| Item | Gemini TTS | ElevenLabs |
|---|---|---|
| Multi-speaker | Up to 2 speakers per request, prebuilt voices | Multi-speaker dialogue on Eleven v4 (announcement) |
| Continuity | Inside one request | Text to Dialogue: preceding/following text and request IDs (Sept 28 changelog) |
| Delivery control | style field and inline angle-bracket tags | Inline audio tags such as [laughs] |
| Per-request limit | Not stated on the page I read | Up to 10,000 characters per generation (v4 page) |
Which shape fits
- Two hosts, one script, short episode: Gemini's single request keeps the pacing between speakers in one render.
- More than two voices, or a long script split across requests: ElevenLabs' continuity inputs exist for that.
- Per-line control (re-record one line without touching the rest): separate renders per line are easier to replace. This is how Sume works.
Building a dialogue in Sume
Each POST /v1/tts-router/generate job takes one voice selector and one transcript of 1 to 20,000 characters. For a dialogue, make one job per turn (or per run of turns by the same speaker), then join the audio on a timeline in order. A few practical rules:
Use transcript_source with script_revision_id and ordered sentence_ids so each job is tied to exact lines of an accepted script, with a receipt per job. Add timestamps when you need word timing for captions. Leave a short gap between turns when you assemble, and listen to the joins, because each job is rendered without knowledge of its neighbours.
Cost is per character: Sonic list x1.25 is $0.0475 per 1K characters, rounded up to the cent per job, so many very short turns each bill at least one cent. Batch consecutive lines by the same speaker into one job where you can.
Sources
Related posts
More in Comparisons
- GitHub Actions schedule or a Sume schedule for a weekly video run?
For a weekly video run, GitHub Actions cron is UTC-first and best-effort at busy times; a Sume schedule has a timezone and a spend cap. You can also chain them.
- GPT-6.1 Sol structured outputs or a Sume Format output_schema?
GPT-6.1 Sol supports structured outputs and function calling in OpenAI's API. A Sume Format run adds media, a spend cap and a receipt. Pick by the job.
- GPT Image 2.5 mask edit, then an Ideogram 4.5 text pass: one chain
Use GPT Image 2.5 for the masked region change and Ideogram 4.5 for the text pass, both on POST /v1/images. When it beats one model, and what each step bills.
- GPT Image 2.5 medium is 6x cheaper than Nano Banana 2 1K on Sume
On Sume, GPT Image 2.5 at medium costs $0.0165 for 1024x1024 and Nano Banana 2 1K costs $0.10, a 6.1x gap. Ratios at 2K and 4K, and what the price omits.
Written by Sume