Two-voice dialogue: Gemini TTS vs ElevenLabs Text to Dialogue

Gemini TTS handles up to 2 speakers per request; ElevenLabs Text to Dialogue keeps continuity with surrounding text and request IDs. In Sume you stitch jobs.

4 min readSume
All posts

For a two-person conversation, Gemini TTS can render both voices in one request (up to 2 speakers with prebuilt voices), while ElevenLabs' Text to Dialogue handles continuity across requests by letting you pass the text before and after a line, or earlier request IDs. In Sume, every TTS job has one voice, so a dialogue is a set of jobs you order and join yourself.

What the vendors document

Vendor pages, read 2026-10-05
ItemGemini TTSElevenLabs
Multi-speakerUp to 2 speakers per request, prebuilt voicesMulti-speaker dialogue on Eleven v4 (announcement)
ContinuityInside one requestText to Dialogue: preceding/following text and request IDs (Sept 28 changelog)
Delivery controlstyle field and inline angle-bracket tagsInline audio tags such as [laughs]
Per-request limitNot stated on the page I readUp to 10,000 characters per generation (v4 page)

Which shape fits

  • Two hosts, one script, short episode: Gemini's single request keeps the pacing between speakers in one render.
  • More than two voices, or a long script split across requests: ElevenLabs' continuity inputs exist for that.
  • Per-line control (re-record one line without touching the rest): separate renders per line are easier to replace. This is how Sume works.

Building a dialogue in Sume

Each POST /v1/tts-router/generate job takes one voice selector and one transcript of 1 to 20,000 characters. For a dialogue, make one job per turn (or per run of turns by the same speaker), then join the audio on a timeline in order. A few practical rules:

Use transcript_source with script_revision_id and ordered sentence_ids so each job is tied to exact lines of an accepted script, with a receipt per job. Add timestamps when you need word timing for captions. Leave a short gap between turns when you assemble, and listen to the joins, because each job is rendered without knowledge of its neighbours.

Cost is per character: Sonic list x1.25 is $0.0475 per 1K characters, rounded up to the cent per job, so many very short turns each bill at least one cent. Batch consecutive lines by the same speaker into one job where you can.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume