How do I make a two-voice dialogue audio for language learners?

Make a two-speaker practice dialogue from one TTS job per line, joined gapless by a $0.01 concat. A ten-line dialogue is about 11 cents on Sume.

4 min readSume
All posts

To make a two-voice dialogue for language learners, send each speaker's lines as separate TTS jobs with two different voices, then join the files in order with a Timeline audio concat. Ten lines of about 120 characters each is 1,200 characters of speech. Sume bills each job to whole cents, so ten jobs cost $0.10 and the concat adds $0.01, about 11 cents in all.

Sume TTS takes one voice per request, so a dialogue is not one call with speaker tags. That is a limit, but for learners it is also a feature: every line is its own file, so you can offer a slow repeat, a line-by-line mode and a version with one speaker muted.

Script the dialogue for the ear

Learners need short turns. Write lines of 60 to 140 characters, one idea each, and put the new vocabulary in the first half of the line, where it is easiest to catch. Use contractions and natural questions, not textbook sentences, and keep the same two names throughout so the listener can follow who is speaking.

Mark each line with a speaker and an index in your own script file, such as A01 or B02. Those labels become the idempotency keys, so a rewrite of a single line is one new job and the rest are untouched.

Choose two voices and set the language

Pick two clearly different voices from the library, and send language with the language you are teaching on every job. The voice and the language should agree; if they do not, Sume returns a 409 tts_voice_language_mismatch before any charge. For Korean and Japanese learners, Sume can infer the language only when the text is Hangul-only or kana-only; send the code explicitly for everything else.

Slow the learner version with generation_config.speed at 0.8 and keep the normal version at 1.0. The price is the same for both, since it follows characters, so a slow set costs another full set of jobs, not a surcharge.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: dialogue-cafe-A01-slow" \
  -d '{
    "transcript": "Bonjour, je voudrais un caf\u00e9, s'"'"'il vous pla\u00eet.",
    "voice": { "id": "'"$VOICE_A"'" },
    "language": "fr",
    "generation_config": { "speed": 0.8 },
    "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 }
  }'

Join the lines

Timeline audio concat takes 1 to 20 parts, so a dialogue of up to 20 lines is one request, joined sample for sample with no re-synthesis, for a flat $0.01. The result returns an offset for each part, which you can use to build a transcript with timings or to cut the dialogue back into lines with a split. For dialogues longer than 20 lines, join in two stages.

The join leaves no gap between speakers, and natural speech has one. Add a short pause as its own part, a few hundred milliseconds of silence imported as an audio file, between the lines. Keep the pause file the same channel layout as the voices; mismatched layouts fail with audio_parts_channel_mismatch.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: dialogue-cafe-joined-slow" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "'"$A01"'" }, { "url": "'"$PAUSE"'" }, { "url": "'"$B02"'" }
    ]
  }'

Versions learners ask for

Because every line is a file, the versions come almost free. A repeat-after-me set plays each line, then a pause long enough for the learner to say it. A role-play set drops one speaker's lines and leaves the pause, so the learner supplies the missing part. A listening quiz shuffles the lines and asks what was said. Each is a different list of parts into the same concat call, not new speech, so only the first generation is billed for the voices.

Cost of a lesson set

A lesson usually ships several versions: normal speed, slow speed, and perhaps a one-sided practice version where the learner speaks one role. Each version is another set of jobs and one concat. The table prices a ten-line dialogue at 1,200 characters.

  • Learners replay lines, so pronunciation quality matters more than a cent or two of price.
  • Have a fluent speaker listen to every line before it ships.
  • Keep the script and audio together so a correction is a one-line change.
Cost of a ten-line, 1,200-character dialogue, vendor list rates read 2026-10-07 and Sume TTS 1.0 catalog rate.
Provider and modelListed rateCost of the speech
Sume TTS 1.0, ten jobs plus concat$0.0475 per 1,000 characters, whole cents per job, $0.01 concat$0.11
ElevenLabs Flash/Turbo$0.04 per 1,000 characters$0.048
ElevenLabs v3$0.08 per 1,000 characters$0.096

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume