How do I make a two-voice dialogue audio for language learners?
Make a two-speaker practice dialogue from one TTS job per line, joined gapless by a $0.01 concat. A ten-line dialogue is about 11 cents on Sume.

To make a two-voice dialogue for language learners, send each speaker's lines as separate TTS jobs with two different voices, then join the files in order with a Timeline audio concat. Ten lines of about 120 characters each is 1,200 characters of speech. Sume bills each job to whole cents, so ten jobs cost $0.10 and the concat adds $0.01, about 11 cents in all.
Sume TTS takes one voice per request, so a dialogue is not one call with speaker tags. That is a limit, but for learners it is also a feature: every line is its own file, so you can offer a slow repeat, a line-by-line mode and a version with one speaker muted.
Script the dialogue for the ear
Learners need short turns. Write lines of 60 to 140 characters, one idea each, and put the new vocabulary in the first half of the line, where it is easiest to catch. Use contractions and natural questions, not textbook sentences, and keep the same two names throughout so the listener can follow who is speaking.
Mark each line with a speaker and an index in your own script file, such as A01 or B02. Those labels become the idempotency keys, so a rewrite of a single line is one new job and the rest are untouched.
Choose two voices and set the language
Pick two clearly different voices from the library, and send language with the language you are teaching on every job. The voice and the language should agree; if they do not, Sume returns a 409 tts_voice_language_mismatch before any charge. For Korean and Japanese learners, Sume can infer the language only when the text is Hangul-only or kana-only; send the code explicitly for everything else.
Slow the learner version with generation_config.speed at 0.8 and keep the normal version at 1.0. The price is the same for both, since it follows characters, so a slow set costs another full set of jobs, not a surcharge.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dialogue-cafe-A01-slow" \
-d '{
"transcript": "Bonjour, je voudrais un caf\u00e9, s'"'"'il vous pla\u00eet.",
"voice": { "id": "'"$VOICE_A"'" },
"language": "fr",
"generation_config": { "speed": 0.8 },
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 }
}'Join the lines
Timeline audio concat takes 1 to 20 parts, so a dialogue of up to 20 lines is one request, joined sample for sample with no re-synthesis, for a flat $0.01. The result returns an offset for each part, which you can use to build a transcript with timings or to cut the dialogue back into lines with a split. For dialogues longer than 20 lines, join in two stages.
The join leaves no gap between speakers, and natural speech has one. Add a short pause as its own part, a few hundred milliseconds of silence imported as an audio file, between the lines. Keep the pause file the same channel layout as the voices; mismatched layouts fail with audio_parts_channel_mismatch.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dialogue-cafe-joined-slow" \
-d '{
"operation": "concat",
"parts": [
{ "url": "'"$A01"'" }, { "url": "'"$PAUSE"'" }, { "url": "'"$B02"'" }
]
}'Versions learners ask for
Because every line is a file, the versions come almost free. A repeat-after-me set plays each line, then a pause long enough for the learner to say it. A role-play set drops one speaker's lines and leaves the pause, so the learner supplies the missing part. A listening quiz shuffles the lines and asks what was said. Each is a different list of parts into the same concat call, not new speech, so only the first generation is billed for the voices.
Cost of a lesson set
A lesson usually ships several versions: normal speed, slow speed, and perhaps a one-sided practice version where the learner speaks one role. Each version is another set of jobs and one concat. The table prices a ten-line dialogue at 1,200 characters.
- Learners replay lines, so pronunciation quality matters more than a cent or two of price.
- Have a fluent speaker listen to every line before it ships.
- Keep the script and audio together so a correction is a one-line change.
| Provider and model | Listed rate | Cost of the speech |
|---|---|---|
| Sume TTS 1.0, ten jobs plus concat | $0.0475 per 1,000 characters, whole cents per job, $0.01 concat | $0.11 |
| ElevenLabs Flash/Turbo | $0.04 per 1,000 characters | $0.048 |
| ElevenLabs v3 | $0.08 per 1,000 characters | $0.096 |
Sources
Related posts
More in Use cases
- UGC-style avatar video: a 9:16 casual scene prompt that works on Sume
How to ask Sume Avatar 1.0 for a phone-filmed look: 9:16, a casual background prompt, silence beats, product image. Request body and what it costs per tier.
- Virtual staging an empty room photo with AI: GPT Image 2.5 via API
Stage an empty room photo with openai/gpt-image-2.5: the room plus up to 15 furniture references, an optional mask, and about $0.08 per staged image at high.
- How do I make a winter tire changeover promo video for my garage?
Three 6-second Wan 3.0 clips, a 420-character voiceover, a duck-under music bed and booking captions for a garage: $2.695 on Sume at 720p.
- Write a lip-sync script in 5 to 14.8 second beats (H3 Max voice-over)
MiniMax H3 Max lip-sync on Sume takes audio of 5 to 14.8 seconds. Write the script in beats that fit the window, then voice each beat; the price per beat.
Written by Sume