AI dubbing for TikTok: one 45-second video in 3 languages, $0.44

A dubbing pipeline on Sume is detach, STT, your translation, TTS per language, then a render: about $0.44 for three 45-second versions.

6 min readSume
All posts

"AI dubbing TikTok" and "AI dubbing video translator" are on Google's autocomplete today (read 2026-10-07). Sume's docs do not describe a single dubbing endpoint. They describe the parts: an audio extractor, speech to text, text to speech with a voice per language, and a timeline render. Put together they dub a video, and the cost is small enough to total on one line.

The pipeline below makes three versions of a 45-second TikTok. It changes the voice track only; the mouth in the picture is the original, so it suits voice-over, screen recordings, product shots and cutaways more than close-up talking heads.

The five steps

  • Detach the audio once: POST /v1/audio-detach, WAV, $0.01.
  • Transcribe it: POST /v1/stt-1.0/transcribe with language_code as a hint, $0.01 for the minute.
  • Translate the text yourself or with your agent. This step is not a Sume media job and carries no line in the media rate card.
  • Speak each translation with POST /v1/tts-1.0/generate, one voice per language.
  • Render each language: Timeline 1.0 with the original video and the new file as audio.url, $0.10 per output minute.

Cost for three languages

Assume about 700 characters of script in each language for 45 seconds of speech, which is a planning figure, not a measured one. At $0.0475 per 1,000 characters, 700 characters is $0.033, rounded up to $0.04 per job.

Three-language dub of one 45-second video, Sume catalog read 2026-10-07
StepCountEachTotal
Audio detach1$0.01$0.01
STT, 45 s billed as 1 minute1$0.01$0.01
TTS, about 700 characters3$0.04$0.12
Timeline render, under 1 minute3$0.10$0.30
Total$0.44

Where it goes wrong

  • Send the right language to TTS. If the voice and the language disagree Sume returns tts_voice_language_warning and wants a confirmation before the retry, which is the guard working as designed.
  • Dubbed speech is a different length from the original. Measure the TTS duration_seconds, and cut or extend the picture to match before the render.
  • Korean, Japanese and other long scripts count by character, spaces and punctuation included, so run a dry_run or max_spend_usd cap on batch jobs.
  • Listen to a name or a number in each language before you ship; a receipt proves the input reached the job, not that it was pronounced right.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume