AI dubbing for TikTok: one 45-second video in 3 languages, $0.44
A dubbing pipeline on Sume is detach, STT, your translation, TTS per language, then a render: about $0.44 for three 45-second versions.

"AI dubbing TikTok" and "AI dubbing video translator" are on Google's autocomplete today (read 2026-10-07). Sume's docs do not describe a single dubbing endpoint. They describe the parts: an audio extractor, speech to text, text to speech with a voice per language, and a timeline render. Put together they dub a video, and the cost is small enough to total on one line.
The pipeline below makes three versions of a 45-second TikTok. It changes the voice track only; the mouth in the picture is the original, so it suits voice-over, screen recordings, product shots and cutaways more than close-up talking heads.
The five steps
- Detach the audio once:
POST /v1/audio-detach, WAV, $0.01. - Transcribe it:
POST /v1/stt-1.0/transcribewithlanguage_codeas a hint, $0.01 for the minute. - Translate the text yourself or with your agent. This step is not a Sume media job and carries no line in the media rate card.
- Speak each translation with
POST /v1/tts-1.0/generate, one voice per language. - Render each language: Timeline 1.0 with the original video and the new file as
audio.url, $0.10 per output minute.
Cost for three languages
Assume about 700 characters of script in each language for 45 seconds of speech, which is a planning figure, not a measured one. At $0.0475 per 1,000 characters, 700 characters is $0.033, rounded up to $0.04 per job.
| Step | Count | Each | Total |
|---|---|---|---|
| Audio detach | 1 | $0.01 | $0.01 |
| STT, 45 s billed as 1 minute | 1 | $0.01 | $0.01 |
| TTS, about 700 characters | 3 | $0.04 | $0.12 |
| Timeline render, under 1 minute | 3 | $0.10 | $0.30 |
| Total | $0.44 |
Where it goes wrong
- Send the right
languageto TTS. If the voice and the language disagree Sume returnstts_voice_language_warningand wants a confirmation before the retry, which is the guard working as designed. - Dubbed speech is a different length from the original. Measure the TTS
duration_seconds, and cut or extend the picture to match before the render. - Korean, Japanese and other long scripts count by character, spaces and punctuation included, so run a
dry_runormax_spend_usdcap on batch jobs. - Listen to a name or a number in each language before you ship; a receipt proves the input reached the job, not that it was pronounced right.
Sources
Related posts
More in Use cases
- AI micro-drama episode cost: 12 lines, two voices, $9.92 to $18.32
A 96-second micro-drama of 12 spoken lines on Sume: TTS at a cent a line, Fabric lip sync at $0.10 or $0.1875 a second, then one timeline render.
- AI music video from audio: a Lyria track, six H3 Max clips, $6.23
Build a 60-second AI music video on Sume: one Lyria track at $0.125, six 10-second MiniMax H3 Max clips at 768p, and one timeline render. About $6.23.
- AI play-by-play voiceover for a sports highlight reel: how to time it
Write one line per play, ask TTS for sentence timings, and place each clip on its line: a 60-second reel costs about 14 cents on Sume. Not live commentary.
- AI spokesperson release notes, 90 seconds: two jobs, cost by tier
A 90-second spokesperson video is two Sume avatar jobs because one job tops out at 60 seconds. Cost at standard, plus and max, and how to split the script.
Written by Sume