Dub a short into Spanish with an API: detach, transcribe, speak, swap
A Spanish dub on Sume is six jobs you control: detach the audio, transcribe with sentence segments, translate yourself, speak each line, join, and render.

To dub a short video into Spanish with Sume, you run the steps yourself. Detach the audio, transcribe it with sentence segments, translate the sentences with whatever tool you trust, speak each Spanish sentence with TTS, join the lines, and render a new MP4 with the Spanish audio over the original picture. Sume does not translate for you, and that is a feature when you want to control the wording.
Steps 1 and 2: detach and transcribe
Import the video, then call POST /v1/audio-detach for a mono 16000 Hz wav. Send the result's audio_url to POST /v1/stt-1.0/transcribe with segmentation: {"mode": "sentence"} and duration_seconds. Each segment has a start and an end, so you know how long each original sentence lasts.
Steps 3 and 4: translate and speak
Translate each sentence outside Sume. Then send each one to POST /v1/tts-1.0/generate with language: "es" and a voice whose language tag is Spanish. A voice with a different tag returns 409 tts_voice_language_mismatch before any charge. Use one idempotency key per sentence, derived from the sentence index and text, so a retry does not create a duplicate.
Check the timing
Spanish often runs longer than English, so compare each spoken file's length with the original segment before you join. Fix the lines that overrun by shortening the translation, not by speeding up the voice.
Steps 5 and 6: join and render
Join the lines with timeline audio concat (1 to 20 parts, sample-domain, no re-synthesis). Then render with Timeline 1.0, passing the joined file as audio.url and the original video as the video[] slot. Run POST /v1/timeline-1.0/plan first for the estimate.
The six jobs, read 2026-10-06:
| Step | Endpoint | Public rate |
|---|---|---|
| Detach audio | POST /v1/audio-detach | $0.01 per job |
| Transcribe | POST /v1/stt-1.0/transcribe | $0.01 per audio minute |
| Speak each line | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Join lines | POST /v1/timeline-1.0/audio | $0.01 per job |
| Render | POST /v1/timeline-1.0/render | $0.10 per output minute |
What a dub still needs from a person
A machine transcript and a machine translation are drafts. Ask a Spanish speaker to read the final lines, and to listen to the voice. Check names, units and anything that sounds formal in one language and casual in the other.
Disclose the dub where a platform asks for it. If the video also has burned English text, plan a Spanish version of that text as well, because the audio swap does not touch the picture.
Captions are a separate job. Burn Spanish captions from the Spanish script text, not from the English source, and confirm the live rates in GET /v1/catalog.
Sources
Related posts
More in Use cases
- Product hero video: pin the first frame to your own packshot
Use frame_images with first_frame on sume/auto so a 6-second product video opens on your exact packshot, not a regenerated lookalike.
- EU Article 50 AI avatar video disclosure checklist for marketers
Tavus says EU Article 50 has required disclosure of AI likenesses since 2 August 2026. A checklist of what to label on an avatar clip, with the Sume fields.
- Feature a YouTube show on your Home tab: single or multiple shelf
YouTube's help adds a show to Home as a Single playlist or Multiple playlists shelf, and shelves can be dragged into order. What to build first.
- Gemini Omni extension only appends: how to add a lead-in clip
Google's Omni docs say extension adds to the end of a clip only. To get a clip before your footage, generate it separately and join on a Sume timeline.
Written by Sume