Dub a short into Spanish with an API: detach, transcribe, speak, swap

A Spanish dub on Sume is six jobs you control: detach the audio, transcribe with sentence segments, translate yourself, speak each line, join, and render.

5 min readSume
All posts

To dub a short video into Spanish with Sume, you run the steps yourself. Detach the audio, transcribe it with sentence segments, translate the sentences with whatever tool you trust, speak each Spanish sentence with TTS, join the lines, and render a new MP4 with the Spanish audio over the original picture. Sume does not translate for you, and that is a feature when you want to control the wording.

Steps 1 and 2: detach and transcribe

Import the video, then call POST /v1/audio-detach for a mono 16000 Hz wav. Send the result's audio_url to POST /v1/stt-1.0/transcribe with segmentation: {"mode": "sentence"} and duration_seconds. Each segment has a start and an end, so you know how long each original sentence lasts.

Steps 3 and 4: translate and speak

Translate each sentence outside Sume. Then send each one to POST /v1/tts-1.0/generate with language: "es" and a voice whose language tag is Spanish. A voice with a different tag returns 409 tts_voice_language_mismatch before any charge. Use one idempotency key per sentence, derived from the sentence index and text, so a retry does not create a duplicate.

Check the timing

Spanish often runs longer than English, so compare each spoken file's length with the original segment before you join. Fix the lines that overrun by shortening the translation, not by speeding up the voice.

Steps 5 and 6: join and render

Join the lines with timeline audio concat (1 to 20 parts, sample-domain, no re-synthesis). Then render with Timeline 1.0, passing the joined file as audio.url and the original video as the video[] slot. Run POST /v1/timeline-1.0/plan first for the estimate.

The six jobs, read 2026-10-06:

Spanish dub jobs, read 2026-10-06
StepEndpointPublic rate
Detach audioPOST /v1/audio-detach$0.01 per job
TranscribePOST /v1/stt-1.0/transcribe$0.01 per audio minute
Speak each linePOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
Join linesPOST /v1/timeline-1.0/audio$0.01 per job
RenderPOST /v1/timeline-1.0/render$0.10 per output minute

What a dub still needs from a person

A machine transcript and a machine translation are drafts. Ask a Spanish speaker to read the final lines, and to listen to the voice. Check names, units and anything that sounds formal in one language and casual in the other.

Disclose the dub where a platform asks for it. If the video also has burned English text, plan a Spanish version of that text as well, because the audio swap does not touch the picture.

Captions are a separate job. Burn Spanish captions from the Spanish script text, not from the English source, and confirm the live rates in GET /v1/catalog.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume