CapCut AI Video Translator not in your region? Script a dub
CapCut's AI Video Translator page says availability depends on region and claims lip sync. Sume documents each step of a dubbing pipeline you can call anywhere.

CapCut's AI Video Translator page says the feature is available in specific regions and advertises lip sync, and it gives no length limits. If it is not offered where you are, Sume's pieces can be called over HTTPS: extract audio, transcribe with stt_create, translate the text, synthesize with tts_create, and caption or assemble on a timeline. Availability of any Sume feature depends on your account and key, not a CapCut region list.
The pieces and what each costs to learn
audio-detach makes a 16000 Hz mono WAV, the speech-to-text shape, for $0.01 per job; the source is limited to 1800 seconds and output to 900. Speech-to-text and text-to-speech are MCP tools in the docs list, and the TTS route checks the voice language against the target before charging. Confirm all rates in GET /v1/catalog.
| Question | CapCut AI Video Translator | Scripted Sume pipeline |
|---|---|---|
| Availability | Specific regions, per the page | API; depends on your key |
| Lip sync | Claimed on the page | Not in this pipeline |
| Length limits | None stated on the page | 1800 s source, 900 s output per detach |
| Steps | One tool | You chain four or five calls |
Translation is yours to choose
Sume does not translate inside a caption or TTS call. You pick the translator, and you can pass the result to captions as script_text or cues. That means you can review the text before it is voiced.
Honest limits
- No lip sync here, so mouth movement will not match the new language.
- CapCut's page gives no pricing or limits in the text read, so none are compared.
- You are responsible for rights to translate and voice the content.
Sources
Related posts
More in Use cases
- Captions when speakers switch languages: Section 508's (in Spanish)
Section 508 says to mark a language change like (in Spanish) and keep the exact words when speech is untranslated. How to burn that with authored Sume cues.
- Captions on AI video with native sound: what Sume actually transcribes
Gemini Omni Flash 1.1 always generates synced audio on Sume. The captions endpoint transcribes speech only, so sound effects need authored cues.
- Should captions censor profanity? Section 508 and YouTube's [ __ ]
Section 508 says caption audible profanity exactly; YouTube's auto captions swap flagged words for [ __ ] by default. How to burn the wording you intend.
- Unreadable captions on bright Reels: dim 0.45, check for free
Burned-in captions vanish on bright footage. Sume video filter dim 0.45 darkens luma; the /video-filter/check endpoint validates the program at no charge.
Written by Sume