CapCut AI Video Translator not in your region? Script a dub

CapCut's AI Video Translator page says availability depends on region and claims lip sync. Sume documents each step of a dubbing pipeline you can call anywhere.

5 min readSume
All posts

CapCut's AI Video Translator page says the feature is available in specific regions and advertises lip sync, and it gives no length limits. If it is not offered where you are, Sume's pieces can be called over HTTPS: extract audio, transcribe with stt_create, translate the text, synthesize with tts_create, and caption or assemble on a timeline. Availability of any Sume feature depends on your account and key, not a CapCut region list.

The pieces and what each costs to learn

audio-detach makes a 16000 Hz mono WAV, the speech-to-text shape, for $0.01 per job; the source is limited to 1800 seconds and output to 900. Speech-to-text and text-to-speech are MCP tools in the docs list, and the TTS route checks the voice language against the target before charging. Confirm all rates in GET /v1/catalog.

CapCut page compared with a scripted pipeline, read 2026-10-03
QuestionCapCut AI Video TranslatorScripted Sume pipeline
AvailabilitySpecific regions, per the pageAPI; depends on your key
Lip syncClaimed on the pageNot in this pipeline
Length limitsNone stated on the page1800 s source, 900 s output per detach
StepsOne toolYou chain four or five calls

Translation is yours to choose

Sume does not translate inside a caption or TTS call. You pick the translator, and you can pass the result to captions as script_text or cues. That means you can review the text before it is voiced.

Honest limits

  • No lip sync here, so mouth movement will not match the new language.
  • CapCut's page gives no pricing or limits in the text read, so none are compared.
  • You are responsible for rights to translate and voice the content.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume