Inworld dubbing API: 200+ languages and rights, vs a Sume pipeline
Inworld's dubbing page claims 200+ languages, clones from 5-15 seconds and says rights clearance sits with you. What a Sume dubbing pipeline covers instead.

Inworld's dubbing page describes cross-lingual dubbing in a cloned voice across 200+ languages and says plainly that cloning requires rights to the source audio, with rights clearance sitting with you (Inworld, read 2026-10-02). Sume has no one-call dubbing route and no cross-lingual clone, so a Sume dub is a chain of separate jobs.
What Inworld says
The page lists instant cloning from 5 to 15 seconds of reference audio and professional clones from 30+ minutes, with timbre, pacing and character kept across languages. The endpoint is described as non-streaming, so it works as an asynchronous dubbing job. The page does not mention lip sync or give file size limits.
| Item | Claim on the page |
|---|---|
| Languages | 200+ with Realtime TTS-2 coverage |
| Instant clone | 5 to 15 seconds of reference audio |
| Professional clone | 30+ minutes of source material |
| Output formats | MP3, WAV, PCM, LINEAR16, OGG_OPUS, mu-law, A-law, FLAC |
| Sample rates | 8 to 48 kHz |
| Rights | Cloning requires rights to the source audio; clearance sits with you |
| Lip sync | Not mentioned |
The Sume chain
Sume has no one-call dubbing route. The pieces are audio detach to pull the audio from a video, speech to text at POST /v1/stt-1.0/transcribe, your own translation, text to speech, and Timeline audio to join the lines.
Each step is a separate job with its own status URL, and text to speech is billed per character. The voice for the target language is a library voice you pick whose language matches the text.
Where the two differ
The differences come down to who owns each step.
- The target voice is not an automatic clone of the original speaker; you pick or clone a voice yourself.
- Translation happens outside Sume's audio API, in whatever model or translator you use.
- Sume's speech to text takes up to 600 seconds per job, so longer videos are split first.
- Sume does not re-sync lips in existing footage; for a talking face, the lip-sync routes start from a still plus audio.
When to pick which
Choose a hosted dubbing API when you want speaker voice preserved across many languages in one call. Choose the Sume chain when you want control of each step, including pause timing from sentence segments. The dubbing pipeline post lists the routes and limits in order.
Sources
Related posts
More in Comparisons
- iStock AI generator: $10,000 indemnity, no resale or print-on-demand
iStock's generator FAQ gives standard legal indemnification up to $10,000 and says the Standard License does not allow resale products or print-on-demand.
- Kapwing subtitle minutes vs Sume caption job price
Kapwing Pro includes 1,000 auto-subtitle minutes for $16 to $24 a month. Sume charges $0.20 per caption job for a video up to 60 seconds. Dated 2026-10-01.
- Kie.ai API alternative: callbacks, credits and price vs Sume
Kie.ai says it prices 30 to 50 percent below official APIs and returns a task id on HTTP 200. Sume reserves list x 1.25. Here is how the two compare.
- Kling 4.0 'more natural lip sync': what you can call on Sume today
Kling's 4.0 page promises more natural lip sync and tighter A/V sync. Sume lists no Kling 4.0 id; for lip sync today use Fabric or H3 Max Lip Sync.
Written by Sume