Inworld dubbing API: 200+ languages and rights, vs a Sume pipeline

Inworld's dubbing page claims 200+ languages, clones from 5-15 seconds and says rights clearance sits with you. What a Sume dubbing pipeline covers instead.

5 min readSume
All posts

Inworld's dubbing page describes cross-lingual dubbing in a cloned voice across 200+ languages and says plainly that cloning requires rights to the source audio, with rights clearance sitting with you (Inworld, read 2026-10-02). Sume has no one-call dubbing route and no cross-lingual clone, so a Sume dub is a chain of separate jobs.

What Inworld says

The page lists instant cloning from 5 to 15 seconds of reference audio and professional clones from 30+ minutes, with timbre, pacing and character kept across languages. The endpoint is described as non-streaming, so it works as an asynchronous dubbing job. The page does not mention lip sync or give file size limits.

Inworld dubbing page, read 2026-10-02
ItemClaim on the page
Languages200+ with Realtime TTS-2 coverage
Instant clone5 to 15 seconds of reference audio
Professional clone30+ minutes of source material
Output formatsMP3, WAV, PCM, LINEAR16, OGG_OPUS, mu-law, A-law, FLAC
Sample rates8 to 48 kHz
RightsCloning requires rights to the source audio; clearance sits with you
Lip syncNot mentioned

The Sume chain

Sume has no one-call dubbing route. The pieces are audio detach to pull the audio from a video, speech to text at POST /v1/stt-1.0/transcribe, your own translation, text to speech, and Timeline audio to join the lines.

Each step is a separate job with its own status URL, and text to speech is billed per character. The voice for the target language is a library voice you pick whose language matches the text.

Where the two differ

The differences come down to who owns each step.

  • The target voice is not an automatic clone of the original speaker; you pick or clone a voice yourself.
  • Translation happens outside Sume's audio API, in whatever model or translator you use.
  • Sume's speech to text takes up to 600 seconds per job, so longer videos are split first.
  • Sume does not re-sync lips in existing footage; for a talking face, the lip-sync routes start from a still plus audio.

When to pick which

Choose a hosted dubbing API when you want speaker voice preserved across many languages in one call. Choose the Sume chain when you want control of each step, including pause timing from sentence segments. The dubbing pipeline post lists the routes and limits in order.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume