Wistia Localize translates the transcript: same order on Sume
Wistia lists $2 per minute per language and translates the editable transcript, not the audio. Here is that transcript-first order built from Sume steps.

What does Wistia Localize do, and how does Sume compare?
Wistia's Localize page lists more than 50 languages, a price of $2 per minute per language with the first 15 minutes free, and three deliverables: dubbed audio with voice cloning, editable transcripts and captions in the target languages. It also lists lip syncing. Sume does not sell a one-click dubbing product; it gives you the individual steps (speech-to-text, text-to-speech, timeline audio, burned captions) as separate jobs you chain yourself.
The detail worth copying from Wistia is how it works: the page says it translates your videos using the editable transcript, not the audio. That means a human can fix the words before any voice is generated. You can run the same order on Sume, and it is the cheapest place to catch a wrong product name.
Why is transcript-first the safer order?
A mistranscribed word costs almost nothing to fix as text and a full re-synthesis to fix as audio. A transcript-first pipeline puts every correction before the expensive steps.
It also gives you a natural review point. The transcript can go to a translator, a glossary check or an agent step, and you only generate the voice once the text is signed off. With Sume that review point is simply the gap between two API calls, and nothing is billed while you wait.
The trade-off is tone. Translating text drops anything that lived only in the delivery: a laugh, a pause, emphasis. If performance carries the video, read the post on audio-to-audio dubbing and decide whether the fidelity is worth the price.
How do you build the transcript step on Sume?
Detach the audio from your hosted video, then transcribe it with sentence segmentation so you get editable lines with start and end times. STT takes a public HTTPS audio_url, an optional language_code hint and an optional duration_seconds (1 to 600) used for the usage reservation; the maximum is 10 minutes per job, so cut longer recordings with an audio-detach range.
Word timings always come back, and segmentation: {"mode": "sentence"} adds sentence segments you can hand to a translator one line at a time.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: wistia-style-transcript-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/talk.wav",
"language_code": "en",
"duration_seconds": 300,
"segmentation": { "mode": "sentence" }
}'What happens after the transcript is approved?
Send each approved line to TTS with language set to the target language, using the voice you chose. Then join the lines with timeline audio concat (up to 20 parts, sample-domain join, no silence at the seams) or place them directly as audio.parts[] in a Timeline 1.0 render. For captions, burn the approved translated lines as authored cues with video captions, which skips a second transcription.
Three things Wistia bundles are not in this chain: voice cloning from the speaker, an editor UI for the transcript, and lip syncing of the original footage. On Sume you pick a voice from the TTS options and you review text in your own tool or with an agent. Our step-by-step cost breakdown puts the pieces side by side.
What does the arithmetic look like?
Wistia's number is easy to project because it is flat: a 10-minute video into four languages is 10 x 4 x $2 = $80 before the free 15 minutes. Sume has no single dubbing price, so you add up the steps. Speech-to-text is listed at $0.01 per audio minute and runs once however many languages follow; each language then adds one TTS job priced per character of translated text, plus a caption job if you burn captions ($0.20 per job up to 60 seconds). Confirm live rates in GET /v1/catalog.
| Step | Wistia Localize | Sume |
|---|---|---|
| Transcribe | Included in the per-language price | Once, per audio minute |
| Translate | Included | Your own step or agent run |
| Voice | Included, with cloning | TTS per character, per language |
| Captions | Included | $0.20 per caption job up to 60 seconds |
| Lip sync | Listed | Not for existing footage |
Which should you pick?
If you want a hosted editor, voice cloning and lip sync with one invoice line, Wistia's flat price is easier to budget. If you already run an automation, want the text reviewed by your own glossary or agent before any audio exists, and are happy to chain jobs with webhooks, building the transcript-first order on Sume gives you a per-step price and keeps every intermediate file.
Sources
Related posts
More in Comparisons
- YouTube Edit with AI limits: Shorts app vs Create app
Edit with AI takes up to 3 minutes of assets in the Shorts camera and 5-second to 5-minute clips in YouTube Create. Limits side by side, plus Sume Timeline.
- YouTube Studio clips and Shorts tool vs Sume trim and captions
YouTube's clips tool cuts Shorts from your long videos inside Studio. Sume does the same file work by API, with trim, captions and Timeline. When to use which.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
Written by Sume