Wistia Localize translates the transcript: same order on Sume

Wistia lists $2 per minute per language and translates the editable transcript, not the audio. Here is that transcript-first order built from Sume steps.

5 min readSume
All posts

What does Wistia Localize do, and how does Sume compare?

Wistia's Localize page lists more than 50 languages, a price of $2 per minute per language with the first 15 minutes free, and three deliverables: dubbed audio with voice cloning, editable transcripts and captions in the target languages. It also lists lip syncing. Sume does not sell a one-click dubbing product; it gives you the individual steps (speech-to-text, text-to-speech, timeline audio, burned captions) as separate jobs you chain yourself.

The detail worth copying from Wistia is how it works: the page says it translates your videos using the editable transcript, not the audio. That means a human can fix the words before any voice is generated. You can run the same order on Sume, and it is the cheapest place to catch a wrong product name.

Why is transcript-first the safer order?

A mistranscribed word costs almost nothing to fix as text and a full re-synthesis to fix as audio. A transcript-first pipeline puts every correction before the expensive steps.

It also gives you a natural review point. The transcript can go to a translator, a glossary check or an agent step, and you only generate the voice once the text is signed off. With Sume that review point is simply the gap between two API calls, and nothing is billed while you wait.

The trade-off is tone. Translating text drops anything that lived only in the delivery: a laugh, a pause, emphasis. If performance carries the video, read the post on audio-to-audio dubbing and decide whether the fidelity is worth the price.

How do you build the transcript step on Sume?

Detach the audio from your hosted video, then transcribe it with sentence segmentation so you get editable lines with start and end times. STT takes a public HTTPS audio_url, an optional language_code hint and an optional duration_seconds (1 to 600) used for the usage reservation; the maximum is 10 minutes per job, so cut longer recordings with an audio-detach range.

Word timings always come back, and segmentation: {"mode": "sentence"} adds sentence segments you can hand to a translator one line at a time.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: wistia-style-transcript-001" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/artf_demo/talk.wav",
    "language_code": "en",
    "duration_seconds": 300,
    "segmentation": { "mode": "sentence" }
  }'

What happens after the transcript is approved?

Send each approved line to TTS with language set to the target language, using the voice you chose. Then join the lines with timeline audio concat (up to 20 parts, sample-domain join, no silence at the seams) or place them directly as audio.parts[] in a Timeline 1.0 render. For captions, burn the approved translated lines as authored cues with video captions, which skips a second transcription.

Three things Wistia bundles are not in this chain: voice cloning from the speaker, an editor UI for the transcript, and lip syncing of the original footage. On Sume you pick a voice from the TTS options and you review text in your own tool or with an agent. Our step-by-step cost breakdown puts the pieces side by side.

What does the arithmetic look like?

Wistia's number is easy to project because it is flat: a 10-minute video into four languages is 10 x 4 x $2 = $80 before the free 15 minutes. Sume has no single dubbing price, so you add up the steps. Speech-to-text is listed at $0.01 per audio minute and runs once however many languages follow; each language then adds one TTS job priced per character of translated text, plus a caption job if you burn captions ($0.20 per job up to 60 seconds). Confirm live rates in GET /v1/catalog.

Cost shape of one 10-minute video into four languages, read 2026-10-02
StepWistia LocalizeSume
TranscribeIncluded in the per-language priceOnce, per audio minute
TranslateIncludedYour own step or agent run
VoiceIncluded, with cloningTTS per character, per language
CaptionsIncluded$0.20 per caption job up to 60 seconds
Lip syncListedNot for existing footage

Which should you pick?

If you want a hosted editor, voice cloning and lip sync with one invoice line, Wistia's flat price is easier to budget. If you already run an automation, want the text reviewed by your own glossary or agent before any audio exists, and are happy to chain jobs with webhooks, building the transcript-first order on Sume gives you a per-step price and keeps every intermediate file.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume