Tavus lists 42 languages; Sume Avatar 1.0 speaks English only

Tavus video generation lists 42 languages. Sume Avatar 1.0 speaks English only. A short decision guide based on who is listening and what you can accept.

4 min readSume
All posts

If the person on screen has to speak Spanish, Hindi or Japanese, Sume Avatar 1.0 is the wrong tool: it speaks English only. Tavus's video generation overview says you can "generate videos in 42 supported languages with your real voice" (Tavus Video Generation overview, read 2026-10-10), which is a different promise.

The Sume side of this comes from the code, not from a docs sentence. The Avatar docs describe scripts and quality tiers but never name a language. The prompt that Sume builds for each clip does: its tests pin the lines "Spoken language: English only." and "No non-English speech.", and the dialogue is passed to the model in quotation marks as exact English words. So treat English as a hard limit.

What each side states

Only facts from the vendor's own page and from Sume's docs and code appear below. The Tavus page does not list the 42 languages on the part I read, and it gives no duration limit or webhook details there, so I make no claims on those.

Language and input facts, vendor page read 2026-10-10
QuestionTavus video generationSume Avatar 1.0
Spoken languages42 supported languages (Tavus overview)English only (pinned in the clip prompt)
Voice"Your real voice" per the overview; custom audio or default text-to-speech as inputsScript-driven; the avatar speaks the script
InputA written script and a face, stock or custom-trainedAn avatar handle plus script or video_inputs
LengthNot stated on the page I read4 to 60 seconds of estimated duration
Billing unitToken usage based on video duration (Tavus overview)Per second of output, by quality tier

What English-only means for your script

Write the whole script in English. A line in another language will not be spoken as written, and the prompt tells the model to produce no non-English speech. Brand names still need care: Sume's own speech normalizer rewrites the spoken form of its own product name so that a letter-pair name is read as letters, which shows that spelling for the ear matters even in English.

A code comment in the avatar workflow notes that speech text is cleaned so caption cue breaks and script markers are not read aloud. Keep cue breaks and bracketed directions out of the script anyway, and put scene direction in the scene prompt instead.

Choose by audience, not by feature list

Language is the first filter, and it removes one of the two options for most non-English audiences. If the audience is English-speaking, the rest of the decision is about cost and workflow, which Sume's docs let you check before you spend. The Avatar Video page lists three quality tiers (standard, plus, max, with plus as the default), five aspect ratios and one resolution, 720p, so the shape of the deliverable is known before you submit.

Be equally clear with stakeholders. A common failure is a global launch plan that assumes one avatar vendor can cover every market, discovered only when the first non-English script is tested. Writing the language constraint into the brief on day one costs nothing, and it keeps the English-only avatar work on the markets where it fits.

  • Audience speaks English: use Sume, and look at the first-frame preview before the full render, which the Avatar video previews page describes.
  • Audience needs another language on a talking face: Sume cannot do it today. Evaluate Tavus on its own docs, including duration and webhook behaviour that its overview page does not state.
  • Audience is mixed: split the campaign. English cuts on Sume, other languages elsewhere, and a shared caption style across both.

Captions are not translation

Sume's Video captions page takes a language hint and styles, and it burns text onto a clip. That helps with silent autoplay, but it does not make the avatar speak. Do not read caption language options as spoken-language support.

One more edge case: Sume's captions page names Hangul styles for Korean speech, yet an avatar clip cannot contain Korean speech to caption. Those styles matter for other clips you caption separately.

Also remember that inline captions on an avatar video are optional and a soft step: the clean video is the primary output, and a caption failure does not fail the job. So an English avatar clip with captions in the same language is the supported pairing.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume