Tavus lists 42 languages; Sume Avatar 1.0 speaks English only
Tavus video generation lists 42 languages. Sume Avatar 1.0 speaks English only. A short decision guide based on who is listening and what you can accept.
If the person on screen has to speak Spanish, Hindi or Japanese, Sume Avatar 1.0 is the wrong tool: it speaks English only. Tavus's video generation overview says you can "generate videos in 42 supported languages with your real voice" (Tavus Video Generation overview, read 2026-10-10), which is a different promise.
The Sume side of this comes from the code, not from a docs sentence. The Avatar docs describe scripts and quality tiers but never name a language. The prompt that Sume builds for each clip does: its tests pin the lines "Spoken language: English only." and "No non-English speech.", and the dialogue is passed to the model in quotation marks as exact English words. So treat English as a hard limit.
What each side states
Only facts from the vendor's own page and from Sume's docs and code appear below. The Tavus page does not list the 42 languages on the part I read, and it gives no duration limit or webhook details there, so I make no claims on those.
| Question | Tavus video generation | Sume Avatar 1.0 |
|---|---|---|
| Spoken languages | 42 supported languages (Tavus overview) | English only (pinned in the clip prompt) |
| Voice | "Your real voice" per the overview; custom audio or default text-to-speech as inputs | Script-driven; the avatar speaks the script |
| Input | A written script and a face, stock or custom-trained | An avatar handle plus script or video_inputs |
| Length | Not stated on the page I read | 4 to 60 seconds of estimated duration |
| Billing unit | Token usage based on video duration (Tavus overview) | Per second of output, by quality tier |
What English-only means for your script
Write the whole script in English. A line in another language will not be spoken as written, and the prompt tells the model to produce no non-English speech. Brand names still need care: Sume's own speech normalizer rewrites the spoken form of its own product name so that a letter-pair name is read as letters, which shows that spelling for the ear matters even in English.
A code comment in the avatar workflow notes that speech text is cleaned so caption cue breaks and script markers are not read aloud. Keep cue breaks and bracketed directions out of the script anyway, and put scene direction in the scene prompt instead.
Choose by audience, not by feature list
Language is the first filter, and it removes one of the two options for most non-English audiences. If the audience is English-speaking, the rest of the decision is about cost and workflow, which Sume's docs let you check before you spend. The Avatar Video page lists three quality tiers (standard, plus, max, with plus as the default), five aspect ratios and one resolution, 720p, so the shape of the deliverable is known before you submit.
Be equally clear with stakeholders. A common failure is a global launch plan that assumes one avatar vendor can cover every market, discovered only when the first non-English script is tested. Writing the language constraint into the brief on day one costs nothing, and it keeps the English-only avatar work on the markets where it fits.
- Audience speaks English: use Sume, and look at the first-frame preview before the full render, which the Avatar video previews page describes.
- Audience needs another language on a talking face: Sume cannot do it today. Evaluate Tavus on its own docs, including duration and webhook behaviour that its overview page does not state.
- Audience is mixed: split the campaign. English cuts on Sume, other languages elsewhere, and a shared caption style across both.
Captions are not translation
Sume's Video captions page takes a language hint and styles, and it burns text onto a clip. That helps with silent autoplay, but it does not make the avatar speak. Do not read caption language options as spoken-language support.
One more edge case: Sume's captions page names Hangul styles for Korean speech, yet an avatar clip cannot contain Korean speech to caption. Those styles matter for other clips you caption separately.
Also remember that inline captions on an avatar video are optional and a soft step: the clean video is the primary output, and a caption failure does not fail the job. So an English avatar clip with captions in the same language is the supported pairing.
Sources
Related posts
More in Comparisons
- Tavus callback_url payloads: what to check before you trust one
Tavus's webhooks page lists conversation events but no signature or retry rule. How to treat them, vs Sume's signed events (read 2026-10-10).
- Text-only hero art: Imagen 4 Ultra or Nano Banana 2.1 on Sume
Imagen 4 Ultra costs 0.075 USD, takes no references and lists 5 ratios. Nano Banana 2.1 costs 0.10 and lists 15. Which to pick for 100 hero images.
- TikTok ad profile photo is 98x98 under 50 KB: what Sume cannot output
TikTok in-feed ads want a 98x98 px profile photo under 50 KB. No Sume route outputs that size, so generate the picture and resize it outside Sume.
- TikTok ad rejected for undisclosed AI: minor vs significant edits
TikTok's ad policy separates minor edits like lighting, background removal and denoising from significant AI changes. See where Sume's trim, crop and dim sit.
Written by Sume