Voice from a short clip: Voxtral 5-25 s, MAI, Sume avatar voice
How much reference audio Mistral and Microsoft ask for to match a voice, and how Sume selects a voice instead: avatar_id, avatar_handle or a voice id.
Mistral says Voxtral TTS can adapt to a new voice from as little as 3 seconds of reference audio, while its model accepts voice prompts between 5 and 25 seconds. Microsoft says MAI-Voice-2.1 offers instant voice matching from short reference clips from a short reference clip without fine-tuning; its page does not give a duration. Sume works the other way around: TTS 1.0 selects an existing voice through avatar_id, avatar_handle or a voice id, and the docs read for this post show no endpoint for cloning a voice from a clip.
What the vendors state
The numbers are from the vendors' own pages. Neither says how good a result from the minimum will be.
| Service | Reference audio stated | Other notes on the page |
|---|---|---|
| Mistral Voxtral TTS | As little as 3 seconds to capture speaker traits; model accepts voice prompts of 5 to 25 seconds | Open weights under CC BY-NC 4.0; API $0.016 per 1,000 characters |
| Microsoft MAI-Voice-2.1 and Flash | A short reference clip, no fine-tuning; no duration stated | $22 (MAI-Voice-2.1) and $15 (Flash) per million characters |
| Sume TTS 1.0 | No clip upload documented; voice comes from an avatar or a voice id | Voice id must be a UUID or a Voices library id |
How Sume picks a voice
The OpenAPI description for POST /v1/tts-1.0/generate says to provide a transcript plus a voice selector: a top-level avatar_id or avatar_handle, or voice.id. The discoverable route is the avatar list, GET /v1/avatar-1.0/avatars, and any avatar whose voice.status is ready can speak. If you pass both an avatar and a voice.id, they must match, otherwise the request fails with 400. A voice name from another ecosystem is rejected before any credit is reserved (invalid_voice_id).
This is a narrower promise than cloning. It gives you a stable voice per avatar, so an episode series keeps the same narrator, as shown in reading the voice from the TTS job.
Which to choose
Pick by what you must do, not by the shortest clip.
- You must speak as a specific real person who consented: use a service that documents cloning and keep the consent record.
- You want a consistent brand narrator across many videos: a stored voice with an id is simpler than re-cloning each time.
- You need non-commercial research on open weights: Voxtral's weights are the option, within CC BY-NC.
- You need the speech inside a larger pipeline of captions and timeline: use a voice Sume can resolve, as described in the basics page.
A note on honesty
Voice cloning raises consent and disclosure questions regardless of the vendor. The wider list of cloning options, and what Sume does not offer, is in best voice cloning TTS in 2026. The models section of the Sume docs lists what the platform generates today.
Sources
Related posts
More in Comparisons
- Vyond Starter $58: 25-minute cap vs Sume 60-second avatar jobs
Vyond lists $58 a month for Starter with 25-40 minute videos. Here is what 25 and 40 minutes of avatar video cost on Sume in 60-second jobs, at each tier.
- Wan 3.0 vs Seedance 2.5 price gap at 720p and 1080p, 5 to 30 seconds
On Sume Wan 3.0 costs a flat $0.125 per second at 720p. Seedance 2.5 costs about 4.6 times as much. Table for 5, 10, 15 and 30 seconds at 720p and 1080p.
- Live captions vs rendered avatar clip captions: WCAG 1.2.4
Full-duplex video AI is live, so WCAG 1.2.4 applies. A rendered avatar clip is prerecorded and follows 1.2.2 instead. Where Sume captions fit.
- xAI batch mixes chat, image and video in one file; Sume uses queues
xAI's Batch API now takes chat, image and video requests in one JSONL file. A Sume bulk queue targets exactly one Format. How to split a holiday job.
Written by Sume