Voice from a short clip: Voxtral 5-25 s, MAI, Sume avatar voice

How much reference audio Mistral and Microsoft ask for to match a voice, and how Sume selects a voice instead: avatar_id, avatar_handle or a voice id.

5 min readSume
All posts

Mistral says Voxtral TTS can adapt to a new voice from as little as 3 seconds of reference audio, while its model accepts voice prompts between 5 and 25 seconds. Microsoft says MAI-Voice-2.1 offers instant voice matching from short reference clips from a short reference clip without fine-tuning; its page does not give a duration. Sume works the other way around: TTS 1.0 selects an existing voice through avatar_id, avatar_handle or a voice id, and the docs read for this post show no endpoint for cloning a voice from a clip.

What the vendors state

The numbers are from the vendors' own pages. Neither says how good a result from the minimum will be.

Reference audio for voice matching, as stated by each vendor (read 2026-10-04)
ServiceReference audio statedOther notes on the page
Mistral Voxtral TTSAs little as 3 seconds to capture speaker traits; model accepts voice prompts of 5 to 25 secondsOpen weights under CC BY-NC 4.0; API $0.016 per 1,000 characters
Microsoft MAI-Voice-2.1 and FlashA short reference clip, no fine-tuning; no duration stated$22 (MAI-Voice-2.1) and $15 (Flash) per million characters
Sume TTS 1.0No clip upload documented; voice comes from an avatar or a voice idVoice id must be a UUID or a Voices library id

How Sume picks a voice

The OpenAPI description for POST /v1/tts-1.0/generate says to provide a transcript plus a voice selector: a top-level avatar_id or avatar_handle, or voice.id. The discoverable route is the avatar list, GET /v1/avatar-1.0/avatars, and any avatar whose voice.status is ready can speak. If you pass both an avatar and a voice.id, they must match, otherwise the request fails with 400. A voice name from another ecosystem is rejected before any credit is reserved (invalid_voice_id).

This is a narrower promise than cloning. It gives you a stable voice per avatar, so an episode series keeps the same narrator, as shown in reading the voice from the TTS job.

Which to choose

Pick by what you must do, not by the shortest clip.

  • You must speak as a specific real person who consented: use a service that documents cloning and keep the consent record.
  • You want a consistent brand narrator across many videos: a stored voice with an id is simpler than re-cloning each time.
  • You need non-commercial research on open weights: Voxtral's weights are the option, within CC BY-NC.
  • You need the speech inside a larger pipeline of captions and timeline: use a voice Sume can resolve, as described in the basics page.

A note on honesty

Voice cloning raises consent and disclosure questions regardless of the vendor. The wider list of cloning options, and what Sume does not offer, is in best voice cloning TTS in 2026. The models section of the Sume docs lists what the platform generates today.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume