Voxtral TTS 2-minute native limit vs Sume TTS 1200-second cap

Mistral says Voxtral TTS natively generates up to two minutes and the API handles longer. Sume TTS fails audio over 1200 seconds with tts_duration_exceeded.

5 min readSume
All posts

Mistral's Voxtral TTS page says the model natively generates up to two minutes of audio and that its API handles longer text. Sume's TTS has no per-call native window to manage but a hard result cap: audio over 1200 seconds fails with tts_duration_exceeded, and the transcript limit is 20,000 characters. Both ceilings mean you plan for splitting long scripts.

Voxtral facts are from Mistral's Voxtral TTS announcement, read on 2026-10-02. Sume's are from the TTS and TTS Router docs.

What does Mistral say about Voxtral's limits?

The page describes a 4B-parameter model with nine languages (English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic), a 70 ms latency figure, API pricing of $0.016 per 1,000 characters at /v1/audio/speech, and voice cloning from a reference as short as three seconds, with 5 to 25 second prompts supported. Open weights are licensed CC BY-NC 4.0. On length it says the model natively generates up to two minutes and the API handles longer generation through smart interleaving.

The page does not say how the API splits text or what the hard maximum is, so that is not claimed here.

What are Sume's limits?

POST /v1/tts-1.0/generate accepts a transcript of 1 to 20,000 characters. If synthesized audio would run past 1200 seconds, the job fails with tts_duration_exceeded. POST /v1/tts-router/generate requires a model from GET /v1/tts-router/models, and every catalog row lists max_characters 20,000 with per-character pricing. The voice comes from avatar_id, avatar_handle or voice.id; the TTS schema has no clone-from-audio field, so a voice must already exist in your library.

How do the limits compare?

Voxtral facts from Mistral's page; Sume facts from the TTS and TTS Router docs, read 2026-10-02.
LimitVoxtral TTSSume TTS
Native audio per generationUp to 2 minutesNot stated per call
Longer than thatAPI handles longer, per MistralFails over 1200 s with tts_duration_exceeded
Text per requestNot stated on the page20,000 characters
Price$0.016 per 1,000 charactersPer-character list price in the router catalog, with Sume's margin
Clone reference3 s minimumNo clone field in the TTS schema

What does the 1200-second cap look like in practice?

Twenty minutes of speech and 20,000 characters are two separate ceilings, and which one hits first depends on the script and the voice's pace. Check both before you submit: the character count at request time and the duration after synthesis.

The failure arrives on the job, not at submit. A tts_duration_exceeded result means the job was accepted and then failed, so a script that submits whole chapters should read the job status, not assume success. The jobs and results page shows what a completed or failed job records.

What about the two-minute figure on Mistral's side?

Two minutes is the native window of the model itself, and Mistral says the API goes beyond it. For a buyer, the practical question is whether seams between internal segments are audible, which the page does not measure and this post does not claim either way. If you split scripts yourself on either service, cut at sentence ends and listen to the join.

How do you split a long script safely?

Cut on paragraph boundaries so each part stays well under 20,000 characters and comfortably under 1200 seconds, submit parts as separate jobs with distinct Idempotency-Key values, then join the audio with Timeline audio concat, which takes up to 20 parts. Keep the same voice and generation_config across parts so the joined chapter sounds consistent.

Sume does not offer an open-weights model or a self-hosted route; Voxtral's CC BY-NC 4.0 weights are for non-commercial use, so check the licence before running them in a product. See the weights licence post.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume