Voxtral TTS 2-minute native limit vs Sume TTS 1200-second cap
Mistral says Voxtral TTS natively generates up to two minutes and the API handles longer. Sume TTS fails audio over 1200 seconds with tts_duration_exceeded.

Mistral's Voxtral TTS page says the model natively generates up to two minutes of audio and that its API handles longer text. Sume's TTS has no per-call native window to manage but a hard result cap: audio over 1200 seconds fails with tts_duration_exceeded, and the transcript limit is 20,000 characters. Both ceilings mean you plan for splitting long scripts.
Voxtral facts are from Mistral's Voxtral TTS announcement, read on 2026-10-02. Sume's are from the TTS and TTS Router docs.
What does Mistral say about Voxtral's limits?
The page describes a 4B-parameter model with nine languages (English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic), a 70 ms latency figure, API pricing of $0.016 per 1,000 characters at /v1/audio/speech, and voice cloning from a reference as short as three seconds, with 5 to 25 second prompts supported. Open weights are licensed CC BY-NC 4.0. On length it says the model natively generates up to two minutes and the API handles longer generation through smart interleaving.
The page does not say how the API splits text or what the hard maximum is, so that is not claimed here.
What are Sume's limits?
POST /v1/tts-1.0/generate accepts a transcript of 1 to 20,000 characters. If synthesized audio would run past 1200 seconds, the job fails with tts_duration_exceeded. POST /v1/tts-router/generate requires a model from GET /v1/tts-router/models, and every catalog row lists max_characters 20,000 with per-character pricing. The voice comes from avatar_id, avatar_handle or voice.id; the TTS schema has no clone-from-audio field, so a voice must already exist in your library.
How do the limits compare?
| Limit | Voxtral TTS | Sume TTS |
|---|---|---|
| Native audio per generation | Up to 2 minutes | Not stated per call |
| Longer than that | API handles longer, per Mistral | Fails over 1200 s with tts_duration_exceeded |
| Text per request | Not stated on the page | 20,000 characters |
| Price | $0.016 per 1,000 characters | Per-character list price in the router catalog, with Sume's margin |
| Clone reference | 3 s minimum | No clone field in the TTS schema |
What does the 1200-second cap look like in practice?
Twenty minutes of speech and 20,000 characters are two separate ceilings, and which one hits first depends on the script and the voice's pace. Check both before you submit: the character count at request time and the duration after synthesis.
The failure arrives on the job, not at submit. A tts_duration_exceeded result means the job was accepted and then failed, so a script that submits whole chapters should read the job status, not assume success. The jobs and results page shows what a completed or failed job records.
What about the two-minute figure on Mistral's side?
Two minutes is the native window of the model itself, and Mistral says the API goes beyond it. For a buyer, the practical question is whether seams between internal segments are audible, which the page does not measure and this post does not claim either way. If you split scripts yourself on either service, cut at sentence ends and listen to the join.
How do you split a long script safely?
Cut on paragraph boundaries so each part stays well under 20,000 characters and comfortably under 1200 seconds, submit parts as separate jobs with distinct Idempotency-Key values, then join the audio with Timeline audio concat, which takes up to 20 parts. Keep the same voice and generation_config across parts so the joined chapter sounds consistent.
Sume does not offer an open-weights model or a self-hosted route; Voxtral's CC BY-NC 4.0 weights are for non-commercial use, so check the licence before running them in a product. See the weights licence post.
Sources
Related posts
More in Comparisons
- WaveSpeed task statuses (timeout, deleted) vs Sume job statuses
WaveSpeed tasks can be created, processing, completed, failed, cancelled, timeout or deleted. Sume has five job statuses. How to map them in a port.
- WaveSpeedAI API alternative: what Sume offers instead
WaveSpeedAI sells 1,000+ models behind one API with tiered rate limits. Sume offers a smaller managed catalog with plan-based concurrency. Compared.
- Best Sume image model for text in images: five catalog ids compared
Vendors pitch text rendering: Ideogram typography, FLUX.2 flex, Qwen-Image 2.0 Pro, Seedream. Which of those ids does Sume list, and how to test them yourself.
- YouTube Studio clips and Shorts tool vs Sume trim and captions
YouTube's clips tool cuts Shorts from your long videos inside Studio. Sume does the same file work by API, with trim, captions and Timeline. When to use which.
Written by Sume