VibeVoice 90-minute TTS vs Sume's 1,200-second jobs: long audio
VibeVoice-TTS-1.5B targets 90 minutes and 4 speakers, and Microsoft removed its TTS code. What that means for long narration, and how Sume chunks it in jobs.

VibeVoice-TTS-1.5B is described as synthesizing up to 90 minutes of conversational speech with up to four speakers in one pass, but Microsoft has removed the TTS code after misuse, so you cannot count on running it. Sume takes the opposite approach to long audio: each job is capped at 20,000 characters and 1,200 seconds, and you assemble a long recording from several jobs.
VibeVoice details come from the Microsoft repository, read on 2026-10-02. Sume details come from the Sume API reference and the timeline audio guide.
What does the VibeVoice repository say?
It describes a family of open-source voice models under the MIT license. VibeVoice-TTS-1.5B synthesizes conversational speech up to 90 minutes with up to 4 distinct speakers. VibeVoice-Realtime-0.5B streams with roughly 300 ms latency and supports about 10 minutes of generation. VibeVoice-ASR-7B handles 60 minutes of audio in a single pass, with speaker identification, timestamps and hotword support across 50 or more languages.
The repository also states that Microsoft removed the VibeVoice-TTS code after discovering misuse inconsistent with the research intent, warns about deepfakes and disinformation, and recommends disclosing AI-generated content. Treat the 90-minute number as a research capability, not a service you can call.
How does Sume handle a long script?
A single Sume TTS request takes 1 to 20,000 characters, and a synthesized result longer than 1,200 seconds fails with tts_duration_exceeded without capturing credit. So a 90-minute recording is several jobs, split at natural breaks such as chapters or sections. The catalog price is $0.0475 per 1,000 characters, and spaces and punctuation count.
Join the finished takes into one file with timeline audio: operation: concat accepts 1 to 20 parts, joins them in the sample domain with no gap at the seams, and costs a flat $0.01 per job. Every part must already be Sume-hosted audio and share one channel layout, and the produced file is capped at 1,800 seconds.
| Limit | Value | What to do |
|---|---|---|
| Characters per TTS request | 20,000 | Split the script at paragraph or chapter breaks |
| Audio per TTS job | 1,200 seconds | Keep each chunk well under 20 minutes of speech |
| Parts per concat job | 20 | Concat in two stages for more than 20 takes |
| Joined file length | 1,800 seconds | Joined files longer than 30 minutes need another plan |
What would a long recording cost?
Take a 60,000-character script. At $0.0475 per 1,000 characters that is 60 x $0.0475 = $2.85 of speech, split into three jobs of 20,000 characters or fewer, plus $0.01 for one concat job, so about $2.86. Treat that as arithmetic on the catalog price, and confirm the live number with the catalog endpoint before you budget. Retakes are billed like any other job, so fix one section and leave the rest alone.
Spaces and punctuation count toward characters, so a script with long stage directions costs more than the spoken words alone. Strip anything the voice should not read before you submit.
What about several speakers?
VibeVoice advertises up to four speakers in one generation. A Sume job has one voice selector, so a dialogue is one job per speaker turn or per block of turns, then a concat. Use the voice id for each speaker, keep the same model id on every job, and set language for non-English text. Check pacing at each seam by listening; concat is gapless, so add silence deliberately if a pause is wanted.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare two ways to get a voiceover.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en"
}'
Does a 90-minute single pass matter?
The VibeVoice-Realtime-0.5B model is a different tool from the 90-minute one: roughly 300 ms latency and about 10 minutes of generation suit live playback. Sume's TTS is an asynchronous job, so it fits narration that is rendered ahead of time, not a live conversation. Pick by workload, not by the largest number on a README.
For consistency, maybe: one pass can keep a single performance arc. For production, chunking has advantages. You re-render one chapter without paying for the other 89 minutes, and each job has its own status, result and idempotency key. You can also request timestamps.words and segmentation.mode: sentence with a wav container to get gapless sentence slices for captions or lip-sync.
If you need open weights today, check the repository for what is actually available before designing around it. If you need a dependable hosted path, plan for chunks and a concat step.
Sources
Related posts
More in Comparisons
- Vimeo AI dubbing in 29 languages vs Sume: one file per language
Vimeo serves dubbed audio in 29 languages from a single link. Sume returns one finished file per language, so the language switch lives in your player or page.
- Voice cloning consent statement example: Azure, Google, OpenAI
Azure, Google and OpenAI each script the consent a speaker records before a clone. Compare the wording, then write your own release for a Sume voice.
- Voxtral TTS 2-minute native limit vs Sume TTS 1200-second cap
Mistral says Voxtral TTS natively generates up to two minutes and the API handles longer. Sume TTS fails audio over 1200 seconds with tts_duration_exceeded.
- Walmart bans seller logos, Coupang wants yours: one prompt per site
Walmart's image guide bars seller logos; Coupang's main-image page says to show yours clearly. Keep a separate prompt per marketplace in a Sume batch.
Written by Sume