VibeVoice 90-minute TTS vs Sume's 1,200-second jobs: long audio

VibeVoice-TTS-1.5B targets 90 minutes and 4 speakers, and Microsoft removed its TTS code. What that means for long narration, and how Sume chunks it in jobs.

5 min readSume
All posts

VibeVoice-TTS-1.5B is described as synthesizing up to 90 minutes of conversational speech with up to four speakers in one pass, but Microsoft has removed the TTS code after misuse, so you cannot count on running it. Sume takes the opposite approach to long audio: each job is capped at 20,000 characters and 1,200 seconds, and you assemble a long recording from several jobs.

VibeVoice details come from the Microsoft repository, read on 2026-10-02. Sume details come from the Sume API reference and the timeline audio guide.

What does the VibeVoice repository say?

It describes a family of open-source voice models under the MIT license. VibeVoice-TTS-1.5B synthesizes conversational speech up to 90 minutes with up to 4 distinct speakers. VibeVoice-Realtime-0.5B streams with roughly 300 ms latency and supports about 10 minutes of generation. VibeVoice-ASR-7B handles 60 minutes of audio in a single pass, with speaker identification, timestamps and hotword support across 50 or more languages.

The repository also states that Microsoft removed the VibeVoice-TTS code after discovering misuse inconsistent with the research intent, warns about deepfakes and disinformation, and recommends disclosing AI-generated content. Treat the 90-minute number as a research capability, not a service you can call.

How does Sume handle a long script?

A single Sume TTS request takes 1 to 20,000 characters, and a synthesized result longer than 1,200 seconds fails with tts_duration_exceeded without capturing credit. So a 90-minute recording is several jobs, split at natural breaks such as chapters or sections. The catalog price is $0.0475 per 1,000 characters, and spaces and punctuation count.

Join the finished takes into one file with timeline audio: operation: concat accepts 1 to 20 parts, joins them in the sample domain with no gap at the seams, and costs a flat $0.01 per job. Every part must already be Sume-hosted audio and share one channel layout, and the produced file is capped at 1,800 seconds.

Long audio on Sume, read 2026-10-02
LimitValueWhat to do
Characters per TTS request20,000Split the script at paragraph or chapter breaks
Audio per TTS job1,200 secondsKeep each chunk well under 20 minutes of speech
Parts per concat job20Concat in two stages for more than 20 takes
Joined file length1,800 secondsJoined files longer than 30 minutes need another plan

What would a long recording cost?

Take a 60,000-character script. At $0.0475 per 1,000 characters that is 60 x $0.0475 = $2.85 of speech, split into three jobs of 20,000 characters or fewer, plus $0.01 for one concat job, so about $2.86. Treat that as arithmetic on the catalog price, and confirm the live number with the catalog endpoint before you budget. Retakes are billed like any other job, so fix one section and leave the rest alone.

Spaces and punctuation count toward characters, so a script with long stage directions costs more than the spoken words alone. Strip anything the voice should not read before you submit.

What about several speakers?

VibeVoice advertises up to four speakers in one generation. A Sume job has one voice selector, so a dialogue is one job per speaker turn or per block of turns, then a concat. Use the voice id for each speaker, keep the same model id on every job, and set language for non-English text. Check pacing at each seam by listening; concat is gapless, so add silence deliberately if a pause is wanted.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: narration-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Welcome back. Today we compare two ways to get a voiceover.",
    "voice": { "id": "'"$VOICE_ID"'" },
    "language": "en"
  }'

Does a 90-minute single pass matter?

The VibeVoice-Realtime-0.5B model is a different tool from the 90-minute one: roughly 300 ms latency and about 10 minutes of generation suit live playback. Sume's TTS is an asynchronous job, so it fits narration that is rendered ahead of time, not a live conversation. Pick by workload, not by the largest number on a README.

For consistency, maybe: one pass can keep a single performance arc. For production, chunking has advantages. You re-render one chapter without paying for the other 89 minutes, and each job has its own status, result and idempotency key. You can also request timestamps.words and segmentation.mode: sentence with a wav container to get gapless sentence slices for captions or lip-sync.

If you need open weights today, check the repository for what is actually available before designing around it. If you need a dependable hosted path, plan for chunks and a concat step.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume