Deepgram Flux TTS: text_spoken and barge-in vs Sume voiceover jobs
Deepgram's Flux TTS reports text_spoken and text_remaining when a caller interrupts. Sume has no barge-in, but word timestamps can mark a cut point.

Deepgram's Flux TTS is built for live voice agents: when a caller interrupts, it reports what was already spoken as text_spoken and what was left as text_remaining. Sume TTS is a different tool, an asynchronous job that returns a finished file, so it has no barge-in. If you want to mark where a listener stopped playback in a finished voiceover, you can request word timestamps and cut the transcript yourself.
AudioXpress reports Flux TTS as generally available on August 13, at $0.045 per 1,000 characters from September 13, with a response time of about 80 ms and English only. Integrations listed include Vapi, Pipecat, LiveKit, Jambonz and Cloudflare.
The comparison
Flux figures come from the AudioXpress report. Sume figures come from the TTS docs and API schema.
| Item | Deepgram Flux TTS | Sume TTS Router |
|---|---|---|
| Delivery | Streaming to a live call | Async job, finished file |
| Interruption | Reports text_spoken and text_remaining | Not supported |
| Languages | English only | Set with a BCP-47 language tag |
| Price | $0.045 per 1,000 characters | About $0.0475 per 1,000 characters (Sonic list x 1.25, rounded up per job) |
| Time to first audio | About 80 ms | Not a latency product |
| Word timing | Not stated in the report | timestamps.words returns start and end seconds |
Mark an interruption point on a finished file
- Request the job with timestamps.words set to true.
- Fetch the result; each word has start and end seconds.
- When your player reports a stop time t, find the last word with end at or before t.
- Treat the words before that as spoken and the rest as remaining text.
import os, requests
r = requests.post(
"https://api.sume.com/v1/tts-router/generate",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "tts-demo-001",
},
json={
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare three prices.",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"timestamps": {"words": True},
},
timeout=30,
)
r.raise_for_status()
print(r.json())Why this is not barge-in
Real barge-in stops generation and changes what the agent does next. Here the audio is already made. You can resume playback or regenerate the rest, but the job itself is not interrupted and a cancelled listener still paid for the full text.
Choosing between them
The deciding question is whether a human is talking to the voice right now. If a caller can interrupt, you need streaming, low latency and a way to know what was heard; that is the job Flux TTS is built for. If a script is written first and the voice is a deliverable, such as a product video, an ad or a lesson, latency matters little and consistency matters a lot. Then an asynchronous job with a pinned model and a stored voice is a better fit.
Many products need both. A support agent may speak live through one vendor while the same company produces its marketing voiceovers through another. Keep the two pipelines separate, log the engine for each, and avoid sharing a voice name between them unless you have tested that the two sound alike.
What Sume does not do
Sume does not stream audio, hold a live call, or react to a caller's voice. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. For a live agent, use a streaming product; use Sume for finished narration, ads and video voice tracks. I could not verify Flux latency claims beyond the AudioXpress report.
Sources
Related posts
More in Comparisons
- Descript Creator 3-person teams vs Sume workspace API keys
Descript lists Creator at $24 for 3-person teams and Business at $50 for 5. Sume API keys are workspace-scoped and spend from one USD balance, no seats.
- Does dictation audio leave your device? Fireflies, Phonon-2, Sume
Fireflies Talk keeps finished dictations local but sends audio to its STT provider. Phonon-2 runs on device. Sume STT is hosted. A sourced privacy comparison.
- Duck a Suno or ElevenLabs track under a voice-over in Sume Timeline
Sume Timeline mixes a soundtrack bed under the voice spine with duck_db from 0 to 20, loop and a fade of up to 10 seconds, at $0.10 per output minute.
- ElevenLabs Free 10,000 credits is 11 minutes of music, no license
ElevenLabs Free gives 10,000 credits a month, about 11 minutes of music, without a commercial license. Sketch there, then ship a track via Sume at $0.125.
Written by Sume