Deepgram Flux TTS: text_spoken and barge-in vs Sume voiceover jobs

Deepgram's Flux TTS reports text_spoken and text_remaining when a caller interrupts. Sume has no barge-in, but word timestamps can mark a cut point.

4 min readSume
All posts

Deepgram's Flux TTS is built for live voice agents: when a caller interrupts, it reports what was already spoken as text_spoken and what was left as text_remaining. Sume TTS is a different tool, an asynchronous job that returns a finished file, so it has no barge-in. If you want to mark where a listener stopped playback in a finished voiceover, you can request word timestamps and cut the transcript yourself.

AudioXpress reports Flux TTS as generally available on August 13, at $0.045 per 1,000 characters from September 13, with a response time of about 80 ms and English only. Integrations listed include Vapi, Pipecat, LiveKit, Jambonz and Cloudflare.

The comparison

Flux figures come from the AudioXpress report. Sume figures come from the TTS docs and API schema.

Live agent TTS versus a voiceover job (read 2026-10-08)
ItemDeepgram Flux TTSSume TTS Router
DeliveryStreaming to a live callAsync job, finished file
InterruptionReports text_spoken and text_remainingNot supported
LanguagesEnglish onlySet with a BCP-47 language tag
Price$0.045 per 1,000 charactersAbout $0.0475 per 1,000 characters (Sonic list x 1.25, rounded up per job)
Time to first audioAbout 80 msNot a latency product
Word timingNot stated in the reporttimestamps.words returns start and end seconds

Mark an interruption point on a finished file

  • Request the job with timestamps.words set to true.
  • Fetch the result; each word has start and end seconds.
  • When your player reports a stop time t, find the last word with end at or before t.
  • Treat the words before that as spoken and the rest as remaining text.
import os, requests

r = requests.post(
    "https://api.sume.com/v1/tts-router/generate",
    headers={
        "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Idempotency-Key": "tts-demo-001",
    },
    json={
        "model": "sonic-3.6",
        "transcript": "Welcome back. Today we compare three prices.",
        "voice": {"id": os.environ["SUME_VOICE_ID"]},
        "timestamps": {"words": True},
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

Why this is not barge-in

Real barge-in stops generation and changes what the agent does next. Here the audio is already made. You can resume playback or regenerate the rest, but the job itself is not interrupted and a cancelled listener still paid for the full text.

Choosing between them

The deciding question is whether a human is talking to the voice right now. If a caller can interrupt, you need streaming, low latency and a way to know what was heard; that is the job Flux TTS is built for. If a script is written first and the voice is a deliverable, such as a product video, an ad or a lesson, latency matters little and consistency matters a lot. Then an asynchronous job with a pinned model and a stored voice is a better fit.

Many products need both. A support agent may speak live through one vendor while the same company produces its marketing voiceovers through another. Keep the two pipelines separate, log the engine for each, and avoid sharing a voice name between them unless you have tested that the two sound alike.

What Sume does not do

Sume does not stream audio, hold a live call, or react to a caller's voice. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. For a live agent, use a streaming product; use Sume for finished narration, ads and video voice tracks. I could not verify Flux latency claims beyond the AudioXpress report.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume