Gradium's sub-50 ms TTS vs Sume's async TTS jobs: which fits

Gradium claims sub-50 ms latency. Sume TTS is a non-streaming job with a 30 second wait window; it suits batch narration and video, not live agent replies.

5 min readSume
All posts

If you need speech to start in tens of milliseconds, Sume TTS is not the right tool: it is an asynchronous job that returns a finished audio file, not a stream. Gradium's October 7 page claims "sub-50ms latency" for its text-to-speech model, which is a different product shape from a Sume job that you submit and read from the result.

This post separates the two shapes so you can place each correctly. Gradium's side comes from its announcement (read 2026-10-10); Sume's from the API reference and Jobs and results.

How does a Sume TTS call actually run?

The Sume schema calls TTS 1.0 an async job with poll or webhook, non-streaming. In mode: async you get a status_url, result_url, events_url and cancel_url back at once. mode: sync and mode: subscribe are aliases for a bounded wait of up to wait_timeout_seconds, with a maximum of 30; if the job is not done, the response carries sync.timed_out or sync.capacity_exhausted and you poll the status URL instead of resubmitting.

The result endpoint answers 409 job_not_completed until result_ready is true. Completed results expose audio artifacts hosted by Sume.

  • Submit with async for anything over a few seconds of audio.
  • Never resubmit on a timed-out wait; poll status_url.
  • Use a signed webhook for terminal delivery only, and keep polling as a backup.

Which workloads suit each shape?

Live voice agents answer in the middle of a conversation, so first-byte time matters and a streaming engine is the right choice. Batch work, such as video voice-over, course narration, ad reads and anything that ends in a rendered file, cares about total cost and repeatability, not the first 50 ms.

Workload fit (Gradium claim read 2026-10-10; Sume behavior from its schema)
WorkloadStreaming engine (Gradium-style)Sume async TTS job
Live phone or chat agentFitsDoes not fit: no streaming
Narration for a rendered videoWorks, no need for the speedFits
Ad reads in bulkWorksFits; one job per script
Captions from word timingsDepends on the vendorFits: timestamps.words on the job
Sentence-sliced wav for lip-syncDepends on the vendorFits: segmentation with wav output

What do you get in exchange for waiting?

The job can return word timings and sentence segments in the same call. With timestamps.words: true the result carries monotonic words[] with start and end seconds. Adding segmentation: {mode: "sentence"} returns gapless segments[], and with a wav or raw container each segment includes a sample-exact audio_url. With mp3 you get timings but no per-segment audio file.

Those outputs feed timeline audio and caption jobs directly, which is where batch pipelines spend their effort. A latency-first streaming engine usually gives you the sound only.

What is a sensible split of responsibilities?

Use a streaming TTS for the live conversation, if you have one. Use Sume for everything that becomes a file: render the narration with a Sonic engine, let the job return timings, and hand the audio to captions, timeline or a lip-sync clip. The two meet at the audio file, so neither side needs to know about the other.

Sources

Related posts

More in Models

All Models posts

Written by Sume