Gradium's sub-50 ms TTS vs Sume's async TTS jobs: which fits
Gradium claims sub-50 ms latency. Sume TTS is a non-streaming job with a 30 second wait window; it suits batch narration and video, not live agent replies.

If you need speech to start in tens of milliseconds, Sume TTS is not the right tool: it is an asynchronous job that returns a finished audio file, not a stream. Gradium's October 7 page claims "sub-50ms latency" for its text-to-speech model, which is a different product shape from a Sume job that you submit and read from the result.
This post separates the two shapes so you can place each correctly. Gradium's side comes from its announcement (read 2026-10-10); Sume's from the API reference and Jobs and results.
How does a Sume TTS call actually run?
The Sume schema calls TTS 1.0 an async job with poll or webhook, non-streaming. In mode: async you get a status_url, result_url, events_url and cancel_url back at once. mode: sync and mode: subscribe are aliases for a bounded wait of up to wait_timeout_seconds, with a maximum of 30; if the job is not done, the response carries sync.timed_out or sync.capacity_exhausted and you poll the status URL instead of resubmitting.
The result endpoint answers 409 job_not_completed until result_ready is true. Completed results expose audio artifacts hosted by Sume.
- Submit with
asyncfor anything over a few seconds of audio. - Never resubmit on a timed-out wait; poll
status_url. - Use a signed webhook for terminal delivery only, and keep polling as a backup.
Which workloads suit each shape?
Live voice agents answer in the middle of a conversation, so first-byte time matters and a streaming engine is the right choice. Batch work, such as video voice-over, course narration, ad reads and anything that ends in a rendered file, cares about total cost and repeatability, not the first 50 ms.
| Workload | Streaming engine (Gradium-style) | Sume async TTS job |
|---|---|---|
| Live phone or chat agent | Fits | Does not fit: no streaming |
| Narration for a rendered video | Works, no need for the speed | Fits |
| Ad reads in bulk | Works | Fits; one job per script |
| Captions from word timings | Depends on the vendor | Fits: timestamps.words on the job |
| Sentence-sliced wav for lip-sync | Depends on the vendor | Fits: segmentation with wav output |
What do you get in exchange for waiting?
The job can return word timings and sentence segments in the same call. With timestamps.words: true the result carries monotonic words[] with start and end seconds. Adding segmentation: {mode: "sentence"} returns gapless segments[], and with a wav or raw container each segment includes a sample-exact audio_url. With mp3 you get timings but no per-segment audio file.
Those outputs feed timeline audio and caption jobs directly, which is where batch pipelines spend their effort. A latency-first streaming engine usually gives you the sound only.
What is a sensible split of responsibilities?
Use a streaming TTS for the live conversation, if you have one. Use Sume for everything that becomes a file: render the narration with a Sonic engine, let the job return timings, and hand the audio to captions, timeline or a lip-sync clip. The two meet at the audio file, so neither side needs to know about the other.
Sources
Related posts
More in Models
- Grok Imagine on Sume lists 9:19.5 and 9:20 but not 4:5
Sume's Grok Imagine row lists tall phone ratios 9:19.5 and 9:20 and wide 20:9 and 19.5:9, but not 4:5, 5:4 or 21:9, and caps n at 1. Here is what to do.
- Grok Imagine multi-image edit: xAI says 5 sources, Sume lists 10
xAI's Imagine docs cap multi-image edits at 5 source images and 10 outputs per request. Sume's Grok Imagine row lists 10 references and one output. Port safely.
- Higgsfield Soul on Sume: n must be 1 or 4, n=2 and n=3 return 400
Sume's Higgsfield Soul row accepts n of 1 or 4 only, rejects image_size, and offers 720p or 1080p at $0.005 or $0.0075 per image. Here is how to batch it.
- Can Imagen 4 use my product photo? No: pick an edit-capable row
Imagen 4 Fast and Ultra on Sume are text-to-image only and reject references. Google also says Imagen is shut down in its API. Six rows take a product photo.
Written by Sume