Realtime voice model or async TTS job for narration?
GPT-Live-1 and Think Fast 2.0 are built for conversation. For narration, an async TTS job is simpler. How to choose, with the facts from vendor pages.

Use an async TTS job for narration and a realtime model for conversation. Narration has a script, no one is waiting mid-sentence, and you want a file you can re-render, store and mix. A realtime model is built to listen, think and speak in a live session, which is more than a voiceover needs.
What the vendors say
OpenAI's model page for gpt-live-1 lists the Live endpoint v1/live/sessions, and its pricing page bills GPT-Live sessions per second, so it is a session API rather than a file-in, file-out speech call. Its price is $0.05 per minute, billed per second. xAI's Think Fast 2.0 post lists a time to first audio of 0.70 seconds and an evaluation across 24 languages, at $0.08 per minute for speech to speech.
| Question | Realtime session | Sume async TTS job |
|---|---|---|
| Interface | Live session | POST a job, fetch the result |
| Price shape | Per minute of session | Per 1,000 characters ($0.0475) |
| Output | Streamed audio | A file you can store |
| Re-render a line | Run it again live | Run the job again; each run is billed |
| Mix with music | Not part of the session | Timeline 1.0 render |
A rule of thumb
- Script known in advance: async TTS.
- Listener can interrupt: realtime.
- Needs captions or timings: async, then STT with word times.
- Needs consistent output across edits: pin the model id, as in pinning a TTS model.
Cost sanity check
A 10-minute narration at 150 words per minute is about 9,000 characters (an assumption: 6 characters per word including spaces). On Sume that is 9 x $0.0475 = $0.4275. The same ten minutes in a realtime session at $0.05 per minute is $0.50, but you would pay again for every retake and could not mix a bed in the session.
Limits
Realtime sessions can respond to what the listener says, and Sume cannot. If that is the product, see the cost view in realtime voice cost per hour and use a live vendor. If it is a voiceover, the Sume audio endpoint map lists the calls to chain.
Sources
Related posts
More in Comparisons
- Replicate output files deleted after an hour: what Sume's URLs do
On Replicate, files from API predictions are deleted after an hour. Sume's Format run docs call media URLs durable, but check expires_at when a URL is signed.
- Replicate Cancel-After header: 5 s to 24 h. Is there a Sume deadline?
Replicate's Cancel-After header sets a prediction deadline from 5 seconds to 24 hours. I found no deadline field in Sume's OpenAPI, so cancel it yourself.
- Replicate Prefer: wait holds 60 s by default; Sume's sync cap is 30 s
Replicate's Prefer: wait header holds the request up to 60 seconds by default; Sume's sync mode waits at most 30 seconds, then returns a job to poll.
- Replicate private models bill idle time: how Sume bills a job
On Replicate, a private model on dedicated hardware bills setup, idle and active time. Sume bills per job: a reserve at submit, then capture or refund.
Written by Sume