Realtime voice model or async TTS job for narration?

GPT-Live-1 and Think Fast 2.0 are built for conversation. For narration, an async TTS job is simpler. How to choose, with the facts from vendor pages.

4 min readSume
All posts

Use an async TTS job for narration and a realtime model for conversation. Narration has a script, no one is waiting mid-sentence, and you want a file you can re-render, store and mix. A realtime model is built to listen, think and speak in a live session, which is more than a voiceover needs.

What the vendors say

OpenAI's model page for gpt-live-1 lists the Live endpoint v1/live/sessions, and its pricing page bills GPT-Live sessions per second, so it is a session API rather than a file-in, file-out speech call. Its price is $0.05 per minute, billed per second. xAI's Think Fast 2.0 post lists a time to first audio of 0.70 seconds and an evaluation across 24 languages, at $0.08 per minute for speech to speech.

Facts from vendor pages named in Sources and the Sume pricing page, read 2026-10-01.
QuestionRealtime sessionSume async TTS job
InterfaceLive sessionPOST a job, fetch the result
Price shapePer minute of sessionPer 1,000 characters ($0.0475)
OutputStreamed audioA file you can store
Re-render a lineRun it again liveRun the job again; each run is billed
Mix with musicNot part of the sessionTimeline 1.0 render

A rule of thumb

  • Script known in advance: async TTS.
  • Listener can interrupt: realtime.
  • Needs captions or timings: async, then STT with word times.
  • Needs consistent output across edits: pin the model id, as in pinning a TTS model.

Cost sanity check

A 10-minute narration at 150 words per minute is about 9,000 characters (an assumption: 6 characters per word including spaces). On Sume that is 9 x $0.0475 = $0.4275. The same ten minutes in a realtime session at $0.05 per minute is $0.50, but you would pay again for every retake and could not mix a bed in the session.

Limits

Realtime sessions can respond to what the listener says, and Sume cannot. If that is the product, see the cost view in realtime voice cost per hour and use a live vendor. If it is a voiceover, the Sume audio endpoint map lists the calls to chain.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume