gpt-realtime-2.1 only supports v1/realtime: Sume TTS is a job
gpt-realtime-2.1 works only on v1/realtime, not Chat Completions or Batch. Sume TTS 1.0 is an async job you poll, which suits narration, not live talk.

OpenAI's model page lists v1/realtime as the only supported endpoint for gpt-realtime-2.1; Chat Completions, Batch and Fine-tuning are marked not supported. Sume's TTS 1.0 route works the other way round: POST /v1/tts-1.0/generate creates an async job that you poll or receive by webhook, and it does not stream.
Facts are from the OpenAI model page and Sume's OpenAPI file and docs, read 2026-10-01.
What does v1/realtime only mean?
You cannot send gpt-realtime-2.1 a one-shot request on the usual text endpoints or queue it as a batch. Every use is a realtime session, so your client holds a live connection for as long as the audio flows.
How does a Sume TTS request differ?
The OpenAPI description of /v1/tts-1.0/generate says phase 1 is "async job + poll/webhook (non-streaming)". With mode: async the response returns status_url, result_url, events_url and cancel_url. Completed results expose mirrored audio artifacts. Related audio routes such as Timeline audio use the same envelope: GET /v1/jobs/:id/status, then GET /v1/jobs/:id/result.
| gpt-realtime-2.1 | Sume TTS 1.0 | |
|---|---|---|
| Endpoint | v1/realtime only | POST /v1/tts-1.0/generate |
| Batch | Not supported | One job per request |
| Delivery | Live session | Poll status_url or webhook |
| Input cap | 128,000 context window | 20000 characters of transcript |
| Output | Streamed audio | Audio file artifact |
What audio format comes back?
A file. The request has an output_format object with a container of mp3, wav or raw and a sample_rate in Hz, defaulting to mp3 at 44100. Sample rates in the schema include 8000, 16000, 22050, 24000, 44100 and 48000. See TTS for video narration for using the file in a timeline.
Which one fits which job?
A live conversation where the user interrupts needs a realtime session, and Sume's TTS 1.0 route is not one: it is non-streaming. Pre-rendered narration, voiceover for a video or a phone-line prompt fits a job: submit the transcript, get the job id, poll, and download the artifact. The transcript is capped at 20000 characters, so split long scripts into several jobs.
Sources
Related posts
More in Developers
- gpt-transcribe expected languages vs Sume STT language_code
OpenAI's gpt-transcribe accepts multiple expected input languages. Sume STT takes one optional language_code hint or auto-detect. What to send for mixed audio.
- Grok Imagine <IMAGE_0> tags shift with a first frame; Sume is 0-based
On xAI, pinning a first frame with image takes <IMAGE_0> and references start at <IMAGE_1>. Sume Omni tags are <IMAGE_REF_0>, 0-based by list order.
- Grok Imagine base64 response_format vs a Sume hosted URL
xAI lets you request base64 instead of a temporary URL. Sume returns Sume-hosted, signed URLs in data[].url, not inline base64, for x-ai/grok-image.
- Grok Imagine video extension duration counts only the extension
In xAI video extension, duration is the added seconds only: a 10s input plus duration 5 returns 15s. Sume has no extend on Omni; chain clips from a last frame.
Written by Sume