Inworld's OpenAI-style /v1/audio/speech vs a Sume TTS job in Python
Inworld added POST /v1/audio/speech on Sept 11, 2026, so OpenAI SDKs work unchanged. The Sume equivalent is a job you poll: a Python sample that runs.

On September 11, 2026, Inworld's realtime TTS gained an OpenAI-format endpoint, POST /v1/audio/speech, so the official OpenAI Python and Node SDKs work against it unchanged, streaming included. Sume has no such endpoint. Its text-to-speech is a job: you submit a request, poll a status URL, and read a hosted audio artifact, which takes about twenty lines of Python.
Inworld facts below are from its TTS release notes, read 2026-10-03. Sume facts are from the OpenAPI reference and the jobs docs.
What did Inworld ship on September 11?
The release note says Inworld's Realtime TTS now supports OpenAI's API format at POST /v1/audio/speech. The official OpenAI Python and Node.js SDKs work unchanged, including streaming responses. Output formats listed are MP3, Opus, FLAC, WAV and PCM, plus server-sent events, and a unified base URL, https://api.inworld.ai/v1, covers both LLM and TTS routing.
For a team already on the OpenAI SDK for speech, the migration is a base URL and a key. The note does not list prices, voice names or limits, so check Inworld's own pricing and voice pages before you commit.
What does the same call look like on Sume?
Sume's route is POST /v1/tts-1.0/generate, or POST /v1/tts-router/generate when you want to name an engine such as sonic-3.6. The body takes transcript (1 to 20,000 characters), one voice selector (avatar_handle, avatar_id or voice.id), and optional language, output_format, timestamps, segmentation and webhook fields. Authentication is a Sume key only. There is no model, voice or input field shaped like OpenAI's, and no raw audio byte stream in the response.
The response is a job envelope. You then read GET /v1/jobs/{id}/status until terminal is true, and read GET /v1/jobs/{id}/result for result.artifacts[], where the audio artifact has type: audio and a media.sume.com URL.
| Item | Inworld /v1/audio/speech | Sume tts-1.0 generate |
|---|---|---|
| Client | OpenAI SDKs, base URL changed | Plain HTTP with a Sume key |
| Response | Audio bytes or SSE stream | Job envelope, then a hosted artifact |
| Formats | MP3, Opus, FLAC, WAV, PCM | Set with output_format; see the OpenAPI reference |
| Voice | As in the OpenAI request shape | avatar_handle, avatar_id or voice.id |
| Text limit | Not stated in the note | 20,000 characters per request |
Can you do it in Python?
This sample submits async with an idempotency key, polls the status, and downloads the first audio artifact. Set SUME_API_KEY, and replace the handle with one of your own avatars. Reusing the same Idempotency-Key on a retry returns the original job rather than billing twice.
import os, time, requests
BASE = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {
"model": "sonic-3.6",
"transcript": "Your order shipped this morning.",
"avatar_handle": "@speaker",
"mode": "async",
}
r = requests.post(f"{BASE}/v1/tts-router/generate", json=body,
headers={**H, "Idempotency-Key": "speech-demo-001"})
r.raise_for_status()
job_id = r.json()["id"]
while True:
s = requests.get(f"{BASE}/v1/jobs/{job_id}/status", headers=H).json()
if s["terminal"]:
break
time.sleep(s.get("next_poll_after_seconds") or 2)
if s["sume_status"] != "completed":
raise SystemExit(f"job ended as {s['sume_status']}")
res = requests.get(f"{BASE}/v1/jobs/{job_id}/result", headers=H).json()
audio = next(a for a in res["result"]["artifacts"] if a["type"] == "audio")
open("speech.mp3", "wb").write(requests.get(audio["url"]).content)
print(audio["url"])What do you give up, and what do you gain?
You give up the drop-in SDK path and any low-latency byte stream: the Sume response is a job envelope and an artifact, with no audio stream, and the job model adds a poll step. You gain a durable record. Each request is a job with an id, a status history, an optional signed webhook, and a hosted artifact you can fetch again later, which suits narration, dubbing and batch work more than live conversation.
A rough price line: at $0.0475 per 1,000 characters, a 600-character line costs about three cents. Run a small test before a large batch, and see the posts below for sync waits and retries.
- Live, interruptible speech: use a streaming endpoint such as Inworld's.
- Narration, dubbing, batch scripts: Sume jobs fit.
- Mixed: stream for conversation, then regenerate final takes as Sume jobs for delivery.
Sources
Related posts
More in Developers
- Is AI avatar video real time? How long a Sume job takes
A Sume avatar video is a job, not a live stream: it queues, renders, and you poll or take a webhook. What the sync wait caps at, and a Python polling loop.
- GET /v1/jobs thread_id filter: why a teammate's job still 404s
The thread_id filter on the Sume jobs list narrows results and never widens what an API key can read. Teammates' jobs stay 404, and only the creator can cancel.
- A kill switch for paid Sume submits: stop new jobs, cancel queued
Add an off switch to code that spends on the Sume API: check a flag before each submit, then cancel queued jobs; a 409 job_generation_already_started will bill.
- Kling 3 input_references 400 unsupported_capability: fix
Sending input_references to kling-3 on Sume returns 400 unsupported_capability. Move the image to frame_images or pick a model that takes references.
Written by Sume