Inworld's OpenAI-style /v1/audio/speech vs a Sume TTS job in Python

Inworld added POST /v1/audio/speech on Sept 11, 2026, so OpenAI SDKs work unchanged. The Sume equivalent is a job you poll: a Python sample that runs.

5 min readSume
All posts

On September 11, 2026, Inworld's realtime TTS gained an OpenAI-format endpoint, POST /v1/audio/speech, so the official OpenAI Python and Node SDKs work against it unchanged, streaming included. Sume has no such endpoint. Its text-to-speech is a job: you submit a request, poll a status URL, and read a hosted audio artifact, which takes about twenty lines of Python.

Inworld facts below are from its TTS release notes, read 2026-10-03. Sume facts are from the OpenAPI reference and the jobs docs.

What did Inworld ship on September 11?

The release note says Inworld's Realtime TTS now supports OpenAI's API format at POST /v1/audio/speech. The official OpenAI Python and Node.js SDKs work unchanged, including streaming responses. Output formats listed are MP3, Opus, FLAC, WAV and PCM, plus server-sent events, and a unified base URL, https://api.inworld.ai/v1, covers both LLM and TTS routing.

For a team already on the OpenAI SDK for speech, the migration is a base URL and a key. The note does not list prices, voice names or limits, so check Inworld's own pricing and voice pages before you commit.

What does the same call look like on Sume?

Sume's route is POST /v1/tts-1.0/generate, or POST /v1/tts-router/generate when you want to name an engine such as sonic-3.6. The body takes transcript (1 to 20,000 characters), one voice selector (avatar_handle, avatar_id or voice.id), and optional language, output_format, timestamps, segmentation and webhook fields. Authentication is a Sume key only. There is no model, voice or input field shaped like OpenAI's, and no raw audio byte stream in the response.

The response is a job envelope. You then read GET /v1/jobs/{id}/status until terminal is true, and read GET /v1/jobs/{id}/result for result.artifacts[], where the audio artifact has type: audio and a media.sume.com URL.

Inworld OpenAI-format speech vs a Sume TTS job (read 2026-10-03)
ItemInworld /v1/audio/speechSume tts-1.0 generate
ClientOpenAI SDKs, base URL changedPlain HTTP with a Sume key
ResponseAudio bytes or SSE streamJob envelope, then a hosted artifact
FormatsMP3, Opus, FLAC, WAV, PCMSet with output_format; see the OpenAPI reference
VoiceAs in the OpenAI request shapeavatar_handle, avatar_id or voice.id
Text limitNot stated in the note20,000 characters per request

Can you do it in Python?

This sample submits async with an idempotency key, polls the status, and downloads the first audio artifact. Set SUME_API_KEY, and replace the handle with one of your own avatars. Reusing the same Idempotency-Key on a retry returns the original job rather than billing twice.

import os, time, requests

BASE = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

body = {
    "model": "sonic-3.6",
    "transcript": "Your order shipped this morning.",
    "avatar_handle": "@speaker",
    "mode": "async",
}
r = requests.post(f"{BASE}/v1/tts-router/generate", json=body,
                  headers={**H, "Idempotency-Key": "speech-demo-001"})
r.raise_for_status()
job_id = r.json()["id"]

while True:
    s = requests.get(f"{BASE}/v1/jobs/{job_id}/status", headers=H).json()
    if s["terminal"]:
        break
    time.sleep(s.get("next_poll_after_seconds") or 2)

if s["sume_status"] != "completed":
    raise SystemExit(f"job ended as {s['sume_status']}")
res = requests.get(f"{BASE}/v1/jobs/{job_id}/result", headers=H).json()
audio = next(a for a in res["result"]["artifacts"] if a["type"] == "audio")
open("speech.mp3", "wb").write(requests.get(audio["url"]).content)
print(audio["url"])

What do you give up, and what do you gain?

You give up the drop-in SDK path and any low-latency byte stream: the Sume response is a job envelope and an artifact, with no audio stream, and the job model adds a poll step. You gain a durable record. Each request is a job with an id, a status history, an optional signed webhook, and a hosted artifact you can fetch again later, which suits narration, dubbing and batch work more than live conversation.

A rough price line: at $0.0475 per 1,000 characters, a 600-character line costs about three cents. Run a small test before a large batch, and see the posts below for sync waits and retries.

  • Live, interruptible speech: use a streaming endpoint such as Inworld's.
  • Narration, dubbing, batch scripts: Sume jobs fit.
  • Mixed: stream for conversation, then regenerate final takes as Sume jobs for delivery.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume