OpenAI-compatible /v1/audio/speech: what Sume's TTS route is

Sume has no `/v1/audio/speech` clone. Its TTS is `POST /v1/tts-1.0/generate` or the Router, a job-based request authenticated with a Sume key.

4 min readSume
All posts

Sume does not expose an OpenAI-format /v1/audio/speech route, so you cannot point an OpenAI SDK at it. Its text to speech is its own request: POST /v1/tts-1.0/generate, or POST /v1/tts-router/generate when you want to name a catalog model. Both create a job, and you authenticate with your Sume API key only.

What did Inworld ship?

Inworld's TTS release notes dated September 11, 2026 (read 2026-09-30) say its realtime TTS now serves POST /v1/audio/speech in OpenAI's format, so apps built on OpenAI's text-to-speech can switch by pointing the SDK at Inworld and choosing an Inworld model and voice. The notes say the official Python and Node.js SDKs work unchanged, including streaming responses.

How does Sume's request differ?

The shapes are different enough that a drop-in swap will not work.

Request shape, Inworld release notes against the Sume API reference, read 2026-09-30.
AspectInworld (per notes)Sume
RoutePOST /v1/audio/speech, OpenAI formatPOST /v1/tts-1.0/generate or POST /v1/tts-router/generate
ModelAn Inworld modelTTS 1.0 has no engine picker; the Router requires model, such as sonic-3.6
VoiceAn Inworld voiceavatar_id, avatar_handle or voice.id
AuthOpenAI SDK pointed at InworldSume API key only; no provider credentials
ResultSpeech responseA job with a status URL and audio

What does a Sume call look like?

Send transcript plus a voice selector. On TTS 1.0, model and model_id are rejected with 400. On the Router, model is required and comes from GET /v1/tts-router/models; an unknown id fails with 400 model_not_found.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: speech-demo-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Thanks for calling.",
    "avatar_handle": "acme",
    "language": "en"
  }'

Is it a streaming response?

No. The route description calls it an async job with poll or webhook delivery, non-streaming. A sync wait is clamped to 0 to 30 seconds, and the schema says that ceiling bounds the HTTP wait, not the job. If the job is not finished, the response is still a 2xx with the current state and polling URLs; keep polling and do not submit a second paid job for the same intent. Audio longer than 1200 seconds fails with tts_duration_exceeded. For latency-sensitive use, read text to speech streaming API.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume