OpenRouter audio/speech returns bytes; Sume TTS returns a job
OpenRouter's /audio/speech streams raw audio bytes. Sume's TTS Router returns a job id and a media URL, so a byte-stream client needs a poll step.

OpenRouter's /api/v1/audio/speech answers with a raw audio byte stream, not JSON. Sume's explicit TTS route, POST /v1/tts-router/generate, answers with a job: you get a job id first, then read a Sume media URL from the result. Porting means replacing "write the response body to a file" with "poll, then download the URL".
OpenRouter facts are from its text-to-speech guide; Sume facts from the OpenAPI file and Jobs and results, read 2026-10-01.
What does OpenRouter return from audio/speech?
The guide describes an endpoint compatible with the OpenAI Audio Speech API: send text, receive a raw audio byte stream in your chosen format. It says the response is not JSON, so you can pipe it to a file or an audio player. Its examples send model, input, voice and response_format.
What does the Sume TTS Router return instead?
The OpenAPI description calls it "Explicit pass-through text-to-speech" and requires a model taken from GET /v1/tts-router/models. Every submit mode returns the job id in its first response, and a 2xx means the job exists, not that audio is ready. The docs tell clients to poll status_url until terminal is true, then read result_url once result_ready is true, and to use the Sume media URLs from the result.
What changes when I port a byte-stream client?
| Step | OpenRouter audio/speech | Sume TTS Router |
|---|---|---|
| First response | Raw audio bytes | Job envelope with the job id |
| Get the audio | Write the body to a file | Poll, then fetch the URL in the result |
| Model choice | model in the body | model from GET /v1/tts-router/models |
| Text field | input | Transcript and voice selector per the OpenAPI schema |
Should I resubmit if my wait times out?
No. In sync and subscribe modes the wait is bounded at 30 seconds; if the job is not terminal, poll rather than resubmit. See also the OpenAI-compatible speech endpoint note and async TTS jobs.
Sources
Related posts
More in Developers
- OpenRouter base64 input_audio vs Sume STT 1.0 audio_url
OpenRouter transcription takes base64 audio inside the JSON body. Sume STT 1.0 takes an audio_url instead, so you host the file and send a link.
- OpenRouter Batch API vs Sume async jobs: which for video?
OpenRouter's Batch API is for text and embeddings with a 24 hour window. For video files, Sume returns a job id per request: async, sync, or webhook mode.
- OpenRouter batch custom_id vs one Sume Idempotency-Key per job
OpenRouter's Batch API needs a custom_id unique within each batch. Sume has no batch envelope for these submits: give each job its own Idempotency-Key.
- OpenRouter batch 24h window and expired vs Sume queued jobs
OpenRouter batches use a 24-hour completion window and can end as expired. Sume jobs have no queue expiry option today: a queued job waits, or you cancel it.
Written by Sume