OpenAI TTS has Opus, AAC and FLAC; Sume TTS has MP3, WAV and raw
OpenAI's speech API lists six output formats. Sume TTS 1.0 returns MP3, WAV or raw PCM, so Opus, AAC and FLAC need a conversion step after generation.

OpenAI's speech endpoint lists MP3, Opus, AAC, FLAC, WAV and PCM as output formats, while Sume's TTS 1.0 route returns an mp3, wav or raw container. If your player needs Opus, AAC or FLAC, generate WAV on Sume and convert it with your own tool afterwards, because Sume does not emit those three.
OpenAI's list is from its text-to-speech guide (read 2026-10-10). The Sume side is from the Sume OpenAPI contract for POST /v1/tts-1.0/generate.
Side by side
The two lists overlap on MP3, WAV and raw samples, and differ on the compressed formats. The OpenAI guide describes PCM as raw 24 kHz 16-bit signed little-endian samples without a header, and recommends wav or pcm for the fastest response. Sume's raw container is paired with an explicit encoding and sample rate that you choose.
| Format | OpenAI speech | Sume TTS 1.0 |
|---|---|---|
| MP3 | Yes (default) | Yes, bit rate 32000 to 192000 |
| WAV | Yes | Yes |
| Raw PCM | Yes, 24 kHz 16-bit | Yes, raw container with pcm_s16le, pcm_f32le, pcm_mulaw or pcm_alaw |
| Opus | Yes | No |
| AAC | Yes | No |
| FLAC | Yes | No |
What Sume lets you tune instead
Sume's output_format block gives you control that the OpenAI list does not describe: sample rates of 8000, 16000, 22050, 24000, 44100 or 48000 Hz, MP3 bit rates from 32,000 to 192,000, and the four raw encodings. For telephony you can pick 8 kHz mu-law or A-law raw. For editing you can pick 48 kHz WAV. That covers the cases most Opus and FLAC requests come from, although it will not give you an Opus stream for WebRTC.
- Streaming to a browser: MP3 or WAV; Sume jobs are request and result, not chunked audio.
- Archiving: WAV, then compress to FLAC on your side.
- Mobile app bundles that want AAC: convert from WAV with your own encoder.
- Phone systems: raw
pcm_mulaworpcm_alawat 8000 Hz.
Other differences that matter when you switch
OpenAI lists 13 built-in voices by name. Sume does not use named stock voices on this route; the voice.id field is a UUID or a voi_ id, and anything else returns 400 invalid_voice_id. Input is capped at 20,000 characters per job on Sume, with a 1,200 second ceiling on the generated audio (tts_duration_exceeded). The OpenAI guide did not state an input limit on the page I read, so I make no comparison there.
Pricing is a flat per-character rate on Sume: $0.0475 per 1,000 characters, with a one-cent minimum per job and a maximum of 95 cents. I did not read an OpenAI price on the page cited above, so there is no price comparison in this post.
A practical rule
Decide your delivery format first. If it is MP3, WAV or telephony raw, Sume covers it in one call. If it is Opus, AAC or FLAC, treat Sume as the generation step and budget a local conversion; do not expect the job to return a file in a format its schema does not list.
Sources
Related posts
More in Comparisons
- Poster headline in the image: Ideogram 4.5 or Nano Banana 2.1 on Sume?
Both vendors pitch readable text in images. On Sume, Ideogram 4.5 starts at $0.0375 and Nano Banana 2.1 at $0.075. Compare price, ratios and reference limits.
- Qwen Image Max vs Qwen Image on Sume: 3.75x the price, no references
On Sume, Qwen Image Max bills $0.09375 per image and is text-only, while Qwen Image bills $0.025 and takes up to 10 references. Both list 13 aspect ratios.
- Recraft V4 vs FLUX.2 Pro on Sume: $0.05 text-only vs $0.0375 with refs
On Sume, Recraft V4 bills $0.05 per image, is text-only and returns webp only. FLUX.2 Pro bills $0.0375 and takes 10 references. Both list 13 ratios.
- Remotion Automator $0.01 a render vs Sume Timeline $0.10 a minute
Remotion lists $0.01 a render with a $100 monthly minimum; Sume Timeline bills $0.10 per output minute. Break-even by volume and what each leaves out.
Written by Sume