OpenAI TTS has Opus, AAC and FLAC; Sume TTS has MP3, WAV and raw

OpenAI's speech API lists six output formats. Sume TTS 1.0 returns MP3, WAV or raw PCM, so Opus, AAC and FLAC need a conversion step after generation.

5 min readSume
All posts

OpenAI's speech endpoint lists MP3, Opus, AAC, FLAC, WAV and PCM as output formats, while Sume's TTS 1.0 route returns an mp3, wav or raw container. If your player needs Opus, AAC or FLAC, generate WAV on Sume and convert it with your own tool afterwards, because Sume does not emit those three.

OpenAI's list is from its text-to-speech guide (read 2026-10-10). The Sume side is from the Sume OpenAPI contract for POST /v1/tts-1.0/generate.

Side by side

The two lists overlap on MP3, WAV and raw samples, and differ on the compressed formats. The OpenAI guide describes PCM as raw 24 kHz 16-bit signed little-endian samples without a header, and recommends wav or pcm for the fastest response. Sume's raw container is paired with an explicit encoding and sample rate that you choose.

Output formats (OpenAI read 2026-10-10; Sume per OpenAPI, read 2026-10-10)
FormatOpenAI speechSume TTS 1.0
MP3Yes (default)Yes, bit rate 32000 to 192000
WAVYesYes
Raw PCMYes, 24 kHz 16-bitYes, raw container with pcm_s16le, pcm_f32le, pcm_mulaw or pcm_alaw
OpusYesNo
AACYesNo
FLACYesNo

What Sume lets you tune instead

Sume's output_format block gives you control that the OpenAI list does not describe: sample rates of 8000, 16000, 22050, 24000, 44100 or 48000 Hz, MP3 bit rates from 32,000 to 192,000, and the four raw encodings. For telephony you can pick 8 kHz mu-law or A-law raw. For editing you can pick 48 kHz WAV. That covers the cases most Opus and FLAC requests come from, although it will not give you an Opus stream for WebRTC.

  • Streaming to a browser: MP3 or WAV; Sume jobs are request and result, not chunked audio.
  • Archiving: WAV, then compress to FLAC on your side.
  • Mobile app bundles that want AAC: convert from WAV with your own encoder.
  • Phone systems: raw pcm_mulaw or pcm_alaw at 8000 Hz.

Other differences that matter when you switch

OpenAI lists 13 built-in voices by name. Sume does not use named stock voices on this route; the voice.id field is a UUID or a voi_ id, and anything else returns 400 invalid_voice_id. Input is capped at 20,000 characters per job on Sume, with a 1,200 second ceiling on the generated audio (tts_duration_exceeded). The OpenAI guide did not state an input limit on the page I read, so I make no comparison there.

Pricing is a flat per-character rate on Sume: $0.0475 per 1,000 characters, with a one-cent minimum per job and a maximum of 95 cents. I did not read an OpenAI price on the page cited above, so there is no price comparison in this post.

A practical rule

Decide your delivery format first. If it is MP3, WAV or telephony raw, Sume covers it in one call. If it is Opus, AAC or FLAC, treat Sume as the generation step and budget a local conversion; do not expect the job to return a file in a format its schema does not list.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume