Read Sume TTS raw PCM in Python: f32le samples at 24 kHz
Ask the TTS Router for container raw, 24000 Hz and pcm_f32le, then load the samples with the standard library for DSP or a custom player.

For raw samples from Sume TTS, set output_format to { "container": "raw", "sample_rate": 24000, "encoding": "pcm_f32le" }. The file you download then has no header: it is a flat run of 32-bit little-endian floats, one per sample, at 24,000 samples per second. In Python, array('f') reads it directly, and dividing the sample count by 24,000 gives the duration. That is the form a DSP script, a feature extractor or a custom audio player wants. The script below does the whole round trip.
The audience for this is the developer who wants the voice as data, not as an MP3. It matters more as speech models get faster: Microsoft's MAI-Voice-2.1 Flash is listed at a vendor-claimed 150 ms end to end (read 2026-10-06 on the October tracker), and once audio arrives that quickly, the next question is what you do with the raw frames. On Sume the answer is simple: ask for them in the format your code uses.
Request raw floats and load them
import array, json, os, sys, time
import urllib.request as u
KEY = os.environ["SUME_API_KEY"]
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
RATE = 24000
job = call("https://api.sume.com/v1/tts-router/generate", {
"model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
"transcript": "Raw samples, no container, ready for analysis.",
"output_format": {"container": "raw", "sample_rate": RATE, "encoding": "pcm_f32le"}},
"raw-demo-v1")
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(job["status_url"])
arts = call(job["result_url"])["result"]["artifacts"]
url = next(a["url"] for a in arts if a["type"] == "audio")
samples = array.array("f")
samples.frombytes(u.urlopen(url).read())
if sys.byteorder == "big":
samples.byteswap()
peak = max(abs(s) for s in samples)
print(f"{len(samples)} samples, {len(samples) / RATE:.2f} s, peak {peak:.3f}")What the format means
Little-endian is the format; big-endian machines need the byteswap() call. pcm_f32le samples are floating point and normally sit between -1 and 1, so the peak printed at the end is a fraction of full scale, not an integer. If you need 16-bit integers instead, ask for pcm_s16le, and read with array('h'). The sample rates the API accepts are 8000, 16000, 22050, 24000, 44100 and 48000 Hz. A raw file has no sample rate inside it. Your code has to remember it, which is why the script holds RATE as a constant that it also sends.
Pick 24,000 Hz when you feed the samples to a speech model, since many of them expect that rate, and 44,100 or 48,000 Hz when the audio will be mixed with music. Resampling later is possible, but doing it yourself adds a step and a place to get the rate wrong. Choose the rate at request time and write it in the same constant the reader uses.
Slices need raw or WAV
Raw and WAV are the two containers that support sentence-audio slices. If you add timestamps: { words: true } and segmentation: { mode: "sentence", emit_audio: true }, each segment gets its own audio_url that is cut sample-exact. MP3 cannot do that, because its frames and priming padding do not line up with sentence edges. The segment options are described in sentence audio from TTS.
| container + encoding | Sample rate | Good for |
|---|---|---|
| mp3 (default), 128 kbps | 44100 (default) | Playing in a page or app |
| wav, pcm_s16le (default) | 44100 | Editors, Timeline joins, lip sync |
| raw, pcm_f32le | 24000 | Analysis, ML features, custom players |
| raw or wav, pcm_mulaw | 8000 (set it) | Phone systems |
Check the byte math
A quick sanity check on the output: a 24,000 Hz float file uses 4 bytes per sample, so one second is 96,000 bytes. If the downloaded size divided by 96,000 does not match the length you expect, the encoding or rate you read with is not the one you requested. That mismatch is the most common mistake with headerless audio, and it produces noise or the wrong pitch, not an error.
Cost and storage
Cost does not depend on the format. A line is billed by character at $0.0475 per 1,000 (pricing, read 2026-10-06), whatever container you pick, and the file is hosted for you as a job artifact (see Jobs and results). Download it once and keep your own copy if you need it long term.
Sources
Related posts
More in Developers
- Record the Idempotency-Key before you POST a Sume video job
A worker that dies after submitting a video job loses the job id. Write the key to a ledger first, then POST; a retry returns the same job. Tested SQLite code.
- Redact sume_live_ API keys from Python logs with a handler filter
A logging.Filter on a logger skips records from child loggers. Put the redaction filter on the handler so a sume_live_ key never reaches the log file.
- Redis sorted set scheduler for AI video job polling in Python
Keep video job ids in a Redis ZSET scored by next-poll time, claim due ids with ZREM so no two workers poll one job, reschedule by next_poll_after_seconds.
- Refuse an empty Sume API key or webhook secret: a fail-fast loader
An empty SUME_API_KEY or webhook secret fails late and confusingly. Load both at boot, reject empty or whitespace values, and never print them. Python, tested.
Written by Sume