Read Sume TTS raw PCM in Python: f32le samples at 24 kHz

Ask the TTS Router for container raw, 24000 Hz and pcm_f32le, then load the samples with the standard library for DSP or a custom player.

5 min readSume
All posts

For raw samples from Sume TTS, set output_format to { "container": "raw", "sample_rate": 24000, "encoding": "pcm_f32le" }. The file you download then has no header: it is a flat run of 32-bit little-endian floats, one per sample, at 24,000 samples per second. In Python, array('f') reads it directly, and dividing the sample count by 24,000 gives the duration. That is the form a DSP script, a feature extractor or a custom audio player wants. The script below does the whole round trip.

The audience for this is the developer who wants the voice as data, not as an MP3. It matters more as speech models get faster: Microsoft's MAI-Voice-2.1 Flash is listed at a vendor-claimed 150 ms end to end (read 2026-10-06 on the October tracker), and once audio arrives that quickly, the next question is what you do with the raw frames. On Sume the answer is simple: ask for them in the format your code uses.

Request raw floats and load them

import array, json, os, sys, time
import urllib.request as u

KEY = os.environ["SUME_API_KEY"]
def call(url, body=None, key=None):
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    if key: h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]

RATE = 24000
job = call("https://api.sume.com/v1/tts-router/generate", {
    "model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
    "transcript": "Raw samples, no container, ready for analysis.",
    "output_format": {"container": "raw", "sample_rate": RATE, "encoding": "pcm_f32le"}},
    "raw-demo-v1")
while not job["terminal"]:
    time.sleep(job.get("next_poll_after_seconds") or 2)
    job = call(job["status_url"])
arts = call(job["result_url"])["result"]["artifacts"]
url = next(a["url"] for a in arts if a["type"] == "audio")
samples = array.array("f")
samples.frombytes(u.urlopen(url).read())
if sys.byteorder == "big":
    samples.byteswap()
peak = max(abs(s) for s in samples)
print(f"{len(samples)} samples, {len(samples) / RATE:.2f} s, peak {peak:.3f}")

What the format means

Little-endian is the format; big-endian machines need the byteswap() call. pcm_f32le samples are floating point and normally sit between -1 and 1, so the peak printed at the end is a fraction of full scale, not an integer. If you need 16-bit integers instead, ask for pcm_s16le, and read with array('h'). The sample rates the API accepts are 8000, 16000, 22050, 24000, 44100 and 48000 Hz. A raw file has no sample rate inside it. Your code has to remember it, which is why the script holds RATE as a constant that it also sends.

Pick 24,000 Hz when you feed the samples to a speech model, since many of them expect that rate, and 44,100 or 48,000 Hz when the audio will be mixed with music. Resampling later is possible, but doing it yourself adds a step and a place to get the rate wrong. Choose the rate at request time and write it in the same constant the reader uses.

Slices need raw or WAV

Raw and WAV are the two containers that support sentence-audio slices. If you add timestamps: { words: true } and segmentation: { mode: "sentence", emit_audio: true }, each segment gets its own audio_url that is cut sample-exact. MP3 cannot do that, because its frames and priming padding do not line up with sentence edges. The segment options are described in sentence audio from TTS.

Output formats and what to use them for (read 2026-10-06)
container + encodingSample rateGood for
mp3 (default), 128 kbps44100 (default)Playing in a page or app
wav, pcm_s16le (default)44100Editors, Timeline joins, lip sync
raw, pcm_f32le24000Analysis, ML features, custom players
raw or wav, pcm_mulaw8000 (set it)Phone systems

Check the byte math

A quick sanity check on the output: a 24,000 Hz float file uses 4 bytes per sample, so one second is 96,000 bytes. If the downloaded size divided by 96,000 does not match the length you expect, the encoding or rate you read with is not the one you requested. That mismatch is the most common mistake with headerless audio, and it produces noise or the wrong pitch, not an error.

Cost and storage

Cost does not depend on the format. A line is billed by character at $0.0475 per 1,000 (pricing, read 2026-10-06), whatever container you pick, and the file is hosted for you as a job artifact (see Jobs and results). Download it once and keep your own copy if you need it long term.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume