TTS mp3 bit rate: 32k to 192k file size per minute for voice audio

Sume TTS 1.0 accepts mp3 bit_rate 32000, 64000, 96000, 128000 or 192000. See the size per minute of each, plus a two-take listening test before you pick one.

4 min readSume
All posts

The default TTS output on Sume is mp3 at 44,100 Hz and 128,000 bits per second. That is a safe default, and not always the smallest choice. A podcast host, an email attachment and an in-app notification do not need the same file. The output_format field lets you choose the bit rate, and a little arithmetic tells you what each choice costs in storage and bandwidth.

The five allowed rates

For container: "mp3", bit_rate accepts exactly 32000, 64000, 96000, 128000 or 192000. Any other number is rejected as a validation error before the job is queued. The file size is bit rate divided by 8, times the seconds. At 128000 that is 16,000 bytes per second, or 960,000 bytes per minute. The table repeats the sum for each rate.

mp3 size per minute of audio by bit rate (arithmetic from the Sume TTS 1.0 schema, read 2026-10-05)
bit_rateKB per minuteMB per 10 minutesMB per hour
320002402.414.4
640004804.828.8
960007207.243.2
128000 (default)9609.657.6
1920001,44014.486.4

What the price does not change

Bit rate is not in the price. TTS 1.0 bills $0.0475 per 1,000 characters whatever the format, so a smaller file saves bandwidth and storage, not credits. Run the comparison on a real sentence, not a theory: one 150-character test line quotes 1 cent per take.

Code

Render the same line at two rates, then listen on the device your audience uses.

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

LINE = "Your table is ready. Please come to the front desk."
for br in (32000, 128000):
    res = run("/tts-1.0/generate", {"transcript": LINE, "language": "en",
              "voice": {"id": os.environ["VOICE_ID"]},
              "output_format": {"container": "mp3", "bit_rate": br}},
              key=f"br-{br}")
    print(br, res["audio_url"], res["output_format"])

How to choose

  • If you will edit the audio later, render wav and export to mp3 last. The wav or mp3 post explains the padding that mp3 adds.
  • For speech played on a phone speaker, try 64000 first. Your ears and speaker decide, so listen.
  • For a download, an archive or a music bed under the voice, keep 128000 or go to 192000.
  • A 20-minute lesson is 19.2 MB at 128000 and 9.6 MB at 64000, so a thousand learners cost 9.6 GB less bandwidth.

Keep the choice

The result echoes output_format. Save it with the audio URL so your next line matches.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume