Apple Podcasts audio specs vs Sume TTS mp3 44.1 kHz 128 kbps

Apple Podcasts wants 44.1 kHz audio, about -16 LKFS and, for WAV or FLAC, stereo. What Sume TTS mp3 output meets and what you must still measure yourself.

5 min readSume
All posts

Apple's audio requirements page for podcasters asks for WAV, FLAC or MP3 uploads at a minimum of 44.1 kHz, and loudness around -16 LKFS with a true peak at or below -1 dBFS. Sume TTS 1.0 defaults to mp3 at 44.1 kHz and 128 kbps, which matches the sample-rate minimum and sits inside Apple's recommended MP3 bit-rate range. Sume's docs do not state a channel layout or loudness normalization, so measure both before you publish.

What Apple's page says

Apple lists separate rules by format. The table keeps the lines relevant to a TTS-made episode.

Apple requirements, read 2026-10-02
FormatRequirement
Accepted filesWAV, FLAC or MP3 in Podcasts Connect; MP3 or AAC through RSS
WAV and FLAC44.1 kHz minimum, 16 or 24 bit, stereo required (single channel rejected)
MP3 mono44.1 kHz and 32 kbps minimum; 96 to 128 kbps recommended
MP3 stereo64 kbps minimum; 128 to 256 kbps recommended
LoudnessAbout -16 dB LKFS plus or minus 1; true peak not above -1 dBFS

Make the file with Sume

Sume's output_format offers 44.1 kHz and 48 kHz, mp3 bit rates up to 192 kbps, and wav. The request below pins the documented mp3 default explicitly.

import os, requests

r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
             "Idempotency-Key": "tts-demo-001"},
    json={"transcript": "Welcome back. Today we cover three updates.",
          "avatar_id": os.environ["SUME_AVATAR_ID"],
          "generation_config": {"speed": 0.95, "emotion": "warm"},
          "output_format": {"container": "mp3", "sample_rate": 44100,
                            "bit_rate": 128000}},
)
print(r.status_code, r.json())

Measure what the docs do not promise

Run the file through ffprobe to read its channel count, and a loudness meter for integrated loudness and true peak. If it is mono and you want WAV or FLAC, convert to stereo in your editor first. If loudness is off, normalize in your editor; do not assume the TTS output lands at -16 LKFS.

Worked example

A quick measurement routine avoids a rejected upload.

  • Run ffprobe on the file and read the sample rate, channels and bit rate.
  • Measure integrated loudness and true peak in your editor or an ebu-r128 tool.
  • If you need WAV, export stereo at 44.1 kHz and 16 or 24 bit.

Checklist before you commit

Apple's loudness lines are targets, not guarantees that every file passes; a mixed episode with music needs its own measurement.

  • Pick MP3 or AAC for RSS delivery.
  • Pick WAV or FLAC only if your files are stereo.
  • Check the page again for changes before each season.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume