Spotify podcast transcript upload: VTT, 5MB and Sume STT parts
Spotify takes VTT or SRT up to 5MB, with timestamps. Build one from Sume STT in 10-minute parts, stitch the cues, and upload from Spotify for Creators.

To upload your own podcast transcript to Spotify, make a VTT or SRT file with timestamps, keep it under 5MB, and add it from the episode's Details page in Spotify for Creators (read 2026-10-03). Sume's speech-to-text returns word timings and sentence segments, so the file is a short script away; the catch is that one Sume STT call covers at most 10 minutes of audio, so a full episode is transcribed in parts and stitched.
This post walks through that route and what Spotify says about editing, syncing and RSS distribution.
What does Spotify require for an uploaded transcript?
Spotify's help page lists the upload tips directly: the transcript must be a VTT or SRT file, the maximum file size is 5MB, and it must include timestamps, otherwise it will not sync during playback. It adds that, depending on your podcast host, you may not be able to sync your transcripts (read 2026-10-03). Transcripts are also not available to every creator, so the Edit control may simply be missing on your show.
Editing is download-and-replace only. Spotify says you cannot edit a transcript inside Spotify for Creators: you open the episode, choose Details, click Edit on the transcript, download the VTT, change it on your device, then choose Upload transcript and Save. That makes the generator script the real editing tool, since fixing a name once in code or in your editor and re-running beats hand-patching timestamps every time a cue moves.
| Topic | What Spotify says | What it means for a Sume-built file |
|---|---|---|
| File type | VTT or SRT | Write WebVTT; it carries the same timing as SRT |
| Size | Max 5MB | Check byte length before uploading |
| Timestamps | Required to sync during playback | Use sentence segments, not plain text |
| Editing | Download the VTT, edit locally, re-upload; no in-app editor | Regenerate and replace the whole file |
| Auto-generated on | Replaces a transcript you uploaded | Leave auto-generated off for that episode |
| Other platforms | Toggle distribution via RSS in Settings | Your RSS host must support transcripts too |
How do I transcribe a full episode with Sume?
The STT route takes a public HTTPS audio_url, returns text, words[] and, when you ask for segmentation.mode: "sentence", gapless segments[] with start and end in seconds from the audio start (STT in the API reference). duration_seconds accepts 1 to 600, so each call tops out at 10 minutes. The public rate is $0.01 per audio minute, so confirm the live price in GET /v1/catalog before a back catalogue run.
Cut the episode into 600-second WAV parts first. On a test file, ffmpeg's segment muxer landed within a few milliseconds of the nominal length, which is fine for transcript cues:
ffmpeg -i episode.mp3 -f segment -segment_time 600 \
-ar 16000 -ac 1 -c:a pcm_s16le ep-%03d.wavHow do I stitch the parts into one VTT?
Host each part at a public HTTPS address, transcribe it, and add part_index * 600 seconds to every segment so the cues land on the episode timeline. The script below does that and prints the byte size so you can compare it with Spotify's 5MB limit. A sentence can straddle a cut point; that shows up as two short cues at the join, so listen to the seam once.
Run the parts in order, not in parallel, if your workspace has a low concurrency limit: each accepted STT job counts toward the queue limits shown in the job response, and a one-hour episode is six jobs. Parallel submission is fine when the limits allow it, since the offsets come from the part index and not from completion order.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PART = 600 # seconds per ffmpeg part
def stamp(s):
ms = round(s * 1000)
return f"{ms // 3600000:02d}:{ms // 60000 % 60:02d}:{ms // 1000 % 60:02d}.{ms % 1000:03d}"
def transcribe(url):
body = {"audio_url": url, "duration_seconds": PART, "segmentation": {"mode": "sentence"}}
job = requests.post(f"{API}/v1/stt-1.0/transcribe", headers=H, json=body).json()["data"]
while not (st := requests.get(job["status_url"], headers=H).json()["data"])["terminal"]:
time.sleep(2)
if not st["result_ready"]:
raise SystemExit(f"{url}: {st['sume_status']}")
return requests.get(job["result_url"], headers=H).json()["data"]["result"]["segments"]
urls = ["https://example.com/public/ep-000.wav", "https://example.com/public/ep-001.wav"]
lines = ["WEBVTT", ""]
for n, url in enumerate(urls):
for g in transcribe(url):
a, b = g["start"] + n * PART, g["end"] + n * PART
lines += [f"{stamp(a)} --> {stamp(b)}", g["text"].strip(), ""]
text = "\n".join(lines)
open("episode.vtt", "w", encoding="utf-8").write(text)
print(len(text.encode()), "bytes; Spotify's limit is 5MB")
What should I check before and after upload?
Spotify has no in-app transcript editor, so each fix is a download, edit and re-upload cycle. Catch problems before the first upload.
- Spot-check names, brands and numbers: Sume returns speech-to-text wording and does not know your guest list.
- Sume STT has no public speaker labels, so add speaker names in your editor if the show needs them.
- Compare the file size to 5MB and confirm the first cue starts near 00:00:00.000.
- Decide per episode whether Spotify's auto-generated transcript or yours should be on, because enabling auto-generated replaces an uploaded file.
- Check that your host passes transcripts through RSS if you want other apps to show them.
Where does this fall short?
If your episodes are video, a transcript does not put words on screen. For short promo clips, burn captions with video captions, which takes a video URL and is priced at $0.20 per job for clips up to 60 seconds under the current estimate. For the speaker-label gap, see how other diarization APIs compare. For one-off URLs rather than parts, podcast transcription from the episode URL covers cost per part.
Sources
Related posts
More in Media tools
- SRT or WebVTT? What Vimeo, Spotify, Apple and Cloudflare accept
Vimeo, Spotify and Apple Podcasts take SRT or WebVTT; Cloudflare Stream documents WebVTT. A table read 2026-10-03 and a script writing both from Sume segments.
- Swap the model in a clothing video with AI: H3 Max Recast
Recast swaps the person in a clip for a person from a photo, 5 to 30 seconds. What it keeps, what the docs leave open about the garment, and the Sume call.
- Turn a ChatGPT try-on image into a video with Seedance 2.5
Saved a try-on image to your ChatGPT Library? Host it, then use it as the first frame of a 9:16 clip on seedance-2.5 through POST /v1/videos. A working script.
- Turn a hum into music with AI: what Sume takes as input
Stability says hum-to-steer is coming. Sume's music API takes text and one optional image, not audio. Here is how to describe a hummed tune in a prompt.
Written by Sume