Spotify podcast transcript upload: VTT, 5MB and Sume STT parts

Spotify takes VTT or SRT up to 5MB, with timestamps. Build one from Sume STT in 10-minute parts, stitch the cues, and upload from Spotify for Creators.

5 min readSume
All posts

To upload your own podcast transcript to Spotify, make a VTT or SRT file with timestamps, keep it under 5MB, and add it from the episode's Details page in Spotify for Creators (read 2026-10-03). Sume's speech-to-text returns word timings and sentence segments, so the file is a short script away; the catch is that one Sume STT call covers at most 10 minutes of audio, so a full episode is transcribed in parts and stitched.

This post walks through that route and what Spotify says about editing, syncing and RSS distribution.

What does Spotify require for an uploaded transcript?

Spotify's help page lists the upload tips directly: the transcript must be a VTT or SRT file, the maximum file size is 5MB, and it must include timestamps, otherwise it will not sync during playback. It adds that, depending on your podcast host, you may not be able to sync your transcripts (read 2026-10-03). Transcripts are also not available to every creator, so the Edit control may simply be missing on your show.

Editing is download-and-replace only. Spotify says you cannot edit a transcript inside Spotify for Creators: you open the episode, choose Details, click Edit on the transcript, download the VTT, change it on your device, then choose Upload transcript and Save. That makes the generator script the real editing tool, since fixing a name once in code or in your editor and re-running beats hand-patching timestamps every time a cue moves.

Spotify transcript handling for creators, read 2026-10-03
TopicWhat Spotify saysWhat it means for a Sume-built file
File typeVTT or SRTWrite WebVTT; it carries the same timing as SRT
SizeMax 5MBCheck byte length before uploading
TimestampsRequired to sync during playbackUse sentence segments, not plain text
EditingDownload the VTT, edit locally, re-upload; no in-app editorRegenerate and replace the whole file
Auto-generated onReplaces a transcript you uploadedLeave auto-generated off for that episode
Other platformsToggle distribution via RSS in SettingsYour RSS host must support transcripts too

How do I transcribe a full episode with Sume?

The STT route takes a public HTTPS audio_url, returns text, words[] and, when you ask for segmentation.mode: "sentence", gapless segments[] with start and end in seconds from the audio start (STT in the API reference). duration_seconds accepts 1 to 600, so each call tops out at 10 minutes. The public rate is $0.01 per audio minute, so confirm the live price in GET /v1/catalog before a back catalogue run.

Cut the episode into 600-second WAV parts first. On a test file, ffmpeg's segment muxer landed within a few milliseconds of the nominal length, which is fine for transcript cues:

ffmpeg -i episode.mp3 -f segment -segment_time 600 \
  -ar 16000 -ac 1 -c:a pcm_s16le ep-%03d.wav

How do I stitch the parts into one VTT?

Host each part at a public HTTPS address, transcribe it, and add part_index * 600 seconds to every segment so the cues land on the episode timeline. The script below does that and prints the byte size so you can compare it with Spotify's 5MB limit. A sentence can straddle a cut point; that shows up as two short cues at the join, so listen to the seam once.

Run the parts in order, not in parallel, if your workspace has a low concurrency limit: each accepted STT job counts toward the queue limits shown in the job response, and a one-hour episode is six jobs. Parallel submission is fine when the limits allow it, since the offsets come from the part index and not from completion order.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PART = 600  # seconds per ffmpeg part

def stamp(s):
    ms = round(s * 1000)
    return f"{ms // 3600000:02d}:{ms // 60000 % 60:02d}:{ms // 1000 % 60:02d}.{ms % 1000:03d}"

def transcribe(url):
    body = {"audio_url": url, "duration_seconds": PART, "segmentation": {"mode": "sentence"}}
    job = requests.post(f"{API}/v1/stt-1.0/transcribe", headers=H, json=body).json()["data"]
    while not (st := requests.get(job["status_url"], headers=H).json()["data"])["terminal"]:
        time.sleep(2)
    if not st["result_ready"]:
        raise SystemExit(f"{url}: {st['sume_status']}")
    return requests.get(job["result_url"], headers=H).json()["data"]["result"]["segments"]

urls = ["https://example.com/public/ep-000.wav", "https://example.com/public/ep-001.wav"]
lines = ["WEBVTT", ""]
for n, url in enumerate(urls):
    for g in transcribe(url):
        a, b = g["start"] + n * PART, g["end"] + n * PART
        lines += [f"{stamp(a)} --> {stamp(b)}", g["text"].strip(), ""]
text = "\n".join(lines)
open("episode.vtt", "w", encoding="utf-8").write(text)
print(len(text.encode()), "bytes; Spotify's limit is 5MB")

What should I check before and after upload?

Spotify has no in-app transcript editor, so each fix is a download, edit and re-upload cycle. Catch problems before the first upload.

  • Spot-check names, brands and numbers: Sume returns speech-to-text wording and does not know your guest list.
  • Sume STT has no public speaker labels, so add speaker names in your editor if the show needs them.
  • Compare the file size to 5MB and confirm the first cue starts near 00:00:00.000.
  • Decide per episode whether Spotify's auto-generated transcript or yours should be on, because enabling auto-generated replaces an uploaded file.
  • Check that your host passes transcripts through RSS if you want other apps to show them.

Where does this fall short?

If your episodes are video, a transcript does not put words on screen. For short promo clips, burn captions with video captions, which takes a video URL and is priced at $0.20 per job for clips up to 60 seconds under the current estimate. For the speaker-label gap, see how other diarization APIs compare. For one-off URLs rather than parts, podcast transcription from the episode URL covers cost per part.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume