Apple Podcasts transcript speaker names: VTT, and Sume STT gaps
Apple shows speaker names when you provide a VTT. Sume STT has no speaker labels, so transcribe each track and merge with voice tags. Script included.

Apple Podcasts shows a speaker's name in the transcript when you provide a VTT file that identifies each speaker; Apple's page names the VTT as the way to do it (read 2026-10-03). Sume's speech-to-text has no public speaker labels, so the workable route is to transcribe each speaker's own recording separately and merge the cues with a voice tag per speaker.
That requires multitrack source audio. If you only have a mixed-down file, speaker names have to be added by hand, and this post says where.
What does Apple say about speaker names and your own transcript?
Apple's creator page says that providing a VTT file lets you identify every speaker with each line, and that any time the speaker changes, the name is displayed in the transcript view. It also says Apple generates a transcript automatically after a new episode is published, that your own transcript is ingested through the RSS transcript tag, and that for subscriber episodes you can upload a VTT or SRT file in Apple Podcasts Connect (read 2026-10-03).
The page does not spell out the markup for a speaker inside the VTT. The W3C WebVTT specification defines a voice span, written as <v Name>text, as the way to say who is talking in a cue (read 2026-10-03), so that is the form to use. Before you push it to a whole show, check one episode in Apple's preview to confirm the names render.
| Topic | What Apple says | Consequence for a Sume-built VTT |
|---|---|---|
| File type | VTT or SRT | Use VTT to carry speaker names |
| Delivery | RSS transcript tag; Connect upload for subscriber episodes | Your host must support the tag |
| Languages | English, Danish, Dutch, Finnish, French, German, Italian, Norwegian, Portuguese, Spanish, Swedish | Other languages need a link in the show notes |
| Quality | Files that do not meet standards are not displayed | Proofread names and punctuation |
| Coverage | Include all segments of the episode | Do not drop the intro or the ad reads |
| Attribution | Shows "Provided by" plus the show name | Expect the label to change from Automatically generated |
Why does Sume STT not label speakers?
The request schema states that provider knobs such as diarize are fixed server-side, and results carry text, words[] and optional sentence segments[] with no speaker field (API reference). Treat speaker attribution as something you supply. Other posts cover the same gap against other vendors: see the Gemini diarization comparison and the MAI-Transcribe-2 comparison.
How do I add speaker names from separate tracks?
If the show was recorded with one file per person, transcribe each file with POST /v1/stt-1.0/transcribe and segmentation.mode: "sentence" (up to 600 seconds per call, $0.01 per audio minute under the current public rate), then merge by start time. Each person's track only contains that person's voice plus a little bleed, so the segments are attributable by construction. The merge below tags each cue with a voice span; two people talking at once simply produce overlapping cues, which WebVTT allows.
The sample uses inline segments so it runs as is; in practice each list is result["segments"] from one STT job.
def stamp(s):
ms = round(s * 1000)
return f"{ms // 3600000:02d}:{ms // 60000 % 60:02d}:{ms // 1000 % 60:02d}.{ms % 1000:03d}"
# Each value is result["segments"] from one Sume STT job on that speaker's own track.
tracks = {
"Host": [{"start": 0.0, "end": 4.1, "text": "Welcome back to the show."}],
"Guest": [{"start": 4.2, "end": 9.0, "text": "Thanks for having me."},
{"start": 9.1, "end": 11.0, "text": " Happy to be here."}],
}
cues = sorted((g["start"], g["end"], who, g["text"])
for who, segs in tracks.items() for g in segs)
lines = ["WEBVTT", ""]
for a, b, who, text in cues:
lines += [f"{stamp(a)} --> {stamp(b)}", f"<v {who}>{text.strip()}", ""]
open("episode.vtt", "w", encoding="utf-8").write("\n".join(lines))
What if I only have a mixed-down episode?
Then Sume gives you timed sentences without names, and the names are manual work. A practical sequence is to generate the VTT, open it beside the audio, and prefix each speaker change yourself. Keep the cue text and timing from Sume and add only the voice tag.
- Start from sentence segments so each cue is one thought and a speaker change falls on a cue boundary.
- Add
<v Host>or<v Guest>at the start of each cue; a voice span that covers the whole cue text needs no closing tag. - Spell names the way you want them displayed, since Apple recommends putting host and guest names in the show and episode descriptions for accurate spelling in its own transcripts.
- Keep the language set correctly in your feed; Apple lists that as a best practice.
When should I just leave Apple's transcript on?
Apple's automatic transcript appears shortly after publication without any file from you. It does not display segments that changed through dynamically inserted audio, and it omits music lyrics, so a show with inserted ads gets gaps. A provided file fixes the gaps and the names, at the price of owning accuracy: Apple says files must be free of spelling and punctuation errors to pass its standards.
For video clips cut from the episode, a transcript is not the deliverable. Burn the words onto the picture with video captions instead, and keep the VTT for the feed.
Sources
Related posts
- Gemini API speaker diarization: 8 speakers; Sume STT has no labels
- MAI-Transcribe-2 speaker diarization vs Sume STT: no speaker field
- Podcast transcription from the episode URL: parts and cost
- How to transcribe an interview and label each speaker
- Apple Podcasts host-read video ads: cut the spot with Sume trim
More in Media tools
- FLUX 21:9 ultrawide image: FLUX.2 on Sume vs FLUX 3 ratios
FLUX 3 Image lists ratios from 21:9 to 9:21. On Sume, FLUX.2 Pro and Flex take 21:9 and 9:21 too, from a 13-ratio list with no auto. A banner request in Python.
- Spotify podcast transcript upload: VTT, 5MB and Sume STT parts
Spotify takes VTT or SRT up to 5MB, with timestamps. Build one from Sume STT in 10-minute parts, stitch the cues, and upload from Spotify for Creators.
- SRT or WebVTT? What Vimeo, Spotify, Apple and Cloudflare accept
Vimeo, Spotify and Apple Podcasts take SRT or WebVTT; Cloudflare Stream documents WebVTT. A table read 2026-10-03 and a script writing both from Sume segments.
- Swap the model in a clothing video with AI: H3 Max Recast
Recast swaps the person in a clip for a person from a photo, 5 to 30 seconds. What it keeps, what the docs leave open about the garment, and the Sume call.
Written by Sume