Apple Podcasts transcript speaker names: VTT, and Sume STT gaps

Apple shows speaker names when you provide a VTT. Sume STT has no speaker labels, so transcribe each track and merge with voice tags. Script included.

5 min readSume
All posts

Apple Podcasts shows a speaker's name in the transcript when you provide a VTT file that identifies each speaker; Apple's page names the VTT as the way to do it (read 2026-10-03). Sume's speech-to-text has no public speaker labels, so the workable route is to transcribe each speaker's own recording separately and merge the cues with a voice tag per speaker.

That requires multitrack source audio. If you only have a mixed-down file, speaker names have to be added by hand, and this post says where.

What does Apple say about speaker names and your own transcript?

Apple's creator page says that providing a VTT file lets you identify every speaker with each line, and that any time the speaker changes, the name is displayed in the transcript view. It also says Apple generates a transcript automatically after a new episode is published, that your own transcript is ingested through the RSS transcript tag, and that for subscriber episodes you can upload a VTT or SRT file in Apple Podcasts Connect (read 2026-10-03).

The page does not spell out the markup for a speaker inside the VTT. The W3C WebVTT specification defines a voice span, written as <v Name>text, as the way to say who is talking in a cue (read 2026-10-03), so that is the form to use. Before you push it to a whole show, check one episode in Apple's preview to confirm the names render.

Apple Podcasts provided-transcript rules, read 2026-10-03
TopicWhat Apple saysConsequence for a Sume-built VTT
File typeVTT or SRTUse VTT to carry speaker names
DeliveryRSS transcript tag; Connect upload for subscriber episodesYour host must support the tag
LanguagesEnglish, Danish, Dutch, Finnish, French, German, Italian, Norwegian, Portuguese, Spanish, SwedishOther languages need a link in the show notes
QualityFiles that do not meet standards are not displayedProofread names and punctuation
CoverageInclude all segments of the episodeDo not drop the intro or the ad reads
AttributionShows "Provided by" plus the show nameExpect the label to change from Automatically generated

Why does Sume STT not label speakers?

The request schema states that provider knobs such as diarize are fixed server-side, and results carry text, words[] and optional sentence segments[] with no speaker field (API reference). Treat speaker attribution as something you supply. Other posts cover the same gap against other vendors: see the Gemini diarization comparison and the MAI-Transcribe-2 comparison.

How do I add speaker names from separate tracks?

If the show was recorded with one file per person, transcribe each file with POST /v1/stt-1.0/transcribe and segmentation.mode: "sentence" (up to 600 seconds per call, $0.01 per audio minute under the current public rate), then merge by start time. Each person's track only contains that person's voice plus a little bleed, so the segments are attributable by construction. The merge below tags each cue with a voice span; two people talking at once simply produce overlapping cues, which WebVTT allows.

The sample uses inline segments so it runs as is; in practice each list is result["segments"] from one STT job.

def stamp(s):
    ms = round(s * 1000)
    return f"{ms // 3600000:02d}:{ms // 60000 % 60:02d}:{ms // 1000 % 60:02d}.{ms % 1000:03d}"

# Each value is result["segments"] from one Sume STT job on that speaker's own track.
tracks = {
    "Host": [{"start": 0.0, "end": 4.1, "text": "Welcome back to the show."}],
    "Guest": [{"start": 4.2, "end": 9.0, "text": "Thanks for having me."},
              {"start": 9.1, "end": 11.0, "text": " Happy to be here."}],
}
cues = sorted((g["start"], g["end"], who, g["text"])
              for who, segs in tracks.items() for g in segs)
lines = ["WEBVTT", ""]
for a, b, who, text in cues:
    lines += [f"{stamp(a)} --> {stamp(b)}", f"<v {who}>{text.strip()}", ""]
open("episode.vtt", "w", encoding="utf-8").write("\n".join(lines))

What if I only have a mixed-down episode?

Then Sume gives you timed sentences without names, and the names are manual work. A practical sequence is to generate the VTT, open it beside the audio, and prefix each speaker change yourself. Keep the cue text and timing from Sume and add only the voice tag.

  • Start from sentence segments so each cue is one thought and a speaker change falls on a cue boundary.
  • Add <v Host> or <v Guest> at the start of each cue; a voice span that covers the whole cue text needs no closing tag.
  • Spell names the way you want them displayed, since Apple recommends putting host and guest names in the show and episode descriptions for accurate spelling in its own transcripts.
  • Keep the language set correctly in your feed; Apple lists that as a best practice.

When should I just leave Apple's transcript on?

Apple's automatic transcript appears shortly after publication without any file from you. It does not display segments that changed through dynamically inserted audio, and it omits music lyrics, so a show with inserted ads gets gaps. A provided file fixes the gaps and the names, at the price of owning accuracy: Apple says files must be free of spelling and punctuation errors to pass its standards.

For video clips cut from the episode, a transcript is not the deliverable. Burn the words onto the picture with video captions instead, and keep the VTT for the feed.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume