How to transcribe an interview and label each speaker
Transcribe an interview in two steps: speech-to-text for the words and their times, then speaker labels from separate tracks or a pass by hand.

To transcribe an interview, run the recording through speech-to-text to get the words and their timestamps, then mark who said what and lay the text out as turns: speaker, time, words. Telling the speakers apart is a separate step, called speaker diarization. Sume's STT 1.0 returns the words with their start and end times but no speaker labels, so the labels come from recording each person on a separate track, or from a pass over the sentences by hand.
The STT facts come from the STT 1.0 schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28; anything described as current behavior is read from Sume's code. The channel split follows the FFmpeg filters documentation.
What is speaker diarization, and does Sume do it?
Diarization splits a recording by voice and tags each stretch Speaker 1, Speaker 2, and so on. It tells voices apart; it doesn't know anyone's name. STT 1.0 doesn't offer it: the API reference says diarize is fixed server-side, and current code runs it off and rejects a request that sends diarize.
What you do get is text, the whole transcript, and words, with start and end in seconds from the start of the audio, ordered by start. Ask for sentence segmentation and you also get segments.
How do I send the recording to speech-to-text?
Send a link, not the file: STT 1.0 takes one public HTTPS audio_url per request and has no field for file bytes. Add segmentation: {"mode": "sentence"} to get the sentence segments you need for labeling by hand; language_code is an optional hint, and the language is detected without it. Speech-to-text API with word timestamps covers the other request fields.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: interview-guest-track-001" \
-d '{
"audio_url": "https://example.com/audio/interview-guest.m4a",
"language_code": "en",
"segmentation": { "mode": "sentence" }
}'How do I label speakers when each person has their own track?
Let the recording do it. If each person was recorded on a separate track, transcribe each track as its own job, tag every word with that track's speaker, then sort all the words by start and begin a new turn whenever the speaker changes.
- When both people talk at once, their words interleave by time; tidy those spots by hand.
- In current code a word carries
startandendonly when the engine returns them, so the merge below skips words without a time. It also skips entries whosetypeisspacing.
// results: one STT 1.0 result per speaker's track, e.g. { Host: a, Guest: b }
const words = Object.entries(results).flatMap(([speaker, result]) =>
result.words
.filter((w) => w.type !== "spacing" && typeof w.start === "number")
.map((w) => ({ speaker, start: w.start, text: w.word })),
);
words.sort((a, b) => a.start - b.start);
const turns = [];
for (const w of words) {
const last = turns[turns.length - 1];
if (last && last.speaker === w.speaker) last.text += " " + w.text;
else turns.push({ speaker: w.speaker, start: w.start, text: w.text });
}How do I transcribe a recorded phone call?
If the recording keeps each side of the call on its own channel of a stereo file, split the channels, put each file at its own public HTTPS URL, and treat them as two tracks. FFmpeg's channelsplit filter puts each channel of an input into its own stream; its default layout is stereo, whose channels FFmpeg names FL (front left) and FR (front right). If both sides are mixed into one channel, label the sentences by hand instead.
ffmpeg -i call.wav -filter_complex "channelsplit=channel_layout=stereo[FL][FR]" \
-map "[FL]" side-a.wav -map "[FR]" side-b.wavHow do I label speakers on a single recording?
Label the sentences. In sentence mode, STT 1.0 groups words on terminal punctuation and splits unpunctuated runs on silence; each segment has index, text, start, end, and duration_seconds. Play the audio from each segment's start, write the speaker next to it, and merge neighbors with the same speaker into one turn. The boundaries come from punctuation and silence, not from voices, so check them where one person cuts in on another.
How long can the interview be, and what does it cost?
A 45-minute interview takes 5 requests per track. Transcribed as one mixed track it costs $0.45, and as two separate tracks $0.90, plus a 5.5% agent fee by default. Transcribe long audio files shows how to split a recording and shift each part's times.
| Limit | STT 1.0 |
|---|---|
| Audio per request | Up to 10 minutes: duration_seconds runs 1–600 |
| Input | One public HTTPS audio_url per request |
| Speaker labels | None; diarize is fixed server-side |
| Price | $0.01 per audio minute, for each track you transcribe |
Sources
Related posts
More in Media tools
- How to upscale video to 4K: scale factors and step order
To upscale video to 4K, enlarge it by 2160 ÷ its height: ×2 from 1080p, ×3 from 720p. How to do it with FFmpeg or Sume, and why to upscale last.
- Hulu ad specs: Disney's video rules for Hulu and Disney+
Hulu and Disney+ ads must be 16:9 (1920×1080 preferred), H.264 or ProRes at 23.98–30 fps, with one two-channel audio track of 128 kbps or more.
- How to insert a video into another video
To insert a video into another video, split the main video at the insert point and play it, the new clip, then the rest, all in one render.
- J cut and L cut: what they are and how to make one
A J cut plays the next shot's sound before its picture; an L cut lets a shot's sound run on under the next picture. Here's how to build each one.
Written by Sume