How to transcribe an interview and label each speaker

Transcribe an interview in two steps: speech-to-text for the words and their times, then speaker labels from separate tracks or a pass by hand.

5 min readSume
All posts

To transcribe an interview, run the recording through speech-to-text to get the words and their timestamps, then mark who said what and lay the text out as turns: speaker, time, words. Telling the speakers apart is a separate step, called speaker diarization. Sume's STT 1.0 returns the words with their start and end times but no speaker labels, so the labels come from recording each person on a separate track, or from a pass over the sentences by hand.

The STT facts come from the STT 1.0 schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28; anything described as current behavior is read from Sume's code. The channel split follows the FFmpeg filters documentation.

What is speaker diarization, and does Sume do it?

Diarization splits a recording by voice and tags each stretch Speaker 1, Speaker 2, and so on. It tells voices apart; it doesn't know anyone's name. STT 1.0 doesn't offer it: the API reference says diarize is fixed server-side, and current code runs it off and rejects a request that sends diarize.

What you do get is text, the whole transcript, and words, with start and end in seconds from the start of the audio, ordered by start. Ask for sentence segmentation and you also get segments.

How do I send the recording to speech-to-text?

Send a link, not the file: STT 1.0 takes one public HTTPS audio_url per request and has no field for file bytes. Add segmentation: {"mode": "sentence"} to get the sentence segments you need for labeling by hand; language_code is an optional hint, and the language is detected without it. Speech-to-text API with word timestamps covers the other request fields.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: interview-guest-track-001" \
  -d '{
    "audio_url": "https://example.com/audio/interview-guest.m4a",
    "language_code": "en",
    "segmentation": { "mode": "sentence" }
  }'

How do I label speakers when each person has their own track?

Let the recording do it. If each person was recorded on a separate track, transcribe each track as its own job, tag every word with that track's speaker, then sort all the words by start and begin a new turn whenever the speaker changes.

  • When both people talk at once, their words interleave by time; tidy those spots by hand.
  • In current code a word carries start and end only when the engine returns them, so the merge below skips words without a time. It also skips entries whose type is spacing.
// results: one STT 1.0 result per speaker's track, e.g. { Host: a, Guest: b }
const words = Object.entries(results).flatMap(([speaker, result]) =>
  result.words
    .filter((w) => w.type !== "spacing" && typeof w.start === "number")
    .map((w) => ({ speaker, start: w.start, text: w.word })),
);
words.sort((a, b) => a.start - b.start);

const turns = [];
for (const w of words) {
  const last = turns[turns.length - 1];
  if (last && last.speaker === w.speaker) last.text += " " + w.text;
  else turns.push({ speaker: w.speaker, start: w.start, text: w.text });
}

How do I transcribe a recorded phone call?

If the recording keeps each side of the call on its own channel of a stereo file, split the channels, put each file at its own public HTTPS URL, and treat them as two tracks. FFmpeg's channelsplit filter puts each channel of an input into its own stream; its default layout is stereo, whose channels FFmpeg names FL (front left) and FR (front right). If both sides are mixed into one channel, label the sentences by hand instead.

ffmpeg -i call.wav -filter_complex "channelsplit=channel_layout=stereo[FL][FR]" \
  -map "[FL]" side-a.wav -map "[FR]" side-b.wav

How do I label speakers on a single recording?

Label the sentences. In sentence mode, STT 1.0 groups words on terminal punctuation and splits unpunctuated runs on silence; each segment has index, text, start, end, and duration_seconds. Play the audio from each segment's start, write the speaker next to it, and merge neighbors with the same speaker into one turn. The boundaries come from punctuation and silence, not from voices, so check them where one person cuts in on another.

How long can the interview be, and what does it cost?

A 45-minute interview takes 5 requests per track. Transcribed as one mixed track it costs $0.45, and as two separate tracks $0.90, plus a 5.5% agent fee by default. Transcribe long audio files shows how to split a recording and shift each part's times.

From the STT 1.0 schema in the Sume API reference and API pricing, read 2026-09-28.
LimitSTT 1.0
Audio per requestUp to 10 minutes: duration_seconds runs 1–600
InputOne public HTTPS audio_url per request
Speaker labelsNone; diarize is fixed server-side
Price$0.01 per audio minute, for each track you transcribe

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume