Label speakers without diarization: transcribe each mic track, merge
Sume STT has no speaker labels. If you record each speaker on a separate track, transcribe both tracks and merge the segments by start time. Python script.

Sume speech-to-text does not label speakers: the documented result has the text, words, sentence segments and a language code, and no speaker field, and the diarization switch is not a request option (Sume fixes it server-side). If you recorded each person on their own track, though, you can get a labeled transcript anyway. Transcribe each track as a separate job with sentence segmentation, tag every segment with the track's speaker name, then sort all segments by start time. The script below does it for two tracks. It only works with separate recordings. A single mixed file with two voices is a different problem that this method cannot solve.
Meeting and interview transcripts are the use case for models like Microsoft's MAI-Transcribe-2-Streaming, listed with 60 languages and a $0.54 per hour introductory price through end 2026 (October tracker, read 2026-10-06). Sume's transcription is $0.01 per audio minute, or $0.60 per hour, and the one thing it does not do is name who spoke.
Transcribe two tracks and merge
import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
TRACKS = {"Host": "https://example.com/host.wav", "Guest": "https://example.com/guest.wav"}
URL = "https://api.sume.com/v1/stt-1.0/transcribe"
def transcribe(name, audio_url):
job = call(URL, {"audio_url": audio_url, "language_code": "en", "duration_seconds": 600,
"segmentation": {"mode": "sentence"}}, f"two-track-v1-{name}")
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(job["status_url"])
return [dict(s, speaker=name) for s in call(job["result_url"])["result"].get("segments", [])]
rows = sorted((s for n, a in TRACKS.items() for s in transcribe(n, a)), key=lambda s: s["start"])
for s in rows:
print(f"[{s['start']:6.1f}] {s['speaker']}: {s['text'].strip()}")Using it on real files
Replace the placeholder URLs with public https links to your own files; each track must be 10 minutes or shorter, and the duration_seconds value (1 to 600) is the amount reserved for billing, so set it close to the true length. For longer recordings, split first and offset the timestamps by each chunk's start, as in transcribing a 45-minute recording. Segments appear only when segmentation.mode is sentence, which is why the request asks for it.
Naming and ordering
Keep the speaker names in one place and pass them in from your recording sheet, not from file names. Use a small tolerance when sorting: if two segments start at the same second, put the earlier-ending one first so short replies stay in order. Finally, write the merged rows to your transcript format once, from one list, so edits happen in one place and never need to be reconciled between the tracks.
Where it breaks
Crosstalk is where the method shows its limits. If the host's microphone also picks up the guest, the same sentence appears on both tracks, and sorting by time puts a duplicate in the transcript. Record on separate microphones or use headphones, and check the first minute of both tracks before you transcribe an hour. Silence on a track is no problem: a track with nothing to say yields no segments for that stretch.
| Recording | Method | Speaker labels |
|---|---|---|
| One mixed file | Sume STT on the file | None |
| One file per speaker | Transcribe each, merge by start | From the track name |
| Needs automatic diarization | A diarizing provider | Not on Sume STT |
Cost and next steps
Cost is per track: two 10-minute tracks are 20 audio minutes, or $0.20 at $0.01 per minute (pricing, read 2026-10-06). That is the same as one 20-minute mixed file, so labeling by track costs nothing extra beyond the second recording. A one-hour interview would be six 10-minute chunks per track, 120 audio minutes or $1.20, and the chunking guide linked above covers the offsets. Remember that every job reserves the duration_seconds you give it, so a track with a long silent tail is cheaper if you pass its true length. For Apple's podcast transcript rules, which ask for speaker names, see speaker names in a VTT; to make subtitles from the merged rows, use STT to SRT in Python. The envelope fields come from Jobs and results.
Sources
Related posts
More in Developers
- Laravel queued job for an AI video API: delay, redispatch, poll
A Laravel ShouldQueue job reads a Sume video job once and redispatches itself with ->delay() from next_poll_after_seconds, so no worker sleeps during a render.
- Let QA override the image model per request, with an allowlist
An internal render endpoint that honors an X-Image-Model header only for allowlisted Sume ids, so QA can test gpt-image-2.5 before the config flips. Python.
- List your transcription jobs: GET /v1/jobs by type, status, cursor
GET /v1/jobs?type=speech_to_text pages 100 jobs at a time, newest first. Join on each row's idempotency_key, not array position, to rebuild a batch.
- Live webinar captions: Sume has no streaming STT, so use chunks
MAI-Transcribe-2-Streaming is $0.54 an hour. Sume STT is $0.60, async and capped at 10 minutes a job. How to caption a webinar recording in chunks.
Written by Sume