Transcribe an interview with timestamps via API (no speaker labels)

Send interview audio to Sume STT for text, word start and end times and sentence segments. It returns no speaker names. 25 minutes costs 28 cents.

5 min readSume
All posts

To transcribe a recorded interview with timestamps through Sume, send the audio to POST /v1/stt-1.0/transcribe as a public HTTPS audio_url and read back text plus words[], where every word has start and end in seconds from the audio start. Add segmentation: {"mode": "sentence"} and you also get sentence segments with no gaps between them. The result does not label who is talking, so an interview needs one extra step on your side to attribute lines.

The price is $0.01 per audio minute, so a 25 minute interview is $0.25 of speech-to-text, plus a cent for each audio slice you cut.

The 10 minute job limit

One STT job takes at most 10 minutes of audio, and duration_seconds accepts 1 to 600. A 25 minute interview is therefore three jobs: two of 600 seconds and one of 300. If the interview is a video, audio detach takes a hosted video of up to 1800 seconds and returns a WAV with a range of { start, end }, output capped at 900 seconds. Detach each 10 minute range as its own job (0 to 600, 600 to 1200, 1200 to 1500) with sample_rate 16000 and channels mono, then send each file to STT. The same approach is in the 12 minute interview walkthrough.

25 minute interview video on Sume, arithmetic from documented rates read 2026-10-07
StepCountRateCost
Detach audio, one range per 10 minutes3 jobs$0.01 per job$0.03
Speech-to-text25 audio minutes$0.01 per minute$0.25
Total$0.28

Timestamps across chunks

Word times are measured from the start of whatever audio you submitted. If you send the second 10 minute slice, its first word starts near 0, not near 600. Keep a list of slice start offsets and add the offset to every start and end before you merge. This is the single most common mistake in a chunked transcript, and the fix is one addition per word.

If you need sentence-level rows for an editor or a quote finder, request the sentence segmentation. Segments are gapless, so each end equals the next start, which makes them easy to turn into a table of timestamp, text pairs.

Putting names on the lines

Sume's STT result documents text, language fields, words and optional segments. It documents no speaker field, and the request schema says provider knobs such as diarization are fixed server-side. That means you attribute lines yourself. Two practical options:

The first is turn-based. You know who asks the questions, so submit separate audio per speaker if you recorded separate tracks, and merge by start. The second is manual: export the segments with timestamps, and have an editor mark speakers once. Neither is automatic, so say so in any workflow you promise to a client.

A third habit helps with checking: keep the audio slice file next to its transcript. When a reader disputes a quote, you can replay the exact slice, because each slice's offset tells you where it sits in the full recording.

  • Separate mic tracks: transcribe each track, tag by file, sort all words by start time.
  • Single mixed track: use sentence segments as rows and label speakers by hand.
  • Quotes for a story: keep the original timestamp so a reader can find the moment in the recording.

Language and accuracy

Leave language_code out and the job auto-detects the language; set it (for example en or ko) when you know it. The completed result carries language_code and, when available, language_probability, so you can flag an interview that detected as something unexpected. Sume publishes no accuracy figure for interviews, so spot check names and numbers against the audio before you publish, and use the word times to jump straight to each questionable line.

Request

Replace the URL with the detached WAV for one range. Keep the idempotency key stable for a retry, and change it when the body changes.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: interview-part-01" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/artf_demo/part-01.wav",
    "duration_seconds": 600,
    "segmentation": { "mode": "sentence" }
  }'

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume