Transcribe an interview with timestamps via API (no speaker labels)
Send interview audio to Sume STT for text, word start and end times and sentence segments. It returns no speaker names. 25 minutes costs 28 cents.

To transcribe a recorded interview with timestamps through Sume, send the audio to POST /v1/stt-1.0/transcribe as a public HTTPS audio_url and read back text plus words[], where every word has start and end in seconds from the audio start. Add segmentation: {"mode": "sentence"} and you also get sentence segments with no gaps between them. The result does not label who is talking, so an interview needs one extra step on your side to attribute lines.
The price is $0.01 per audio minute, so a 25 minute interview is $0.25 of speech-to-text, plus a cent for each audio slice you cut.
The 10 minute job limit
One STT job takes at most 10 minutes of audio, and duration_seconds accepts 1 to 600. A 25 minute interview is therefore three jobs: two of 600 seconds and one of 300. If the interview is a video, audio detach takes a hosted video of up to 1800 seconds and returns a WAV with a range of { start, end }, output capped at 900 seconds. Detach each 10 minute range as its own job (0 to 600, 600 to 1200, 1200 to 1500) with sample_rate 16000 and channels mono, then send each file to STT. The same approach is in the 12 minute interview walkthrough.
| Step | Count | Rate | Cost |
|---|---|---|---|
| Detach audio, one range per 10 minutes | 3 jobs | $0.01 per job | $0.03 |
| Speech-to-text | 25 audio minutes | $0.01 per minute | $0.25 |
| Total | $0.28 |
Timestamps across chunks
Word times are measured from the start of whatever audio you submitted. If you send the second 10 minute slice, its first word starts near 0, not near 600. Keep a list of slice start offsets and add the offset to every start and end before you merge. This is the single most common mistake in a chunked transcript, and the fix is one addition per word.
If you need sentence-level rows for an editor or a quote finder, request the sentence segmentation. Segments are gapless, so each end equals the next start, which makes them easy to turn into a table of timestamp, text pairs.
Putting names on the lines
Sume's STT result documents text, language fields, words and optional segments. It documents no speaker field, and the request schema says provider knobs such as diarization are fixed server-side. That means you attribute lines yourself. Two practical options:
The first is turn-based. You know who asks the questions, so submit separate audio per speaker if you recorded separate tracks, and merge by start. The second is manual: export the segments with timestamps, and have an editor mark speakers once. Neither is automatic, so say so in any workflow you promise to a client.
A third habit helps with checking: keep the audio slice file next to its transcript. When a reader disputes a quote, you can replay the exact slice, because each slice's offset tells you where it sits in the full recording.
- Separate mic tracks: transcribe each track, tag by file, sort all words by start time.
- Single mixed track: use sentence segments as rows and label speakers by hand.
- Quotes for a story: keep the original timestamp so a reader can find the moment in the recording.
Language and accuracy
Leave language_code out and the job auto-detects the language; set it (for example en or ko) when you know it. The completed result carries language_code and, when available, language_probability, so you can flag an interview that detected as something unexpected. Sume publishes no accuracy figure for interviews, so spot check names and numbers against the audio before you publish, and use the word times to jump straight to each questionable line.
Request
Replace the URL with the detached WAV for one range. Keep the idempotency key stable for a retry, and change it when the body changes.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: interview-part-01" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/part-01.wav",
"duration_seconds": 600,
"segmentation": { "mode": "sentence" }
}'Sources
Related posts
More in Use cases
- Translate a coffee bag label to Japanese with Ideogram 4.5
Translate a packaging label into Japanese with ideogram/ideogram-v4.5: give the exact string, one language per call, check characters. $0.075 per bag at medium.
- Turn summer footage into autumn with an edit prompt (Omni Flash)
Change the season of a clip you already shot with a video_url edit on Gemini Omni Flash 1.1 at Sume: prompt wording, what to verify, and a 6-second price table.
- How do I make a two-voice dialogue audio for language learners?
Make a two-speaker practice dialogue from one TTS job per line, joined gapless by a $0.01 concat. A ten-line dialogue is about 11 cents on Sume.
- UGC-style avatar video: a 9:16 casual scene prompt that works on Sume
How to ask Sume Avatar 1.0 for a phone-filmed look: 9:16, a casual background prompt, silence beats, product image. Request body and what it costs per tier.
Written by Sume