AssemblyAI diarization +$0.02/hour: a 12-hour panel vs Sume STT

AssemblyAI prices standard diarization at +$0.02 an hour on a $0.21 base. Twelve hours of panel audio is $2.76; Sume STT is $7.20 with no diarization field.

3 min readSume
All posts

For a 12-hour recording of panel discussions, AssemblyAI's published prices come to $2.76 with Universal-3.5 Pro plus standard speaker diarization ($0.21 + $0.02 = $0.23 per hour, times 12). Sume STT is $0.01 per audio minute, which is $7.20 for 12 hours. Sume's request body does not let you turn diarization on or off; its docs describe that setting as fixed server-side.

Prices were read on AssemblyAI's page and Sume's catalog on 2026-10-09.

The prices

AssemblyAI lists two diarization lines for asynchronous audio: standard at +$0.02 per hour and experimental at +$0.065 per hour. The base models are Universal-3.5 Pro at $0.21 per hour and Universal-2 at $0.15 per hour.

12 hours of panel audio, rates read 2026-10-09
OptionPer hour12 hours
Universal-2 only$0.15$1.80
Universal-3.5 Pro only$0.21$2.52
Universal-3.5 Pro plus standard diarization$0.23$2.76
Universal-3.5 Pro plus experimental diarization$0.275$3.30
Sume STT 1.0$0.60$7.20

What Sume STT gives you instead

Sume STT returns the text, language fields and word-level timings. You can also request sentence segmentation. The request takes a public HTTPS audio_url, an optional language_code, and an optional duration_seconds up to 600. Provider knobs such as diarize and tag_audio_events are fixed server-side, so you cannot add a speaker-label option to a request.

If you need to know who is speaking in a panel, that is a gap. A common workaround is to give each speaker a separate audio channel or file when recording, then transcribe each file on its own. With Sume that is a job per speaker file, at $0.01 per minute each.

Deciding between them

For multi-speaker audio where labels matter, the AssemblyAI row is cheaper and does the job. For single-speaker audio, such as a lecture or a voice memo, speaker labels add nothing and the price difference comes down to the base rate. Test on a short stretch of your real recording, since overlapping speech is where diarization differs most between services.

  • Panels and interviews: pay for diarization.
  • Lectures and narration: skip it.
  • Sume pipelines that already use Sume media: keep STT there and record speakers separately.

Per-speaker files on Sume, priced

Suppose the panel has four speakers and you recorded each on a separate track. Transcribing four 12-hour tracks on Sume STT is 4 x $7.20 = $28.80, because each track is billed by its own audio length even though most of it is silence. That is far above the AssemblyAI row. Per-speaker tracks only make sense if the speakers are on screen for short stretches, or if you cut the silence out first.

You can pull a track from a hosted video with audio detach at $0.01 per job, but detach accepts a source of at most 1,800 seconds and outputs at most 900 seconds, so a 12-hour recording does not fit that tool at all. For an event this long the vendor with a diarization option is the practical choice, and Sume is better kept for the clips you cut afterward.

Whichever you pick, spot-check the first ten minutes for speaker mix-ups before you run all 12 hours.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume