AssemblyAI Universal-3.5 Pro $0.21 per hour vs Sume STT
AssemblyAI lists Universal-3.5 Pro async at $0.21 an hour plus $0.02 for speaker labels. Sume STT is $0.60 an hour with no add-on line. What to compare.

AssemblyAI's pricing page lists Universal-3.5 Pro for asynchronous transcription at $0.21 per hour, Universal-2 at $0.15 per hour, and a speaker-diarization add-on at $0.02 per hour. Sume STT 1.0 is $0.01 per audio minute ($0.60 per hour). At list price AssemblyAI is roughly a third of Sume's rate.
What the page lists
AssemblyAI also lists a realtime model and $50 in free credits. The table covers only the async lines.
| Item | AssemblyAI | Sume STT 1.0 |
|---|---|---|
| Newest async model | Universal-3.5 Pro $0.21 per hour | $0.60 per hour |
| Older async model | Universal-2 $0.15 per hour | Single public rate |
| Speaker labels | +$0.02 per hour | Fixed server-side; no separate line in the docs |
| Free credit | $50 | Not claimed |
Sume job example
Sume's transcript comes back as text plus words[] with start and end times.
import os, requests
r = requests.post(
"https://api.sume.com/v1/stt-1.0/transcribe",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "stt-demo-001"},
json={"audio_url": "https://example.com/call.mp3",
"language_code": "en",
"duration_seconds": 300},
)
print(r.status_code, r.json())Where this leaves you
If transcription cost is your largest line and you do not need other media steps, AssemblyAI's published rate is lower. If the transcript is an input to captions, dubbing or timeline audio, staying inside one job model can save integration work. Measure on your own audio before deciding.
Worked example
Speaker labels change the total, so here is the full line.
- 100 hours of Universal-3.5 Pro: $21.00; with speaker labels at +$0.02 per hour: $23.00.
- 100 hours of Universal-2: $15.00.
- Sume STT for 100 hours: $60.00.
Checklist before you commit
Sume's docs say diarization and event tags are fixed server-side, so there is no separate switch to price. Check a sample result to see how speakers appear in your output.
- Decide whether you need labelled speakers at all.
- Check AssemblyAI's streaming billing if you will also stream: it is per session duration.
- Compare accuracy on noisy audio from your own use case.
Sources
Related posts
More in Comparisons
- Asset Studio 1-Click A/B Testing vs a Sume bulk run of hook variants
Google's Asset Studio adds Gemini Omni video and 1-Click A/B Testing. If your Q4 test spans channels, queue the hook variants as a Sume bulk run instead.
- Azure personal voice consent: name and company must match audio
Azure personal voice requires a recorded consent statement whose talent name and company match the audio. The fields, formats and what Sume collects instead.
- Bannerbear 60 POSTs per 10 seconds vs Sume's per-minute budgets
Bannerbear allows 60 POST requests per 10-second window. Sume budgets writes and reads per minute by plan: 120 to 1200 writes, reads 40 times higher.
- Bannerbear video tools mapped to Sume endpoints, gap by gap
Bannerbear lists eleven video tools. Sume covers trim, crop, concat, captions and color via filters; picture-in-picture and GIF previews are not documented.
Written by Sume