AssemblyAI Voice Agent API at $4.50 an hour all-in
AssemblyAI's Voice Agent API is $4.50 an hour ($0.075 a minute) and includes STT, LLM, TTS and orchestration. Streaming bills on session time, not audio time.

AssemblyAI's Voice Agent API is listed at $4.50 an hour, which is $0.075 a minute, and the price includes speech-to-text, the LLM, text-to-speech and orchestration. A 10,000-minute month at that rate is $750. One detail changes your budget: streaming is billed on session duration, not on how much audio was spoken.
What the page lists
From AssemblyAI's pricing page, read 2026-10-05, with Cartesia's agent price from its own page for context.
| Item | Listed price | Per minute or hour |
|---|---|---|
| AssemblyAI Voice Agent API | $4.50 per hour, includes STT, LLM, TTS and orchestration | $0.075 per minute |
| AssemblyAI Universal-Streaming (STT only) | $0.15 per hour | $0.0025 per minute |
| AssemblyAI Universal-3.6 Pro Realtime (STT only) | $0.45 per hour | $0.0075 per minute |
| Cartesia agents | $0.06 per minute | $3.60 per hour |
Session time, not audio time
The page says streaming is billed on session duration. A session that sits open while the caller thinks, or while the line is silent, costs the same per minute as one full of speech. For a support line with long pauses, the effective price per minute of actual speech is higher than $0.075. Model it with your own call logs: billable minutes equal call length, not talk time.
Hold times, voicemail greetings and ringing that your system keeps open all count if the session is open. Closing sessions promptly is the cheapest optimisation.
What a month costs
Arithmetic on $0.075 a minute. It assumes every minute of call time is session time.
| Call minutes per month | Hours | Cost |
|---|---|---|
| 1,000 | 16.7 | $75 |
| 10,000 | 166.7 | $750 |
| 100,000 | 1,666.7 | $7,500 |
All-in versus assembling the parts
A bundled price is simple but hides the split. Building your own agent means paying separately for streaming STT, an LLM, TTS and orchestration, as with OpenAI's gpt-realtime-2.1, whose page lists per-token audio prices ($32 input and $64 output per million audio tokens). The bundled number is easy to forecast; the unbundled one is easier to tune. The sources for this post give no latency or quality figures for the Voice Agent API, so test a call before you commit.
Sume does not offer a voice-agent product. Its audio services are asynchronous jobs: Cartesia Sonic text-to-speech, speech-to-text on video clips, music and lip-sync. They suit produced content, such as an ad read or a talking-head clip, not live calls.
Sources
Related posts
More in Pricing
- Audio description for 200 videos at 400 characters: TTS cost
Voicing a 400-character description track for each of 200 videos is 80,000 characters: $1.20 on MAI Flash, $1.76 on MAI-Voice-2.1, $3.80 on Sume TTS.
- Why your Sume balance dropped for a job that is still queued
Sume reserves the estimated cost when it accepts a job, not when it starts. Six queued 10-second Wan 720p jobs hold $7.50 at once. How to read it.
- Black Friday hook tests: 20 ten-second variants on seven Sume rates
What 20 ten-second Black Friday ad variants cost on Wan 3.0, Gemini Omni Flash 1.1 and MiniMax H3 Max at Sume's per-second rates, from $12.50 to $50.00.
- Budget for 12 UGC ad variants: draft on Mini, finish on Seedance 2.5
What 12 vertical UGC ad variants cost on Sume when you draft on seedance-2-mini at 480p and finish the winners on seedance-2.5 at 720p. Full arithmetic inside.
Written by Sume