GPT-Live 1 pricing per second vs Sume's per-job audio billing
GPT-Live 1 voice sessions cost $0.05 per minute billed per second. Sume bills TTS per character and STT per audio minute, each reserved at submit.

OpenAI's changelog (Sep 10) says GPT-Live 1 is generally available, with voice sessions at $0.05 per minute billed per second, and backend model and tool usage charged separately. Sume has no live voice session; its audio is two jobs: text to speech billed per character, and speech transcription billed per audio minute, each reserved when you submit.
This page compares billing units only. Sume's rates come from the rate card (TTS $0.0475 per 1,000 characters; STT $0.01 per audio minute), read 2026-09-30.
What is a GPT-Live 1 session billed on?
Time: $0.05 per minute, metered per second while the session runs. The same entry says GPT-Live 1 builds full-duplex voice conversations that can continue while a backend model or agent handles reasoning and tools, and those backend costs are separate.
What are Sume's audio billing units?
TTS 1.0 is priced per transcript character after Sume margin. Spaces and punctuation count, and the maximum is 20000 characters per request. STT 1.0 is priced per audio minute, with 1 minute reserved if you send no duration and a maximum of 10 minutes. Paid generation reserves the estimated amount when the request is accepted, and successful completion captures it (generation admission).
| GPT-Live 1 session | Sume TTS 1.0 | Sume STT 1.0 | |
|---|---|---|---|
| Unit | Per minute, billed per second | Per character | Per audio minute |
| Listed | $0.05 per minute | $0.0475 per 1,000 characters | $0.01 per audio minute |
| Cap | Not stated | 20000 characters | 10 minutes |
| When charged | While the session runs | Reserved at submit | Reserved at submit |
Can I build a voice agent on Sume?
You can chain the jobs: transcribe a recorded turn, produce a reply with your own logic, then synthesize it. That is turn-based, with job latency between turns, not full-duplex audio. For narration use see TTS for video narration.
How should I estimate cost for each?
For a session, multiply minutes by the listed rate and add backend model usage. For Sume jobs, count the characters you will speak and the minutes you will transcribe; both are known before submit. If balance cannot cover the reservation, submit fails with 402 insufficient_credits before provider work starts.
Sources
Related posts
More in Pricing
- Hedra API pricing: estimate endpoint vs Sume dry_run
Hedra says you can price the exact request with an estimate endpoint. Sume's hosted MCP has dry_run and max_spend_usd gates that preview cost before a job.
- HeyGen API concurrency limit: 1.5x burst vs Sume queued jobs
HeyGen Enterprise burst adds up to 50 slots billed at 1.5x. Sume queues jobs above plan concurrency and returns 429 only when the queue is full.
- HeyGen voice clone API: 20+ minutes, paid slot, no Sume clone
HeyGen's professional voice clone needs 1-10 recordings totaling 20+ minutes and a paid slot. Sume's TTS tool selects voices and has no clone upload.
- Lyria 3 Clip Preview 30-second price vs Sume Music fixed price
Google lists Lyria 3 Clip Preview (30s) at a per-song price. Sume Music charges one fixed price per accepted generation, whatever length the prompt asks for.
Written by Sume