Speech to text API pricing: cost per minute and per hour
Speech to text API pricing is a rate per minute of audio. Sume's STT rate, what it includes, what an hour costs, and the per-request limits.

Speech to text API pricing is a rate per minute of audio, so an hour of recordings costs 60 times the per-minute rate. Sume's speech-to-text API (STT 1.0) charges $0.01 per audio minute, plus a 5.5% agent fee by default, which puts one hour of audio at $0.60 before the fee.
The rate is read from the code behind API pricing. What a request takes and returns comes from the STT schema in the Sume API reference, the OpenAPI document behind the API reference docs; billing behavior comes from Usage and Generation admission. All read on 2026-09-29.
How much does speech to text cost at my volume?
One request takes at most 10 minutes of audio, so longer recordings go in several requests at the same per-minute rate. The request counts below assume each part is 10 minutes or less; How to transcribe long audio files covers splitting and re-basing timestamps.
| Audio | Minutes | Requests | STT price |
|---|---|---|---|
| 1 minute | 1 | 1 | $0.01 |
| 10 minutes (one maximum request) | 10 | 1 | $0.10 |
| 1 hour | 60 | 6 | $0.60 |
| 10 hours | 600 | 60 | $6.00 |
| 100 hours | 6,000 | 600 | $60.00 |
What does the per-minute price include?
A completed job returns the transcript text, language fields when available, and word-level timings in words[], in seconds from the audio start. The API reference says word timings are always returned, with no flag to turn them on and no extra line for them on the rate card; in current code a word's start and end are optional, so check for them. Speech to text API with word timestamps covers the request.
- Sentences: send
segmentation: { "mode": "sentence" }to also get sentence segments built from the word timings. - Language:
language_codeis optional; omit it for auto-detect. - Speakers: not included.
diarizeis fixed server-side and a request that sends it is refused, so the transcript has no speaker labels. How to transcribe an interview covers labeling speakers yourself.
How is a request billed?
When Sume accepts a request, it reserves the estimated cost from your balance; a successful job captures it, and a failed job releases or refunds the reservation. The estimate comes from duration_seconds, an optional field from 1 to 600. Omit it and Sume reserves for 1 minute.
- Send
duration_secondswhen you know the length, so the reservation matches the file. GET /v1/usage?job_id=…sums what one job cost;debited_usdis what the wallet deducted.- A submit that can't reserve the estimate fails with
402 insufficient_creditsbefore any work starts.
What are the limits?
- Input:
audio_url, a public HTTPS URL to the audio. There is no file upload in the request. - Length: at most 10 minutes of audio per request.
- Waiting:
mode: syncwaits at most 30 seconds; after that, poll the job or use a webhook. - Results:
/resultreturns409 job_not_completeduntilresult_readyis true.
Sources
Related posts
More in Pricing
- Synthesia AI dubbing costs: minutes per plan and lip sync
Synthesia AI dubbing costs plan credits: up to 48 minutes a month on Starter ($29) and 140 on Creator ($89), and lip sync uses twice the credits.
- Text to speech commercial use: what the terms allow
Whether you can use text to speech audio commercially is set by the tool's terms and plan. What Sume's terms say, and what to check for a cloned voice.
- Unlimited AI video generator subscription: does one exist?
An unlimited AI video plan still limits something. Sume has no unlimited or lifetime plan: paid plans include some usage, then bill per clip.
- How Sume pricing works: plans, one wallet, published model rates
Sume plans set access and concurrency. Usage draws from one prepaid wallet at each model's published USD rate, for generation, the Agent, Formats, and the API.
Written by Sume