Speech-to-text price per audio hour: xAI, OpenAI, Sume
Per audio hour xAI lists $0.10, OpenAI mini $0.18, OpenAI 4o $0.36 and Sume video_inspect transcripts $0.60. Dated 2026-10-01, with honest limits.

Per audio hour, xAI lists speech to text at $0.10 over REST, OpenAI's gpt-4o-mini-transcribe at $0.18 and gpt-4o-transcribe at $0.36. Sume's video_inspect transcript costs $0.01 per audio minute, or $0.60 per hour, which is the highest of the five. Sume's value is that the transcript comes with the video work it belongs to, not that it is cheapest.
Rates per hour
Per-minute rates times 60. Vendor rows are public list prices.
| Option | Listed price | Per hour | Per 1,000 hours |
|---|---|---|---|
| xAI speech to text, REST | $0.10 per hour | $0.10 | $100.00 |
| xAI speech to text, streaming | $0.20 per hour | $0.20 | $200.00 |
| OpenAI gpt-4o-mini-transcribe | $0.003 per minute | $0.18 | $180.00 |
| OpenAI gpt-4o-transcribe | $0.006 per minute | $0.36 | $360.00 |
| Sume video_inspect transcript | $0.01 per audio minute | $0.60 | $600.00 |
Where Sume fits
video_inspect takes a video, returns a transcript along with frames and timing, and bills $0.01 per audio minute. The default reserve is one minute and the maximum duration hint is 600 seconds, so it is built for clips, not for a day of recordings.
If all you need is a transcript of long audio at the lowest cost, the vendor rows above are cheaper. If you need to inspect a clip you generated or received and then act on it, one call and one wallet is the point.
- Long audio: use a dedicated speech-to-text model.
- Short clips with visual context: use video_inspect.
- Always compare accuracy on your own audio, not just price.
A concrete month
Take 200 hours of recordings in a month. xAI REST is $20, OpenAI mini $36, OpenAI 4o $72 and Sume video_inspect $120. The spread is $100 a month, which is a real cost but a small one beside most production budgets.
If those same 200 hours were 12,000 one-minute clips that you also want to inspect visually, the single-call value of a clip-level tool is likelier to outweigh the difference.
Limits
Accuracy, language coverage, diarization and latency differ among these services and are not compared here. The tables show the model amount. A workspace Agent Fee line can be added on top, so treat the job's own usage cost as the truth and check the live catalog before you budget; limits and rows change.
Sources
Related posts
More in Comparisons
- Speechify API vs Sume: speech marks and streaming vs video jobs
Speechify's API streams speech with word-level marks and voice cloning. Sume documents no standalone speech endpoint, but accepts audio for talking-video jobs.
- StepAudio 3 Gen: voice, SFX and music in one clip vs Sume jobs
StepFun's stepaudio-3-gen-preview makes voice, effects, ambience and music in one audio output. Sume uses separate music, speech and Timeline mix steps.
- StepAudio 3 Music from a dry vocal or reference audio vs Sume
StepAudio 3 Music accepts lyrics, vocals or reference audio, even scoring a dry vocal. Sume Music takes a text prompt and one optional image, no audio.
- Synthesia burned-in captions on dubbed videos, and Sume captions
Synthesia added one-click burned-in captions to dubbed videos on 9/30/2026. Here is what that means, and how to burn captions onto a finished video URL on Sume.
Written by Sume