300 UGC videos of 45 seconds: the transcript bill

300 creator videos of 45 seconds are 3.75 hours: $2.03 at MAI-Transcribe-2-Streaming's intro price, $2.25 on Sume STT. Detach and word-timing notes.

3 min readSume
All posts

Transcribing 300 user-generated videos of 45 seconds each is 225 audio minutes, or 3.75 hours, which costs $2.03 on MAI-Transcribe-2-Streaming ($2.025 before rounding) and $2.25 on Sume STT 1.0.

Short clips are a different shape from calls: each is one job, and the useful output is word timings for captions.

The arithmetic

MAI-Transcribe-2-Streaming is priced per hour, Sume STT 1.0 per audio minute, so the sums convert minutes to hours first. The Sume rate is the public rate on the Sume docs; confirm the live number in GET /v1/catalog.

Per minute, MAI works out to $0.009 ($0.54 / 60) against Sume's $0.01, so Sume is about 11 percent higher. On this run the difference is $0.225. Microsoft's rate is an introductory price that its page says runs through the end of the year, and the page does not say what the price becomes afterward, so a budget past December should not assume it.

Transcription cost for 225.0 audio minutes, rates as of 2026-10-08
OptionRateArithmeticResult
MAI-Transcribe-2-Streaming$0.54 per hour of audio, intro price through the end of 2026225.0 min = 3.750 h x $0.54$2.025
Sume STT 1.0$0.01 per audio minute225.0 min x $0.01$2.25
Sume STT 1.0 job count10 minutes (600 s) per job300 x ceil(0.75 / 10) = 300 jobs300 jobs

Words for captions

Sume STT always returns word timings, so the same job that gives you text gives you the data a caption renderer needs. If the final goal is burned-in captions, the standalone caption job ($0.20 reserved for a video of up to 60 seconds) runs speech recognition inside the job, so do not run both.

Audio detach works on a Sume-hosted video; for STT itself, a public HTTPS audio_url is enough.

  • Check each clip has audible speech first; a silent clip fails caption jobs with caption_no_speech.
  • Submit the 300 clips with a concurrency limit and one idempotency key per clip.

Limits that shape a Sume run

Sume does not list MAI-Transcribe-2-Streaming, so the MAI row is a price reference. Sume STT 1.0 takes a public HTTPS audio_url and returns the transcript with word timings; there is no flag for timings because they always come back. Sentence segments are available when you send segmentation.

  • duration_seconds is an optional hint from 1 to 600; leave it out and Sume reserves one minute, so send it for a long file.
  • language_code is a hint such as en or ko; leave it out for auto-detect.
  • diarize and tag_audio_events are fixed server-side, so do not send them; there are no speaker-label controls.
  • The MAI price is an introductory rate that Microsoft says runs through the end of the year.

Running it

Submit with stt_create and an idempotency_key, wait with jobs_wait, then read jobs_result for the text and words. Cut recordings longer than 10 minutes at a pause first; audio detach pulls the track out of a video and Timeline audio split slices a Sume-hosted file into ranges.

The figures are list arithmetic as of 2026-10-08. A one-time check of your own audio is worth more than a benchmark, so run one representative file through each option before you commit a budget.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume