MiniMax speech-to-text $0.38 per hour vs Sume STT $0.01 per minute

MiniMax lists speech-to-text at $0.38 per hour. Sume's STT is $0.01 per audio minute ($0.60 per hour), reached through video inspect on a hosted clip.

4 min readSume
All posts

MiniMax lists its speech-to-text endpoint at $0.38 per hour on its pay-as-you-go page. Sume's transcript rate is $0.01 per audio minute, which is $0.60 per hour by arithmetic. On rate alone MiniMax is lower; the Sume docs describe a narrower route, because transcription there is an option on video inspect for one clip already hosted on media.sume.com.

What does each page list?

The MiniMax Audio section lists one ASR row: Speech-to-Text, described as audio-to-text transcription with streaming, speaker diarization and subtitle (srt/vtt) export, at $0.38 / hour. The Sume docs give the public STT rate on the video inspect page as $0.01 per audio minute and say to confirm the live value in GET /v1/catalog.

Speech-to-text as listed by MiniMax (read 2026-10-03) and in the Sume video inspect docs
ItemMiniMaxSume
Listed rate$0.38 / hour$0.01 per audio minute
Same rate per hour (arithmetic)$0.38$0.60
Cost of 10 hours of audio (arithmetic)$3.80$6.00
Features named on the pageStreaming, speaker diarization, srt/vtt exportWord list, optional sentence segments, language_code hint

What is the Sume route, exactly?

Sume's transcript is not a standalone upload endpoint in the pages I read. POST /v1/video-inspect reads one clip from your workspace's media.sume.com storage, and transcribe: true runs Sume STT 1.0 on its audio. The docs give the source cap as 1800 seconds. A silent clip returns inspect_source_has_no_audio, so check probe.has_audio first.

The response carries transcript with text, words[], optional sentence segments[] and an audio_url. Setting segmentation.mode to sentence returns gapless sentence segments shaped like caption lines.

How is the Sume transcript billed?

Every inspect reserves its compute ceiling and is captured at its own container seconds, never above the hold, and the transcript adds its per-minute rate to that reservation. If you omit duration_seconds the reservation assumes one minute; the documented maximum hint is 600 seconds. So a transcript request has a small compute component on top of the $0.01 per minute, which the MiniMax per-hour figure does not have to match.

Which should you pick?

Choose by input and output, not by the hourly rate alone. If you have raw audio files, need diarization, or need streaming, the MiniMax page lists those and prices them lower per hour. If the audio is already inside a video you generated or imported on Sume, and you want words, sentence segments and stills from the same call, the Sume route avoids moving the file elsewhere.

Whichever you use, run a short test on your own audio and measure word errors yourself: neither vendor page publishes an accuracy figure that this comparison can rely on.

What else is on the MiniMax audio page?

The same Audio section lists text-to-speech by character, for example speech-2.8-hd at $100 per million characters and speech-2.8-turbo at $60 per million characters, so speech in and speech out are priced on one page. That is a useful reminder that a transcription rate is rarely the whole audio bill: a pipeline that transcribes, rewrites and re-voices pays for each stage.

Sume splits the same work across surfaces too. Transcripts come from video inspect, audio can be pulled out of a clip with the $0.01 flat audio-detach route, and captions are a separate caption job. Price each stage you actually use, then add them.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume