Clean vs verbatim transcripts: MAI-Transcribe-2 style vs Sume captions
MAI-Transcribe-2 batch has transcribeStyle clean or verbatim. Sume STT has no style flag; for polished captions supply script_text and keep your wording.

MAI-Transcribe-2 batch exposes a transcribeStyle setting with two values, clean and verbatim, so you choose whether fillers and false starts survive in the text. Sume STT 1.0 has no equivalent field: its request takes audio_url, language_code, duration_seconds, segmentation and metadata. If you want captions that read cleanly, the Sume route is to give the caption job your own script_text and let it align that wording to the speech.
What each side lets you control
The Microsoft batch page lists formats WAV, MP3 and FLAC, diarization, timestamps of word, segment or none, a phraseList for biasing, and the transcribeStyle choice. Style is the decision that matters for subtitles: verbatim is right for legal or research records, clean is right for viewers.
On Sume, the speech-to-text model returns text, language_code and always-on words[], and the knobs that exist in the provider, such as diarization and audio event tags, are fixed server-side; sending them is rejected.
| Need | MAI-Transcribe-2 batch | Sume |
|---|---|---|
| Keep fillers | transcribeStyle: verbatim | No flag; check your sample |
| Drop fillers | transcribeStyle: clean | Supply script_text to the caption job |
| Word timing | timestamps: word | words[] always returned |
| Own wording on screen | Edit the output yourself | script_text aligned onto speech |
| Brand-name bias | phraseList hints | script_text carries exact spelling |
The Sume route for clean captions
The video captions endpoint accepts script_text, which is aligned onto what STT hears, so the on-screen wording is yours and the timing is the speaker's. It is mutually exclusive with words, cues and segments, which skip STT entirely.
If your wording drifts too far from the speech, the job fails with script_alignment_mismatch or script_alignment_failed, and the documented next action is to simplify the script or omit it. That is the honest limit of the approach: it is for tidying, not rewriting.
Which to choose
Choose verbatim output when the text itself is the record. Choose a script when the text is a product. For an interview where you want the speaker's actual words minus filler, edit the transcript once and send it as script_text; for a pre-written voiceover, send the original script. Do not use this to translate: the script has to match the spoken language.
A practical split by use
The right setting follows from what happens to the text next.
- Court, medical or research record: verbatim, and keep the timestamps next to it.
- Subtitles for viewers: clean text, with the speaker's timing preserved.
- Dubbing source: clean text, since filler words become awkward extra syllables in a target language.
- Search and indexing: either works; verbatim can help when people search for exact phrasing.
Sources
Related posts
More in Comparisons
- Clef or Clef-flash in front of a Sume agent: which tier to use
Cloudflare lists a median 38.8 ms for Clef-flash and 209.3 ms for Clef. Use the fast tier to gate runs and the larger one for the choices that cost real money.
- Cloudflare Workers AI TTS: Aura-1, Aura-2, MeloTTS vs Sume TTS
Cloudflare lists Aura-1 at $0.015 and Aura-2 at $0.030 per 1,000 characters and MeloTTS at $0.0002 per minute. Sume TTS 1.0 is $0.0475. Costs at three volumes.
- Cloudflare Workers AI image prices per tile vs Sume per image
Cloudflare prices images per 512x512 tile: $0.0000528 for flux-1-schnell, $0.0070 lucid-origin. A 1024x1024 is four tiles. Sume image models start at $0.025.
- Cloudflare Workers AI Whisper vs Sume STT: cost of 1,000 hours
Cloudflare lists Whisper at $0.0005 per audio minute, so 1,000 hours is $30. Sume STT 1.0 is $0.01 per minute, or $600. What the 20x gap does and does not mean.
Written by Sume