Clean vs verbatim transcripts: MAI-Transcribe-2 style vs Sume captions

MAI-Transcribe-2 batch has transcribeStyle clean or verbatim. Sume STT has no style flag; for polished captions supply script_text and keep your wording.

4 min readSume
All posts

MAI-Transcribe-2 batch exposes a transcribeStyle setting with two values, clean and verbatim, so you choose whether fillers and false starts survive in the text. Sume STT 1.0 has no equivalent field: its request takes audio_url, language_code, duration_seconds, segmentation and metadata. If you want captions that read cleanly, the Sume route is to give the caption job your own script_text and let it align that wording to the speech.

What each side lets you control

The Microsoft batch page lists formats WAV, MP3 and FLAC, diarization, timestamps of word, segment or none, a phraseList for biasing, and the transcribeStyle choice. Style is the decision that matters for subtitles: verbatim is right for legal or research records, clean is right for viewers.

On Sume, the speech-to-text model returns text, language_code and always-on words[], and the knobs that exist in the provider, such as diarization and audio event tags, are fixed server-side; sending them is rejected.

Transcript style controls (Microsoft batch page read 2026-10-05; Sume docs and repo)
NeedMAI-Transcribe-2 batchSume
Keep fillerstranscribeStyle: verbatimNo flag; check your sample
Drop fillerstranscribeStyle: cleanSupply script_text to the caption job
Word timingtimestamps: wordwords[] always returned
Own wording on screenEdit the output yourselfscript_text aligned onto speech
Brand-name biasphraseList hintsscript_text carries exact spelling

The Sume route for clean captions

The video captions endpoint accepts script_text, which is aligned onto what STT hears, so the on-screen wording is yours and the timing is the speaker's. It is mutually exclusive with words, cues and segments, which skip STT entirely.

If your wording drifts too far from the speech, the job fails with script_alignment_mismatch or script_alignment_failed, and the documented next action is to simplify the script or omit it. That is the honest limit of the approach: it is for tidying, not rewriting.

Which to choose

Choose verbatim output when the text itself is the record. Choose a script when the text is a product. For an interview where you want the speaker's actual words minus filler, edit the transcript once and send it as script_text; for a pre-written voiceover, send the original script. Do not use this to translate: the script has to match the spoken language.

A practical split by use

The right setting follows from what happens to the text next.

  • Court, medical or research record: verbatim, and keep the timestamps next to it.
  • Subtitles for viewers: clean text, with the speaker's timing preserved.
  • Dubbing source: clean text, since filler words become awkward extra syllables in a target language.
  • Search and indexing: either works; verbatim can help when people search for exact phrasing.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume