Sume STT returns 400 for diarize: a two-speaker fix at $0.60 an hour

Sume STT 1.0 fixes diarize and tag_audio_events server-side and rejects them with 400. For two speakers, transcribe each track and merge by word times.

5 min readSume
All posts

If you send diarize or tag_audio_events to POST /v1/stt-1.0/transcribe, Sume returns a 400: STT 1.0 fixes both settings server-side and asks you not to send them. If you need two speakers kept apart, record or export each speaker as a separate file, transcribe each at $0.01 per audio minute, and merge the results by word start time. Two 30-minute tracks cost 60 audio minutes x $0.01 = $0.60.

Why you see the 400

Microsoft's MAI-Transcribe-2 page lists diarisation, timestamps and keyword biasing as built in, so developers coming from it may carry a diarize: true flag into their Sume request. Sume's schema is strict, and the two names are explicitly refused with a message that says the fields are fixed server-side. The documented STT result exposes text, language fields when available, and words[] with { word, start, end } in seconds from the start of the audio. This post does not claim anything about speaker labels beyond that.

The one-job-per-track workflow

Most interview tools can export each microphone separately. Import each file, detach or normalise it to 16 kHz mono, and send one STT job per file. Then merge by sorting the words from both results on start, and label each word with its track name.

Two speakers, 30 minutes each, rates read 2026-10-09
StepJobsBasisCost
STT, speaker A1 (or 3 at 10 minutes each)30 audio minutes x $0.01$0.30
STT, speaker B1 (or 3)30 audio minutes x $0.01$0.30
Detach, only if the files are inside videos2$0.01 per job$0.02
Total with no detach$0.60

Limits that still apply

Three limits shape the plan.

  • STT reserves from audio-minute estimates. Omit duration_seconds and it reserves 1 minute; the schema accepts 1 to 600 seconds and the catalog cap is 10 minutes, so split a long file into ranges of 600 seconds or less.
  • The audio must be reachable at a public HTTPS URL. Detach output is the easy route if the source is a video.
  • Crosstalk is where this approach shines and fails: if both microphones pick up both voices, the words repeat in two tracks. Mute the other speaker in each export, or keep the louder transcript for overlapping words.

Merging the words

Each result's words[] times start at zero for that file, so the two tracks must start together. Export both microphones from the same session at the same sample rate, and keep any trim identical on both. Then the merge is a sort on start with a track label per word. If one export had a head trim of 1.5 seconds, add 1.5 to that track's times before merging. Keep the metadata field on each STT job to record the track name and the trim you applied.

When you need labels inside one file

If you only have a single mixed recording, there is no documented Sume STT setting that labels speakers. Say so in your pipeline, and do not synthesise speaker names from timestamps. Microsoft's page lists the feature for MAI-Transcribe-2; the page does not list a price, so compare cost only after you have fetched Microsoft's current rate card.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume