Grok STT vad_threshold for quiet audio vs Sume STT, which has no gate

xAI added a vad_threshold knob to speech-to-text for quiet or noisy audio. Sume STT exposes no such gate; here is what you can tune instead.

5 min readSume
All posts

If you need to transcribe very quiet or noisy speech and you are looking for a voice-activity threshold, Sume STT does not have one. Its request schema takes the audio, an optional language hint, a duration hint and sentence segmentation, and nothing that gates non-speech audio. xAI's Speech to Text does expose that control, so the two are not interchangeable on this one knob.

This post lays out what xAI documents, what Sume's schema actually accepts, and what to do to your audio before it reaches Sume if your recordings are faint.

What did xAI add to speech-to-text?

xAI's release notes say Speech to Text now accepts a vad_threshold parameter, as a streaming query parameter and as a batch multipart field, to tune the voice-activity gate that skips non-speech audio. Per the same notes, lower values transcribe quieter or noisy speech, it is useful for narrowband telephony, and setting it to 0 disables the gate. The notes also describe a separate smart_turn end-of-turn option for streaming. Read at xAI release notes on 2026-10-04.

Voice-activity control, xAI release notes and Sume spec, read 2026-10-04
QuestionxAI Speech to TextSume STT 1.0
Tunable voice-activity thresholdYes, vad_thresholdNo such field in the request schema
Value 0Disables the gateNot applicable
Language hintNot covered in these notesOptional language_code, omit for auto-detect
Sentence segmentation from word timingsNot covered in these notesOptional segmentation with mode sentence

What does Sume STT let you set?

In the Sume OpenAPI spec, the STT 1.0 request requires audio_url, a public HTTPS link. Optional fields are language_code, duration_seconds (used for usage reservation, up to 10 minutes), segmentation, metadata, mode, webhook_url and wait_timeout_seconds. Diarization and audio-event tagging are fixed server-side, and word timings are always returned.

Because there is no gate to loosen, Sume's answer to faint speech is upstream: fix the level before you submit. Sume STT is priced at $0.01 per audio minute, so a retry on a cleaned-up file costs a cent a minute, not a rebuild.

How do you prepare quiet audio for Sume STT?

Treat it as a two-step job. First normalise the level where you host the file, then submit the new public URL. Sume does not do this step for you, and we are not claiming that it improves accuracy for every recording.

  • Raise the gain on the source file with your own audio tool and re-host it at a public HTTPS URL.
  • Pass language_code when you know it, for example ko, instead of relying on auto-detect for short clips.
  • Keep duration_seconds close to the real length so the reservation matches the work.
  • Compare the transcript against a known passage before you run a whole archive.

Which one fits your pipeline?

Choose xAI's endpoint when the threshold itself is the thing you need to tune, for example narrowband call audio where the gate swallows words. Choose Sume STT when you want a batch job with polling URLs, sentence segments and one price, and you can fix levels before upload. For a cost comparison across providers see cheapest speech-to-text API per hour, and for language hints see single language_code vs multiple expected languages.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume