Grok STT vad_threshold for quiet audio vs Sume STT, which has no gate
xAI added a vad_threshold knob to speech-to-text for quiet or noisy audio. Sume STT exposes no such gate; here is what you can tune instead.

If you need to transcribe very quiet or noisy speech and you are looking for a voice-activity threshold, Sume STT does not have one. Its request schema takes the audio, an optional language hint, a duration hint and sentence segmentation, and nothing that gates non-speech audio. xAI's Speech to Text does expose that control, so the two are not interchangeable on this one knob.
This post lays out what xAI documents, what Sume's schema actually accepts, and what to do to your audio before it reaches Sume if your recordings are faint.
What did xAI add to speech-to-text?
xAI's release notes say Speech to Text now accepts a vad_threshold parameter, as a streaming query parameter and as a batch multipart field, to tune the voice-activity gate that skips non-speech audio. Per the same notes, lower values transcribe quieter or noisy speech, it is useful for narrowband telephony, and setting it to 0 disables the gate. The notes also describe a separate smart_turn end-of-turn option for streaming. Read at xAI release notes on 2026-10-04.
| Question | xAI Speech to Text | Sume STT 1.0 |
|---|---|---|
| Tunable voice-activity threshold | Yes, vad_threshold | No such field in the request schema |
| Value 0 | Disables the gate | Not applicable |
| Language hint | Not covered in these notes | Optional language_code, omit for auto-detect |
| Sentence segmentation from word timings | Not covered in these notes | Optional segmentation with mode sentence |
What does Sume STT let you set?
In the Sume OpenAPI spec, the STT 1.0 request requires audio_url, a public HTTPS link. Optional fields are language_code, duration_seconds (used for usage reservation, up to 10 minutes), segmentation, metadata, mode, webhook_url and wait_timeout_seconds. Diarization and audio-event tagging are fixed server-side, and word timings are always returned.
Because there is no gate to loosen, Sume's answer to faint speech is upstream: fix the level before you submit. Sume STT is priced at $0.01 per audio minute, so a retry on a cleaned-up file costs a cent a minute, not a rebuild.
How do you prepare quiet audio for Sume STT?
Treat it as a two-step job. First normalise the level where you host the file, then submit the new public URL. Sume does not do this step for you, and we are not claiming that it improves accuracy for every recording.
- Raise the gain on the source file with your own audio tool and re-host it at a public HTTPS URL.
- Pass
language_codewhen you know it, for exampleko, instead of relying on auto-detect for short clips. - Keep
duration_secondsclose to the real length so the reservation matches the work. - Compare the transcript against a known passage before you run a whole archive.
Which one fits your pipeline?
Choose xAI's endpoint when the threshold itself is the thing you need to tune, for example narrowband call audio where the gate swallows words. Choose Sume STT when you want a batch job with polling URLs, sentence segments and one price, and you can fix levels before upload. For a cost comparison across providers see cheapest speech-to-text API per hour, and for language hints see single language_code vs multiple expected languages.
Sources
Related posts
More in Comparisons
- HeyGen translates 10 languages at once: what Sume does
HeyGen translate handles up to 10 target languages per job. Sume captions burn authored text per language, one $0.20 job per clip up to 60 seconds.
- HeyGen speaker rules as a checklist for a Sume face swap clip
HeyGen asks for a face within 45 degrees, one speaker at a time and under 10 feet. Use that as a pre-check, then Sume's own 4 to 15 second rules.
- HeyGen Video Agent prompt-to-video vs a Sume script-driven avatar job
HeyGen's v3 quick start creates a video from a prompt and returns a session_id. Sume's avatar route takes a script and avatar handle and returns a job.
- Ideogram Plus or Pro vs API per image: what the page shows
Ideogram sells Plus at $15 and Pro at $42 a month in credits, and its page lists no per-image API price. How to compare it with a per-image row, with a script.
Written by Sume