OpenAI gpt-transcribe keywords vs Sume STT: no vocabulary field

OpenAI's STT guide lists keywords, prompt and languages parameters. Sume STT has none: only audio_url, language_code, duration and segmentation.

3 min readSume
All posts

OpenAI's speech-to-text guide lists keywords, prompt and languages parameters for gpt-transcribe, which let you bias recognition toward product names and jargon (read 2026-10-03). Sume's POST /v1/stt-1.0/transcribe has no keyword or vocabulary field, so names have to be fixed after the fact.

Parameters side by side

Sume's request takes audio_url, language_code, duration_seconds between 1 and 600, optional segmentation.mode set to sentence, and metadata. That is the whole input.

STT request controls (read 2026-10-03)
ControlOpenAI gpt-transcribeSume STT 1.0
Vocabulary or keyword hintskeywords and promptNone
Languagelanguageslanguage_code
Speaker separationSeparate diarize modelNot documented
Word timestampswhisper-1 for timestampswords[] on every result
Max input25 MB file600 seconds

Fixing names after the transcript

Because the result includes words[] with start and end offsets, a find-and-replace pass keeps timing intact. Build a small map of the misheard form to the correct spelling, apply it to text and to each word, and leave the offsets alone. For captions, the Sume caption endpoint accepts script_text, so you can supply the correct script instead of correcting a transcript.

Choosing between them

If your audio is full of brand terms, product codes or names, a keyword hint is worth real accuracy and OpenAI's parameter is the better fit. If you need word timings that feed a Sume caption or timeline job, Sume's result is already in the right shape. For per-minute cost see the 1,000-hour comparison.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume