Speech to text custom vocabulary: Sume STT takes a language hint only
Gemini 3.5 Transcribe biases up to 1,000 custom terms. Sume STT has no vocabulary field, only a language_code hint, so fix names after transcription.

Custom vocabulary for speech to text means passing a list of terms so the model spells names and jargon correctly. Gemini 3.5 Transcribe supports up to 1,000 such terms; Sume STT (sume/stt-1.0) has no vocabulary field. Its request accepts audio_url, language_code, duration_seconds, segmentation, metadata and the job-mode fields (mode, webhook_url, wait_timeout_seconds), and rejects anything else.
Vendor facts are from the Gemini model page, read 2026-09-30; Sume facts from the OpenAPI document.
How does Gemini describe custom vocabulary?
The model page lists custom vocabulary biasing at up to 1,000 terms, adding that customers typically see best results with up to 100. It is incompatible with diarization and with word-level timestamps, so you choose between biasing and those features on a single request. OpenAI's speech-to-text guide is in the same area; its quickstart model, gpt-transcribe, is the one the guide recommends for recorded speech.
What does Sume STT accept?
The request schema is closed (additionalProperties: false), so a vocabulary or keywords key would fail validation. The one language control is language_code, described as an optional BCP-47 or provider language hint; omit it for auto-detect. Word timings are always returned.
| Option | Gemini 3.5 Transcribe | Sume STT 1.0 |
|---|---|---|
| Custom term list | Up to 1,000 terms | Not accepted |
| Language hint | Auto-detects 85+ languages | language_code or auto-detect |
| Combined with word timings | Incompatible | Timings always returned |
How do I get product names right on Sume?
Set language_code when you know the language, then correct terms after transcription with a replacement list you maintain, matched against text or against words[] so the timings stay valid. See language hint versus auto-detect.
Is there a vocabulary control on the text-to-speech side?
Yes, in the other direction: the TTS request has an optional pronunciation_dict_id, described as an optional pronunciation dictionary id. That changes how words are spoken, not how they are recognised. See text-to-speech pronunciation.
Sources
Related posts
More in Developers
- Stripe idempotency key length (255) and Sume's key rules
Stripe allows keys up to 255 characters. Sume generation keys are 1-255 printable ASCII and optional; crawl keys are required and capped at 200.
- Submitting 20 avatar videos at once: what each Sume plan accepts
Sume accepts as many paid jobs as concurrency plus queue allows: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Past that a submit returns 429 queue_full.
- Which model does Sume Agent Completions run? Only sume-agent
The model field on POST /v1/agent/completions accepts only sume-agent. Omit it for the same agent; any other value returns 400 invalid_request.
- sume --agent --json: what it redacts before you paste output
Sume CLI --agent mode redacts or summarizes URL-like and account fields where supported. What to still strip before pasting output into a ticket or chat.
Written by Sume