Cartesia STT keyterms for brand names, and Sume STT without keyterms
Cartesia STT adds keyterm prompting for brand names and invented words. Sume STT has no keyterm field, so here is how to fix brand names after transcription.

Cartesia's speech-to-text now accepts keyterms: domain-specific terms, brand or product names, and rare or invented words that the model should transcribe correctly. Sume STT's request schema has no keyterm field: its provider knobs are fixed server-side, so brand names are corrected after transcription, using the words[] timings that every result carries.
Sources: the Cartesia 2026 changelog and Sume's API reference, read 2026-10-01.
What did Cartesia add?
The changelog entry is called keyterm prompting for better accuracy, with a keyterms guide and controls in the Speech-to-Text Playground and the API. It does not state limits or accuracy numbers on the page read here.
What does the Sume request accept?
Word timings are always returned, so there is no flag to enable them.
| Field | Role |
|---|---|
audio_url | Public HTTPS audio |
language_code | Optional hint; omit for auto-detect |
duration_seconds | Optional, 1 to 600, improves usage reservation |
segmentation | Optional sentence segments |
| Provider knobs | Fixed server-side, such as diarize and tag_audio_events |
How do I fix brand names without keyterms?
Keep a short list of the names you care about and their common mis-spellings, then replace them in the returned text. Because each word has start and end in seconds, you can also flag a word for a human check and jump straight to that moment in the audio. This is your own post-processing, not a Sume feature, and it only works for errors you can predict.
The language_code hint can help when a name is in a different language from the rest of the clip, but the docs describe it only as a language hint.
When is a keyterm field worth having?
When the vocabulary is large or changes weekly, and a replace list would lag behind. For drug-name accuracy specifically see Nova-3 Pharma.
Sources
Related posts
More in Models
- Creatify Aurora 2.0 max audio length: 59.7 s, and how to split
Creatify Aurora 2.0 takes up to 59.7 s of audio; Aurora v1 takes 5 minutes. Sume Avatar Video accepts 4-60 s per job, so split longer scripts across jobs.
- Creatify Boreal talking clips vs Sume's still-plus-audio route
Creatify says Boreal's gains are smallest on single-person talking clips. Sume makes every speaking shot from an accepted still plus TTS audio via Fabric.
- D-ID V4 expressives sentiment_id vs Sume's emotion string
D-ID picks a delivery with a sentiment_id preset; Sume TTS takes a free-text emotion string plus a 0.6-1.5 speed multiplier. How each one is set.
- Deepgram Nova-3 de-CH, Lithuanian, Portuguese: language_code on Sume
Deepgram improved Nova-3 for Swiss German, Lithuanian and Portuguese on September 24. What to send as language_code on Sume STT for regional variants.
Written by Sume