Deepgram Nova-3 keyterms mid-stream: update vocabulary live vs

Deepgram's Oct 2 update swaps Nova-3 keyterms during a stream with a Configure message. What the 500-token limit means, and why Sume STT jobs have no such.

5 min readSume
All posts

On October 2, 2026, Deepgram added a way to change Nova-3 keyterms while a stream is open: send a Configure message on the /v1/listen connection and the keyterm list is replaced without reconnecting. Sume's speech-to-text route works differently. It is a batch job on a finished audio URL with no vocabulary field, so the way to improve a name the transcript got wrong is to fix it in text afterwards.

Deepgram details here are from its changelog (read 2026-10-03). Sume details come from the OpenAPI reference and the jobs docs.

How does the Deepgram mid-stream update work?

The message is a JSON object of type Configure carrying a keyterms array, for example { "type": "Configure", "keyterms": ["Deepgram", "customer service"] }. According to the changelog, each array replaces the entire keyterm list, including terms set through the query parameter, and an empty array clears all keyterms. The limit is 500 tokens per update, the same as the query parameter, and an update over the limit returns an Error and leaves the stream unchanged.

You can combine it with feature toggles in the same message, such as "features": { "numerals": true }. It works on all Nova-3 streaming models including multilingual and Medical. Two caveats from the page: it is available on the global endpoint only, not yet on the EU, Australia or India endpoints, and any other model returns a KeytermsNotSupported error.

Deepgram Nova-3 mid-stream keyterms, from the Oct 2 2026 changelog entry (read 2026-10-03)
ItemDetail
MessageConfigure with a keyterms array on /v1/listen
EffectReplaces the whole list, query-parameter terms included
ClearSend an empty array
Limit500 tokens per update; over-limit returns Error, stream unchanged
ModelsAll Nova-3 streaming models; others return KeytermsNotSupported
EndpointsGlobal only for now, not EU, Australia or India

When is a live vocabulary swap useful?

It helps when the right words change mid-call: a support line that moves from billing to a drug name, a meeting that switches from one product to another, an agent that has just looked up a customer's surname. Without it you either carry every possible term from the start, hitting the token cap, or you reconnect and lose a moment of audio.

Because the update replaces the list, your client needs to hold the full list it wants at each moment, not only the new terms. Remember that the 500 tokens are tokens, not words, so long compound names eat the budget faster.

What does Sume STT offer instead?

Sume STT is POST /v1/stt-1.0/transcribe with a public HTTPS audio_url. The request accepts an optional language_code hint, an optional duration_seconds between 1 and 600 for usage reservation, an optional sentence segmentation, plus metadata and the usual mode fields. Word timings are always returned as words[] with start and end seconds. The reference says provider knobs such as diarization are fixed server-side, and there is no vocabulary or keyterm field.

It is also not a streaming socket. A job runs on a recorded file, with a sync wait of at most 30 seconds, so for a live call you would transcribe a recording after the fact. The price is $0.01 per audio minute at the public rate, reserved at one minute when you omit duration_seconds.

How do you fix names in a Sume transcript?

Run the job, then correct known terms in text. Keep a small map of misheard forms to the right spelling for your brand and product names, apply it to the returned text, and keep the word timings untouched so captions stay in sync. If a term is wrong in a way a map cannot catch, review that line by hand and keep the original job id next to the corrected text.

If live vocabulary swapping is the requirement, Deepgram's streaming API is built for it and Sume's batch job is not. If you transcribe recordings and want a stable job record, per-minute pricing and word timings, the Sume route fits.

  • Live call, vocabulary changes mid-stream: Deepgram Configure message.
  • Recorded file, fixed spelling pass afterwards: Sume STT job and a text map.
  • Both: store the raw transcript and the corrected one separately.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume