Deepgram Flux TTS inline IPA vs Sume's pronunciation dictionary id
Deepgram Flux TTS now takes inline IPA overrides in Early Access. Sume TTS takes a pronunciation_dict_id instead. How the two approaches differ in practice.

Deepgram's September 30, 2026 entry adds inline pronunciation control to Flux TTS: you override one word in the text with IPA, on both batch and streaming requests, as an Early Access feature. Sume TTS takes a different route, an optional pronunciation_dict_id on the request, so the fix lives in a dictionary rather than inside each script.
Sources: the Deepgram changelog and Sume's API reference, read 2026-10-01.
How does the Flux inline override work?
The changelog's example puts a JSON object in the text, with the word and its IPA, for the drug name dupilumab. It says results can vary between generations while the feature is in Early Access, so generate each term several times before relying on it in production. Applied overrides are reported in controls_applied on streaming and the dg-pronunciations-applied header on batch.
What does Sume offer instead?
The TTS request schema has pronunciation_dict_id, described as an optional pronunciation dictionary id, up to 128 characters. The script text is a separate transcript field, capped at 20000 characters, so the same dictionary id can be reused across scripts. The cited schema does not describe how a dictionary is built, so confirm that in the docs before planning around it.
Which one fits which job?
A name that appears once suits an inline override. A brand or drug list that appears in every script suits a reusable dictionary.
| Question | Flux inline IPA | Sume dictionary id |
|---|---|---|
| Where the fix lives | In the text, per occurrence | On the request, by id |
| Status stated by source | Early Access, results can vary | Optional field in the schema |
| Combine with other controls | Not with pause or speed other than 1.0 | Not stated in the cited schema |
| Proof step | Generate several times | Listen to the output, then reuse the id |
What should I test first?
Pick the three words most often misread, run them through a short script, and listen. The broader walkthrough is in text to speech pronunciation.
Sources
Related posts
More in Developers
- Flux TTS pause markers: 8 per request, 500-3000 ms, and Sume scripts
Flux TTS allows up to 8 pause markers of 500 to 3000 ms per request, batch only. Plan a script around that limit and see what Sume gives you for pacing.
- Deepgram self-hosted 261001 drops Whisper: what Sume STT callers do
Deepgram Self-Hosted release 261001 removes Whisper support and asks you to move to Nova-3 first. What it means for Sume STT, which has no model to migrate.
- Deno fetch stops retrying fresh-connection POSTs: Sume keys
Deno now retries fetch transport errors only on reused connections. That cuts duplicate POSTs, but a paid Sume submit still needs an Idempotency-Key.
- Remove video speckle noise by API: median filter radius
Sume's video filter allowlists median, which replaces each pixel with the middle value of its neighbours. Radius runs 1 to 127; start at 1 and compare.
Written by Sume