Deepgram Flux TTS inline IPA vs Sume's pronunciation dictionary id

Deepgram Flux TTS now takes inline IPA overrides in Early Access. Sume TTS takes a pronunciation_dict_id instead. How the two approaches differ in practice.

4 min readSume
All posts

Deepgram's September 30, 2026 entry adds inline pronunciation control to Flux TTS: you override one word in the text with IPA, on both batch and streaming requests, as an Early Access feature. Sume TTS takes a different route, an optional pronunciation_dict_id on the request, so the fix lives in a dictionary rather than inside each script.

Sources: the Deepgram changelog and Sume's API reference, read 2026-10-01.

How does the Flux inline override work?

The changelog's example puts a JSON object in the text, with the word and its IPA, for the drug name dupilumab. It says results can vary between generations while the feature is in Early Access, so generate each term several times before relying on it in production. Applied overrides are reported in controls_applied on streaming and the dg-pronunciations-applied header on batch.

What does Sume offer instead?

The TTS request schema has pronunciation_dict_id, described as an optional pronunciation dictionary id, up to 128 characters. The script text is a separate transcript field, capped at 20000 characters, so the same dictionary id can be reused across scripts. The cited schema does not describe how a dictionary is built, so confirm that in the docs before planning around it.

Which one fits which job?

A name that appears once suits an inline override. A brand or drug list that appears in every script suits a reusable dictionary.

Inline override versus dictionary id, from sources read 2026-10-01.
QuestionFlux inline IPASume dictionary id
Where the fix livesIn the text, per occurrenceOn the request, by id
Status stated by sourceEarly Access, results can varyOptional field in the schema
Combine with other controlsNot with pause or speed other than 1.0Not stated in the cited schema
Proof stepGenerate several timesListen to the output, then reuse the id

What should I test first?

Pick the three words most often misread, run them through a short script, and listen. The broader walkthrough is in text to speech pronunciation.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume