Deepgram Flux TTS watermark is mandatory in self-hosted 261001
Deepgram's self-hosted release 261001 makes Flux TTS watermarking mandatory. What changes for self-hosters, and what Sume's hosted TTS jobs do and do not say.

Deepgram's self-hosted release 261001, dated October 1, 2026, makes watermarking of Flux TTS audio mandatory: the engine needs a watermarker.<uuid>.dgv2 file and a watermarker_uuid setting under [flux_tts]. If you run Flux TTS on your own GPUs, plan that file into your deploy. If you call a hosted TTS API instead, the thing to check is what that vendor says about marks on its audio.
Everything about Deepgram below comes from its changelog, read on 2026-10-03. Everything about Sume comes from the Sume OpenAPI reference and the job docs. The changelog entry does not say how the watermark is detected, how audible it is, or whether it survives re-encoding, so this post does not guess.
What exactly changed in self-hosted release 261001?
The changelog lists four container images tagged release-261001: the API, the engine, the license proxy and billing. FIPS variants carry a -fips suffix. The engine image requires an NVIDIA driver at version 580 or newer. Besides the TTS watermark, the release removes Whisper model support, so self-hosters on Whisper must move to Nova-3. It also aligns FIPS engine metrics and certificates with the standard images, and improves Japanese numerals formatting and German number, date and measurement reading.
| Change | What the changelog says |
|---|---|
| Flux TTS watermarking | Now mandatory; needs a watermarker file and a watermarker_uuid under [flux_tts] |
| Whisper | Support removed; migrate to Nova-3 |
| Engine host | NVIDIA driver 580 or newer |
| FIPS | -fips images for all four services |
| Formatting | Japanese numerals; German numbers, dates and measurements |
What does mandatory mean for your deploy?
Read it as a configuration gate: the changelog says the watermark is required, and names the two things you supply. It does not say the engine refuses to start without them, so test a staging engine on 261001 before you roll it to production, and keep the watermarker file in your secret store like a license file. Then run your usual TTS acceptance clips and listen, because a new engine build can shift output even when the release notes do not mention it.
If you resell the audio, the mark is also a fact to tell customers. A mandatory mark on the engine's output is a provenance signal you did not add yourself, which can matter for the disclosure rules in your market.
Does Sume's TTS carry a watermark?
Sume's documentation does not describe a watermark or an embedded provenance record on TTS audio. What it documents is the job: you submit text to POST /v1/tts-1.0/generate or to the TTS Router, and you get a Sume job with a hosted audio artifact on media.sume.com. The router's catalog ids today are Cartesia Sonic ones (sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview), and billing is per character, $0.0475 per 1,000 characters at the public rate.
So the honest comparison is not watermark against watermark. It is a self-hosted engine whose vendor now mandates a mark, against a hosted route whose docs say nothing either way. If a client asks you to prove an audio file is marked, Sume's docs give you a job id and artifact, not a mark you can point to. Ask the provider you rely on, and do not promise marks Sume has not documented.
Which should you pick?
Pick self-hosting if you need audio never to leave your own infrastructure, you can run the NVIDIA hardware, and you want the engine vendor's mark on every file. Pick a hosted job route if you would rather pay per character, avoid GPU operations, and keep a per-request record. Both can sit in the same pipeline: generate, then log the job id or deploy build, then edit.
Whichever you choose, keep a ledger entry per file: engine and version, the request, and the date. That ledger is what you hand over when someone asks where a clip came from.
- Self-hosted Flux TTS: record the release tag (for example 261001) and the watermarker uuid you configured.
- Sume TTS: record the job id, the model id you sent and the voice selector.
- After any edit, record the tool and the new file, since an edit may change what a mark does.
Sources
Related posts
More in Comparisons
- Deepgram India endpoint GA: voice data localization vs a US-hosted API
Deepgram's api.in.deepgram.com is generally available for STT, TTS and voice agents. If data must stay in India, what Sume's docs say about where its API runs.
- Deepgram Nova-3 keyterms mid-stream: update vocabulary live vs
Deepgram's Oct 2 update swaps Nova-3 keyterms during a stream with a Configure message. What the 500-token limit means, and why Sume STT jobs have no such.
- Deepgram self-hosted drops Whisper in 261001: a hosted STT job instead
Deepgram's release 261001 removes Whisper support from self-hosted. Your options: move to Nova-3, or send recordings to a hosted STT job such as Sume's.
- DeepL Voice keeps speakers' voices in 14 languages: Sume dubs
DeepL's September 15 release keeps each speaker's voice across 14 languages in live talk. A Sume dub picks a TTS voice id per line and does not clone.
Written by Sume