Deepgram Flux TTS watermark is mandatory in self-hosted 261001

Deepgram's self-hosted release 261001 makes Flux TTS watermarking mandatory. What changes for self-hosters, and what Sume's hosted TTS jobs do and do not say.

5 min readSume
All posts

Deepgram's self-hosted release 261001, dated October 1, 2026, makes watermarking of Flux TTS audio mandatory: the engine needs a watermarker.<uuid>.dgv2 file and a watermarker_uuid setting under [flux_tts]. If you run Flux TTS on your own GPUs, plan that file into your deploy. If you call a hosted TTS API instead, the thing to check is what that vendor says about marks on its audio.

Everything about Deepgram below comes from its changelog, read on 2026-10-03. Everything about Sume comes from the Sume OpenAPI reference and the job docs. The changelog entry does not say how the watermark is detected, how audible it is, or whether it survives re-encoding, so this post does not guess.

What exactly changed in self-hosted release 261001?

The changelog lists four container images tagged release-261001: the API, the engine, the license proxy and billing. FIPS variants carry a -fips suffix. The engine image requires an NVIDIA driver at version 580 or newer. Besides the TTS watermark, the release removes Whisper model support, so self-hosters on Whisper must move to Nova-3. It also aligns FIPS engine metrics and certificates with the standard images, and improves Japanese numerals formatting and German number, date and measurement reading.

Deepgram self-hosted release 261001 items that touch audio (read 2026-10-03)
ChangeWhat the changelog says
Flux TTS watermarkingNow mandatory; needs a watermarker file and a watermarker_uuid under [flux_tts]
WhisperSupport removed; migrate to Nova-3
Engine hostNVIDIA driver 580 or newer
FIPS-fips images for all four services
FormattingJapanese numerals; German numbers, dates and measurements

What does mandatory mean for your deploy?

Read it as a configuration gate: the changelog says the watermark is required, and names the two things you supply. It does not say the engine refuses to start without them, so test a staging engine on 261001 before you roll it to production, and keep the watermarker file in your secret store like a license file. Then run your usual TTS acceptance clips and listen, because a new engine build can shift output even when the release notes do not mention it.

If you resell the audio, the mark is also a fact to tell customers. A mandatory mark on the engine's output is a provenance signal you did not add yourself, which can matter for the disclosure rules in your market.

Does Sume's TTS carry a watermark?

Sume's documentation does not describe a watermark or an embedded provenance record on TTS audio. What it documents is the job: you submit text to POST /v1/tts-1.0/generate or to the TTS Router, and you get a Sume job with a hosted audio artifact on media.sume.com. The router's catalog ids today are Cartesia Sonic ones (sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview), and billing is per character, $0.0475 per 1,000 characters at the public rate.

So the honest comparison is not watermark against watermark. It is a self-hosted engine whose vendor now mandates a mark, against a hosted route whose docs say nothing either way. If a client asks you to prove an audio file is marked, Sume's docs give you a job id and artifact, not a mark you can point to. Ask the provider you rely on, and do not promise marks Sume has not documented.

Which should you pick?

Pick self-hosting if you need audio never to leave your own infrastructure, you can run the NVIDIA hardware, and you want the engine vendor's mark on every file. Pick a hosted job route if you would rather pay per character, avoid GPU operations, and keep a per-request record. Both can sit in the same pipeline: generate, then log the job id or deploy build, then edit.

Whichever you choose, keep a ledger entry per file: engine and version, the request, and the date. That ledger is what you hand over when someone asks where a clip came from.

  • Self-hosted Flux TTS: record the release tag (for example 261001) and the watermarker uuid you configured.
  • Sume TTS: record the job id, the model id you sent and the voice selector.
  • After any edit, record the tool and the new file, since an edit may change what a mark does.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume