Resolve 21 Speech Generation vs a text-to-speech API call
Resolve 21 generates speech from text with Blackmagic voice models. Sume's tts_create does it from a script, and a voice-language mismatch returns 409 first.

DaVinci Resolve 21 adds Speech Generation, which turns text into speech with Blackmagic voice models. If you need the same thing in a script, Sume exposes tts_create as an MCP tool and a text-to-speech REST route. One behavior worth knowing: Sume compares the voice's primary language with the request's language and returns HTTP 409 tts_voice_language_mismatch before any job or charge when they disagree.
How the language check behaves
The voice's primary language is stored with the voice, and TTS reads the target from language on REST or payload.language on MCP. A known mismatch stops before billing. Over MCP, tts_create returns a non-error tts_voice_language_warning result with confirmation_required: true. Ask the user, then retry the same request and idempotency key with confirm_language_mismatch: true. Regional tags compare by primary language, and fil and tl match. If the library has no language metadata for a raw voice ID, no check blocks the submission.
| Question | Resolve 21 Speech Generation | Sume text to speech |
|---|---|---|
| Where it runs | Inside Resolve | REST route or MCP tool tts_create |
| Voices | Blackmagic voice models | Sume voice library and avatar voices |
| Wrong-language guard | Not described on the page | 409 before any charge, or a confirm step over MCP |
| Billing | Included with the app | Per job, shown in GET /v1/catalog |
Where speech goes next
A generated voice track is usually only half the job. Pair it with Sume's Timeline 1.0 render, which takes a spine of up to 1800 seconds and 1 to 200 video slots, or with a caption call that burns text in. The documented dubbing pipeline stores the steps as detach, speech-to-text, translate, then text-to-speech.
Honest limits
- Blackmagic's page names the feature but gives no language or length limits, so none are compared.
- Sume does not edit Resolve projects; this is a separate route for files.
- The Sume language check is a safeguard, not a pronunciation check.
Sources
Related posts
More in Media tools
- Resolve 21 UltraSharpen and Motion Deblur vs an upscale API
Resolve 21 adds UltraSharpen and Motion Deblur. Sume's video upscale API takes 1 to 30 seconds, a 1.1 to 4 factor and a fast, standard or pro tier.
- Runway ACEScg EXR: what the September 12 plate-referenced change did
Runway changed its ACEScg EXR build on September 12: from the source plate, not the generated frames. What changed, what did not, and Sume's MP4 output.
- Runway Ruby SDR-to-HDR API: 30-second limit, 4096 px, 20 credits/s
Runway's Ruby turns SDR video into HDR for 20 credits per second, 40 above 4 megapixels, with a 30-second input cap. Sume outputs MP4, so trim first.
- SCC caption file: YouTube's preferred format vs Sume burn-in
YouTube calls .scc its preferred caption file format. Sume burns captions into the video and takes no SRT upload. When each one is right, plus a workflow.
Written by Sume