Descript's ElevenLabs voice library vs Sume TTS voices

Descript's voice picker now uses the ElevenLabs voice library. Sume text-to-speech uses its own voices, and blocks a voice-language mismatch with a 409.

5 min readSume
All posts

Descript's 2026-09-17 changelog says its voice picker moved to the ElevenLabs voice library. If you want a script-driven voice instead, Sume's text-to-speech takes a Sume voice from its own library, and the route guards against choosing a voice in the wrong language: a known mismatch returns HTTP 409 tts_voice_language_mismatch before any job or charge.

Voices and language

Each voice stores a primary language. The request sets the target language with language on REST or payload.language on the MCP tool tts_create. Regional tags compare by primary language. Over MCP the mismatch arrives as a warning result with confirmation_required: true; retry with confirm_language_mismatch: true after asking the user. If the library has no metadata for a raw voice id, nothing blocks the request.

Voice sources compared, read 2026-10-03
QuestionDescript pickerSume text to speech
Voice sourceElevenLabs voice librarySume voice library
Where you use itDescript editorREST or MCP tool
Language guardNot described in the changelog409 before any charge
ResultAudio in the projectJob with an audio artifact

Why the guard matters

A batch job that pairs a French voice with Korean text would otherwise waste credit. Failing before billing makes a bulk script safe to retry: use the same Idempotency-Key for the retry.

Honest limits

  • Sume's voices are not ElevenLabs voices, and this post does not compare quality.
  • The language check does not test pronunciation.
  • Descript's changelog does not give voice counts or limits.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume