No language warning on Sume TTS? Raw voice ids skip the check

Sume's voice-language check needs library metadata. Send a raw voice id it cannot look up and no mismatch warning appears. How to add your own safety net.

3 min readSume
All posts

Sume TTS 1.0 stops a Korean voice reading a Japanese script, but only when it knows what language the voice speaks. If library metadata is unknown or unavailable you can still submit a raw voice id. In that case there is nothing to compare, so no 409 and no warning.

What counts as a raw voice id

The voice.id field takes a TTS voice UUID (8-4-4-4-12 hex) or a Voices library id (voi_ plus 32 hex). Anything else, such as a voice name from another vendor, is rejected at once with 400 invalid_voice_id before a job is queued or credits are reserved. A UUID that Sume has no language record for passes shape validation, and nothing is compared.

Why it matters

The mistake this check exists for is quiet. A voice reading text in a language it does not speak produces audio, and you pay for it: at $0.0475 per 1,000 characters, a 20,000-character script is $0.95. Nothing errors. You find out when you listen.

Build your own safety net

  • Keep a small table in your project: voice id, language, region, who approved it.
  • Send the language field on every non-English request, so the request states the language you intend.
  • Render the first 200 characters of any new voice and script pair as a 1 cent test before the full job.
  • Prefer an avatar or library voice, where Sume knows the language, over a bare UUID.

A cheap test

Submit the same short script with the new raw id and with a library voice in the same language. If the raw id sounds off, you have spent about 2 cents. If you skip the test and the full script is wrong, you have spent up to 95 cents per 20,000 characters, plus any video render that used the audio.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume