TTS invalid_voice_id 400: a voice name from another vendor
Sume returns 400 `invalid_voice_id` for a voice name that is not a UUID or `voi_` id, before any job or credit. Copy an id verbatim or send an avatar.

The 400 means voice.id was not in a shape Sume accepts. It must be a TTS voice UUID (8-4-4-4-12 hex) or a Voices library id (voi_ plus 32 hex); any other shape, such as a voice name copied from another vendor, is rejected synchronously with public_reason=invalid_voice_id. No job is queued and no credits are reserved. Fix it by copying an id verbatim or sending avatar_id or avatar_handle instead.
Why do other vendors' voices fail?
Each vendor names its own voices. The ElevenLabs changelog entry for September 28, 2026 (read 2026-09-30) announces Eleven v4 and Eleven v4 Turbo and mentions voice cloning across its models; the voices it uses are ElevenLabs voices. Sume's schema says plainly that voice.id is not a voice name from another TTS ecosystem, so a name or id from that ecosystem is not a Sume id.
What is accepted and what is rejected?
| Value sent | Result |
|---|---|
| TTS voice UUID, 8-4-4-4-12 hex | Accepted |
Voices library id, voi_ plus 32 hex | Accepted; resolved to the stored TTS voice before queueing |
| Any other shape, such as a vendor voice name | 400 invalid_voice_id, nothing queued, no credits reserved |
avatar_id or avatar_handle | Sume resolves that avatar's TTS voice at submit time |
How do I fix it?
Either copy the id exactly as you received it, or skip the voice id. The schema's discoverable selector is the avatar: list GET /v1/avatar-1.0/avatars and use any avatar whose voice.status is ready. If you set both an avatar and voice.id, they must match or you get a 400 for that instead.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: voice-fix-001" \
-d '{"transcript": "Your order has shipped.", "avatar_handle": "acme"}'Is it safe to retry?
Yes, in the sense that the rejection happens before a job exists, so the failed attempt cost nothing. Change the voice field and resend. For where voice ids come from, see HeyGen API avatar id and voice id and finding a voice id on Sume.
Sources
Related posts
More in Developers
- TTS locale vs language field: getting an en-GB accent
Cartesia's locale field picks a regional accent such as en-GB on Sonic 3.6. Sume's TTS request has a language field only, so pick the accent via the voice.
- TTS reads 1999-2000 wrong: write ranges and fractions as words
Cartesia does not normalize ranges like 1999-2000 or fractions like 2/3. Sume sends your transcript literally, so write them as words before you submit.
- TTS voice-language mismatch warning: confirm, then retry
Sume's TTS warns when a voice's language differs from `language`. No job or charge exists yet; after the user agrees, retry with `confirm_language_mismatch`.
- Twitch 2K clips: trim a 1440p clip without setting output
Video trim's output width and height are limited to 256-2160, so a 2560-wide 1440p clip cannot be conformed. Omit output and the source size is kept.
Written by Sume