"Avatar does not have a usable TTS voice": the 400 and its fixes

Sume TTS with avatar_id or avatar_handle returns 400 when the avatar has no TTS voice, or when voice.id disagrees with it. What each message means and the fix.

4 min readSume
All posts

When a Sume text-to-speech request names an avatar with avatar_id or avatar_handle, the API resolves that avatar to its TTS voice before it queues anything. If the avatar has no usable voice, the answer is HTTP 400 invalid_request with the message "Avatar does not have a usable TTS voice. Clone or transfer a voice for this avatar first." The fix is on the avatar, not on your request: give the avatar a voice, or skip the avatar selector and send a voice id directly.

The same resolution step produces three other 400s and one 404, so it pays to read the message before you change anything.

Which message means what?

The table separates the five outcomes by the message text, so match your error log against the first column before you change any code.

Avatar selector errors on TTS 1.0 and TTS Router, from the Sume API source and OpenAPI document (read 2026-10-03)
SymptomStatus and codeWhat it meansFix
Avatar does not have a usable TTS voice400 invalid_request, field avatar_id or avatar_handleThe avatar exists in your workspace but carries no Cartesia-backed TTS voiceClone or transfer a voice for the avatar, or pass voice.id yourself
voice.id conflicts with the TTS voice resolved from the avatar reference400 invalid_request, field voiceYou sent both an avatar and a voice id and they differOmit voice.id, or send the avatar's own voice id
avatar_id and avatar_handle must resolve to the same avatar400 invalid_request, field avatar_handleThe two selectors point at different avatarsSend one selector, not both
avatar_handle must be 2-30 lowercase letters, numbers, underscores, or periods400 invalid_request, field avatar_handleHandle grammar: periods and underscores cannot be first, last, or consecutiveFix the handle, or use avatar_id
Avatar not found404The id or handle is not in the calling workspace (another workspace's avatar looks the same)List avatars with the same key you generate with

How do I tell whether an avatar has a voice before I call TTS?

The OpenAPI description for the avatar selector says it directly: GET /v1/avatar-1.0/avatars lists the workspace's avatars, and an avatar's voice.status is ready exactly when the selector works for TTS. If it is anything else, you will get the first row of the table above. Check it once when you onboard an avatar instead of discovering it at the first narration job.

What if I only have a voice id?

Then you do not need the avatar at all. The voice selector is { "mode": "id", "id": "..." } and accepts a TTS voice UUID or a Voices library id (voi_ followed by 32 hex characters). A value of any other shape fails with invalid_voice_id before a job is queued. When you send both an avatar and a voice, the voice id has to equal the avatar's resolved one, which is why the second row in the table exists.

Sume's tool guidance for agents takes the same position: Assets and Voices are optional references, not an admission gate, so a raw voice id can be passed straight to tts_create without a workspace row.

import os
import uuid

import requests

r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={
        "x-api-key": os.environ["SUME_API_KEY"],
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "transcript": "Welcome back to the weekly update.",
        "avatar_handle": "studio_host",
        "mode": "async",
    },
    timeout=30,
)
print(r.status_code)
print(r.json().get("error", {}).get("message"))

Does a failed selector cost anything?

No job is created for these 400s, and TTS is only charged once a job exists. When a request does go through, text to speech is billed per transcript character at $0.0475 per 1,000 characters, with spaces and punctuation counting and a maximum of 20,000 characters per request. A 3,000-character narration is $0.1425.

Treat the error as a setup problem rather than a transient one. It will not clear on retry, and it will not clear because the avatar's video is ready: a talking-head avatar can render video while its TTS voice is still unset, since the two are separate fields.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume