TTS with only an API key: list avatars and pick one with voice ready
You do not need a voice id for Sume TTS. List your avatars, pick one whose voice status is ready, and send its handle as avatar_handle. Python example.
Sume TTS wants a voice. If you only have an API key, you do not need to hunt for a voice id: GET /v1/avatar-1.0/avatars lists avatars in your workspace, and an avatar whose voice.status is ready can be used directly as the selector with avatar_id or avatar_handle.
List and pick
The response puts the avatars in data.avatars. Each summary has an id, a handle, and a voice object whose status is processing, ready or failed, or null when the avatar has no voice. This Python picks the first ready one.
import json
import os
import urllib.request
key = os.environ.get("SUME_API_KEY")
if not key:
raise SystemExit("set SUME_API_KEY")
req = urllib.request.Request(
"https://api.sume.com/v1/avatar-1.0/avatars",
headers={"Authorization": "Bearer " + key},
)
with urllib.request.urlopen(req) as r:
avatars = json.load(r)["data"]["avatars"]
ready = [a for a in avatars if (a.get("voice") or {}).get("status") == "ready"]
print(ready[0]["handle"] if ready else "no avatar with a ready voice")
Then speak
Send avatar_handle with transcript and language to the TTS generate route. Do not also send voice.id unless it equals the avatar's own voice id, or the request is rejected. Cost is the usual $0.0475 per 1,000 characters.
If nothing is ready
An empty list means you have no avatar with a finished voice. Create one first; a voice with processing status should not be used yet, and failed needs a new attempt.
Related posts
More in Developers
- TTS word timings to karaoke captions: map words[] to text
Sume TTS returns words[] with start and end; captions take words as text, start, end. A short Python map skips speech-to-text and burns your exact script.
- Turn STT words into paragraphs: break at pauses over 1.2 seconds
Sume STT returns a flat text string. Use word start and end times to break it into paragraphs at long pauses. Python, no extra API call or cost.
- AI music API in TypeScript: generate, wait and save an MP3
Call Sume's Music Router from TypeScript: generateMusicRouter, waitForJob, then download the audio artifact. $0.125 per track, Google lists Lyria 3.5 at $0.08.
- TypeScript exhaustive switch over a Sume run's terminal status
A run ends as completed, failed, canceled or skipped, and the last two send no webhook. Use a never check so a new status fails the build.
Written by Sume