Audition a saved Sume voice with sample_url before a paid TTS take
voices_list and voices_get return a sample_url for each ready workspace voice. Play it first, then spend on tts_create with the cartesia_voice_id.

Call voices_list, then play the sample_url on the voice you like before you pay for a take. Each ready row returns its id (voi_...), slug, gender, language, sample_url, source_kind and cartesia_voice_id. Pass the cartesia_voice_id as voice.id to tts_create. Listening costs nothing; a TTS take costs $0.0475 per 1,000 characters.
What the list returns
voices_list returns ready workspace reference voices. You can filter by language (ISO 639-1) and by slug, and limit defaults to 100. Rows for voices that are still being made are not speakable yet, so they should not be used.
| Field | Use |
|---|---|
| id | The voi_... row id |
| slug | A stable short name to filter by |
| language | The language the voice is tuned for |
| sample_url | A short audio sample to play |
| source_kind | Whether the voice was generated or cloned |
| cartesia_voice_id | The value to pass as voice.id |
Check one voice closely
voices_get inspects one ready voice by voi_ id, Cartesia id, slug or name. It returns the same sample, the source kind, and the language the voice is optimised for. Read the language before you send a script: a script in a different language from the voice's can return a language-mismatch response instead of audio.
On the API, a mismatch is a 409 tts_voice_language_mismatch. In the agent tool it comes back as tts_voice_language_warning with a confirmation step, and you confirm only if you meant it.
A short audition routine
Do this once per voice and keep the result.
- List voices for the script's language.
- Play the
sample_urlfor the top two or three. - Send one 100-character line to each finalist, which costs about half a cent each, and compare them on your real script's tone.
- Record the winner's
cartesia_voice_idin your project notes or Format so every episode uses it.
You do not need a library row
The library is a convenience, not a gate. A raw voice id that you already know can go straight into tts_create without a lookup, and a library miss does not mean you must create anything. Use the audition flow when you have several candidates and want to hear them; skip it when you already have an id. Freezing the preset afterwards keeps every episode consistent.
Reading the language first
Filter voices_list by the language of your script before you listen. A voice tuned for another language may sound fine on its sample but be rejected for your script, and a mismatch costs you a round trip. Tagalog is tl; fil is normalised to it.
Pick the voice, note its cartesia_voice_id, and you are ready to render.
Related posts
More in Agents
- ChatGPT confirms write tools; Sume MCP also gates them by scope
ChatGPT developer mode asks before write actions. Sume's hosted MCP adds a second gate: OAuth mcp:write, an idempotency_key and an optional max_spend_usd.
- Claude 'Allow always' on a paid Sume tool: what still caps spend
Claude custom connectors let you approve a tool once and keep approving it. For a paid Sume tool, the scope, idempotency key and max_spend_usd still apply.
- Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0: cap Sume calls anyway
A benchmark score says how well Sonnet 5.5 finishes tasks, not what a Sume call may cost. Put generation_spend_cap_usd on every Agent Completion it starts.
- Clef-flash as a yes/no gate before a paid Sume Agent Completion
Cloudflare's Clef-flash returns typed answers with probabilities. Put one in front of POST /v1/agent/completions as a filter, keep the spend cap as the guard.
Written by Sume