Audition a saved Sume voice with sample_url before a paid TTS take

voices_list and voices_get return a sample_url for each ready workspace voice. Play it first, then spend on tts_create with the cartesia_voice_id.

5 min readSume
All posts

Call voices_list, then play the sample_url on the voice you like before you pay for a take. Each ready row returns its id (voi_...), slug, gender, language, sample_url, source_kind and cartesia_voice_id. Pass the cartesia_voice_id as voice.id to tts_create. Listening costs nothing; a TTS take costs $0.0475 per 1,000 characters.

What the list returns

voices_list returns ready workspace reference voices. You can filter by language (ISO 639-1) and by slug, and limit defaults to 100. Rows for voices that are still being made are not speakable yet, so they should not be used.

voices_list row fields from packages/studio-agent-core (read 2026-10-05)
FieldUse
idThe voi_... row id
slugA stable short name to filter by
languageThe language the voice is tuned for
sample_urlA short audio sample to play
source_kindWhether the voice was generated or cloned
cartesia_voice_idThe value to pass as voice.id

Check one voice closely

voices_get inspects one ready voice by voi_ id, Cartesia id, slug or name. It returns the same sample, the source kind, and the language the voice is optimised for. Read the language before you send a script: a script in a different language from the voice's can return a language-mismatch response instead of audio.

On the API, a mismatch is a 409 tts_voice_language_mismatch. In the agent tool it comes back as tts_voice_language_warning with a confirmation step, and you confirm only if you meant it.

A short audition routine

Do this once per voice and keep the result.

  • List voices for the script's language.
  • Play the sample_url for the top two or three.
  • Send one 100-character line to each finalist, which costs about half a cent each, and compare them on your real script's tone.
  • Record the winner's cartesia_voice_id in your project notes or Format so every episode uses it.

You do not need a library row

The library is a convenience, not a gate. A raw voice id that you already know can go straight into tts_create without a lookup, and a library miss does not mean you must create anything. Use the audition flow when you have several candidates and want to hear them; skip it when you already have an id. Freezing the preset afterwards keeps every episode consistent.

Reading the language first

Filter voices_list by the language of your script before you listen. A voice tuned for another language may sound fine on its sample but be rejected for your script, and a mismatch costs you a round trip. Tagalog is tl; fil is normalised to it.

Pick the voice, note its cartesia_voice_id, and you are ready to render.

Related posts

More in Agents

All Agents posts

Written by Sume