voices_create persona: 20 characters minimum, one variant in chat

To invent a Sume voice from words, voices_create needs a persona of 20+ characters or invent:true. Chat makes one variant by default, three only if you ask.

5 min readSume
All posts

voices_create makes a new workspace voice from a text description. In generate mode it needs a persona of at least 20 characters, or invent: true to let the agent write one. In chat it returns one voice by default and three only when you ask for variants: 3. The voice is saved to Assets, then Voices, and is usable as voice.id in tts_create.

Two modes, one tool

The tool either generates a voice from a description or clones one from audio. The two are exclusive: a clone request sends audio_url and nothing from the generate side, and sending both audio_url and persona, invent or variants is invalid.

  • **Generate:** persona (at least 20 characters) or invent: true, plus the transcript language.
  • **Clone:** an https audio_url. For a video, run audio detach first to get a clean audio file.
  • Optional for both: name and gender. Gender defaults to female.

Writing a persona that clears the bar

Twenty characters is short, but a persona that short gives the generator little to work with. A usable one says age range, register, pace and setting in one sentence. For example: "Calm woman in her thirties, warm and unhurried, like a museum audio guide."

If you give no persona, or ask the agent to pick one at random, it invents one. That is the invent: true path, and it is meant for the case where you want a voice quickly and do not mind what it is.

One variant or three

The agent tool defaults to one variant. The Assets page in the app offers three to pick from. Ask for three in chat only when you want to compare, since each variant is a paid generation. The tool description tells the agent to set variants: 3 only if you ask.

Because the call is paid, expect the agent to confirm before running it when the spend is not trivial.

After the call

A voice is not speakable until the row is ready and carries a cartesia_voice_id. If the call returns processing, the agent calls again with the same voice_id (see what to do when voices_create says processing). When ready, use cartesia_voice_id as voice.id on the next tts_create in the same turn.

The row stays in Assets, then Voices, so teammates in the workspace can reference it later.

What not to put in a persona

Keep it to how the voice sounds and the role it plays. Do not name a real person or ask for a celebrity voice; if you want a specific person's voice, clone it from a sample you have the right to use. A persona describes a fictional speaker, and the generator treats it as such.

If the first voice is close but not right, change one attribute in the persona (older, slower, warmer) and generate again.

Related posts

More in Agents

All Agents posts

Written by Sume