voices_create persona: 20 characters minimum, one variant in chat
To invent a Sume voice from words, voices_create needs a persona of 20+ characters or invent:true. Chat makes one variant by default, three only if you ask.

voices_create makes a new workspace voice from a text description. In generate mode it needs a persona of at least 20 characters, or invent: true to let the agent write one. In chat it returns one voice by default and three only when you ask for variants: 3. The voice is saved to Assets, then Voices, and is usable as voice.id in tts_create.
Two modes, one tool
The tool either generates a voice from a description or clones one from audio. The two are exclusive: a clone request sends audio_url and nothing from the generate side, and sending both audio_url and persona, invent or variants is invalid.
- **Generate:**
persona(at least 20 characters) orinvent: true, plus the transcript language. - **Clone:** an https
audio_url. For a video, run audio detach first to get a clean audio file. - Optional for both:
nameandgender. Gender defaults to female.
Writing a persona that clears the bar
Twenty characters is short, but a persona that short gives the generator little to work with. A usable one says age range, register, pace and setting in one sentence. For example: "Calm woman in her thirties, warm and unhurried, like a museum audio guide."
If you give no persona, or ask the agent to pick one at random, it invents one. That is the invent: true path, and it is meant for the case where you want a voice quickly and do not mind what it is.
One variant or three
The agent tool defaults to one variant. The Assets page in the app offers three to pick from. Ask for three in chat only when you want to compare, since each variant is a paid generation. The tool description tells the agent to set variants: 3 only if you ask.
Because the call is paid, expect the agent to confirm before running it when the spend is not trivial.
After the call
A voice is not speakable until the row is ready and carries a cartesia_voice_id. If the call returns processing, the agent calls again with the same voice_id (see what to do when voices_create says processing). When ready, use cartesia_voice_id as voice.id on the next tts_create in the same turn.
The row stays in Assets, then Voices, so teammates in the workspace can reference it later.
What not to put in a persona
Keep it to how the voice sounds and the role it plays. Do not name a real person or ask for a celebrity voice; if you want a specific person's voice, clone it from a sample you have the right to use. A persona describes a fictional speaker, and the generator treats it as such.
If the first voice is close but not right, change one attribute in the persona (older, slower, warmer) and generate again.
Related posts
More in Agents
- voices_create says processing: call again with voice_id
When voices_create returns processing, the voice is still being made. Call it again with the same voice_id to wait. A resubmit starts a second paid generation.
- Weekly new-model watcher as a Sume schedule: schema without links
A scheduled watcher that lists new AI models fails if its output schema holds vendor page URLs. Use host-only fields and read each run with the actions API.
- Weekly trend video schedule stops early: the $1.00 default cap
A Sume schedule with no spend cap runs with $1.00 of generation per run, so a weekly video run can stop after a clip. Set the cap to fit the plan.
- Which Sume key scope starts a Format, schedule or agent run
formats:write starts a Format run, actions:write a schedule run, agent_completions:write an agent run. Older keys lack them, and service keys can't start some.
Written by Sume