voices_create says processing: call again with voice_id
When voices_create returns processing, the voice is still being made. Call it again with the same voice_id to wait. A resubmit starts a second paid generation.

If voices_create returns an outcome of processing, the voice has been queued but is not ready. Call voices_create again with the same voice_id (and optionally the job_id) to wait for it. Never send the original persona or audio again: that submits a second paid generation. Speak only once the row is ready and has a cartesia_voice_id.
What each outcome means
| Outcome | What it means | Your next call |
|---|---|---|
| ready | The row exists with a cartesia_voice_id | tts_create with that id as voice.id |
| processing | The job is still running | voices_create again with voice_id |
| failed | The job ended in error | Report the actual error; do not retry blindly |
Why the id matters
The first response gives you an identifier for the work in flight. The second call with that identifier attaches to the same job and waits. A second call without it looks like a brand-new request, and the tool is marked as a paid create. The tool's own instructions say it directly: resume with voice_id, never resubmit.
This is the same rule as for audio detach in sync mode: an unfinished job is polled, not repeated.
The same-turn flow
When someone asks to make a voice and speak in the same message, the agent should not ask them to come back later. The loop is: voices_create until ready, read cartesia_voice_id, call tts_create with that value as voice.id, and the transcript is what the person asked to say (or a short default if they only asked for a voice).
If you are driving the tools from your own agent, copy that loop and cap the number of resume calls so a stuck job fails clearly rather than spinning.
Where the voice ends up
The row lands in Assets, then Voices, in the workspace. Teammates can @ it later and voices_list returns it with a sample_url. A voice that failed does not appear as a speakable row, and a pending row is not speakable yet, so check the status before you build a script around it.
A defensive loop
If you drive this from code, store the voice_id from the first response before doing anything else. On each processing outcome, call again with that id and wait a short interval. Stop after a sensible number of attempts and report the last status. Do not create a second voice while the first is unresolved.
This keeps one generation per voice and one bill per voice.
Related posts
More in Agents
- Weekly new-model watcher as a Sume schedule: schema without links
A scheduled watcher that lists new AI models fails if its output schema holds vendor page URLs. Use host-only fields and read each run with the actions API.
- Weekly trend video schedule stops early: the $1.00 default cap
A Sume schedule with no spend cap runs with $1.00 of generation per run, so a weekly video run can stop after a clip. Set the cap to fit the plan.
- Which Sume key scope starts a Format, schedule or agent run
formats:write starts a Format run, actions:write a schedule run, agent_completions:write an agent run. Older keys lack them, and service keys can't start some.
- Why the Sume API hides a schedule's instructions text
GET /v1/actions returns the cron, model, cap and schema of a Sume schedule but not its instructions text. Read and edit instructions in the dashboard.
Written by Sume