voices_create says processing: call again with voice_id

When voices_create returns processing, the voice is still being made. Call it again with the same voice_id to wait. A resubmit starts a second paid generation.

5 min readSume
All posts

If voices_create returns an outcome of processing, the voice has been queued but is not ready. Call voices_create again with the same voice_id (and optionally the job_id) to wait for it. Never send the original persona or audio again: that submits a second paid generation. Speak only once the row is ready and has a cartesia_voice_id.

What each outcome means

voices_create outcomes from packages/studio-agent-core/src/voice-mcp-text.ts (read 2026-10-05)
OutcomeWhat it meansYour next call
readyThe row exists with a cartesia_voice_idtts_create with that id as voice.id
processingThe job is still runningvoices_create again with voice_id
failedThe job ended in errorReport the actual error; do not retry blindly

Why the id matters

The first response gives you an identifier for the work in flight. The second call with that identifier attaches to the same job and waits. A second call without it looks like a brand-new request, and the tool is marked as a paid create. The tool's own instructions say it directly: resume with voice_id, never resubmit.

This is the same rule as for audio detach in sync mode: an unfinished job is polled, not repeated.

The same-turn flow

When someone asks to make a voice and speak in the same message, the agent should not ask them to come back later. The loop is: voices_create until ready, read cartesia_voice_id, call tts_create with that value as voice.id, and the transcript is what the person asked to say (or a short default if they only asked for a voice).

If you are driving the tools from your own agent, copy that loop and cap the number of resume calls so a stuck job fails clearly rather than spinning.

Where the voice ends up

The row lands in Assets, then Voices, in the workspace. Teammates can @ it later and voices_list returns it with a sample_url. A voice that failed does not appear as a speakable row, and a pending row is not speakable yet, so check the status before you build a script around it.

A defensive loop

If you drive this from code, store the voice_id from the first response before doing anything else. On each processing outcome, call again with that id and wait a short interval. Stop after a sensible number of attempts and report the last status. Do not create a second voice while the first is unresolved.

This keeps one generation per voice and one bill per voice.

Related posts

More in Agents

All Agents posts

Written by Sume