Use a saved workspace voice in Sume chat: @ mention it in the turn
Tag a saved voice with @ in a Sume chat turn and the agent prefers that voice for speech. Here is what it reads from the voice and what a mismatch looks like.

In a Sume chat, type @ and pick a saved voice from Assets, then Voices, and the agent treats it as your choice for speech in that turn. The agent reads the voice's sample_url, source_kind and language through voices_get, then calls tts_create with its cartesia_voice_id as voice.id. Teammates in the workspace can @ the same voice.
How the mention is used
The tool guidance for voices_get says it inspects a ready saved voice and that the read includes @voice mentions. It also says that when a user tags a voice with @, the agent should prefer it for speech in the turn. So the tag is a strong hint, and the agent is still the one that calls tts_create.
Because the voice is a workspace row, it persists. You do not re-describe it each time; you tag it.
What it checks
The voice has a language it is tuned for. If your script is in a different language, tts_create returns a language warning with a confirmation step rather than silently producing odd audio. Confirm only if you meant it. The raw API equivalent is 409 tts_voice_language_mismatch.
fil and tl are treated as matching, and the voice tool normalises Tagalog to tl. For other languages, keep the voice and the script on the same code.
Ways to set up a voice you can tag
- **Describe it:** ask the agent to make a voice from a persona of at least 20 characters. See voices_create persona rules.
- **Clone it:** give an https audio sample. For a video, extract the audio first.
- **Use an existing id:** a raw voice id works in
tts_createwithout a library row, so tagging is optional.
Practical habits
Name voices by role, not by date, so a teammate knows which to tag. Play the sample_url before you commit to a voice for a series; the audition routine takes a few minutes. A pending row is not speakable yet, so wait for it to become ready, and do not resubmit its creation.
Each TTS take costs $0.0475 per 1,000 characters. The tag does not change that, but it does change who can repeat the take with the same voice.
When the tag is not enough
If the agent speaks in a different voice, say which voice you mean and tag it again. A voice that is still pending is not speakable, so wait for it to be ready. And if you give a raw voice id in the message, the agent can use it directly without a library row.
The tag changes the voice, not the script: the transcript is still whatever you asked the agent to say.
Related posts
More in Agents
- voices_create persona: 20 characters minimum, one variant in chat
To invent a Sume voice from words, voices_create needs a persona of 20+ characters or invent:true. Chat makes one variant by default, three only if you ask.
- voices_create says processing: call again with voice_id
When voices_create returns processing, the voice is still being made. Call it again with the same voice_id to wait. A resubmit starts a second paid generation.
- Weekly new-model watcher as a Sume schedule: schema without links
A scheduled watcher that lists new AI models fails if its output schema holds vendor page URLs. Use host-only fields and read each run with the actions API.
- Weekly trend video schedule stops early: the $1.00 default cap
A Sume schedule with no spend cap runs with $1.00 of generation per run, so a weekly video run can stop after a clip. Set the cap to fit the plan.
Written by Sume