Cartesia acceptable use: the explicit consent clause for clones
Cartesia's policy allows only your own voice or others' with explicit consent. What that means for a voice you clone on Sume, which runs on Sonic.

Cartesia's acceptable use policy puts the consent duty on the person submitting audio: you may submit your own voice, or recordings of others with explicit consent, and you are solely responsible for having the rights and consents for that input. It also bars impersonating another person, including celebrities, and impersonating political candidates or government officials. The wording is on Cartesia's policy page (read 2026-10-02).
Sume's text to speech runs on Cartesia Sonic models, so the same expectation is the safe baseline for any clip you upload to Sume's Voices library.
The clauses that matter
The page does not say that synthetic audio must be labelled, and it does not single out deceased people. It is a rights-and-impersonation policy, not a disclosure policy, so labels still come from the law and the platform where you publish.
| Topic | What the policy says |
|---|---|
| Whose audio | Your own, or others with explicit consent |
| Who is responsible | You, for all necessary rights and consents for the input |
| Impersonation | Not allowed, including celebrities and businesses |
| Elections | Impersonating candidates or officials, or election disinformation, not allowed |
| Disclosure of synthetic audio | No explicit clause on the page |
Why this applies to a Sume clone
Sume's TTS Router catalog lists Cartesia Sonic ids such as sonic-3.6, and library voices carry a Cartesia voice id. Cartesia's 2026 changelog (read 2026-10-02) shows sonic-3.6 as the current model from 27 August 2026. Sume's Voices library in the app has a create dialog that takes a name, a gender, a language and an uploaded or microphone-recorded clip. That dialog has no consent-recording step, and the public API has no route that creates a clone.
That means the dialog will not stop you from uploading someone else's clip. Your release is the control.
A one-page consent record
A short written record covers most of what the policy expects.
- Name the speaker, the date and the project the voice may be used for.
- State that synthetic speech from the voice may be published, and where.
- Say how the speaker can withdraw, and what happens to existing audio.
- Store the signed record next to the voice id in your own files.
- Do not clone anyone who is not reachable to sign, including a public figure.
What to do next
Check the policy page again before a campaign, since it can change. For the Sume side of the clone flow, read Voice cloning API, and for what Sonic 3.6 changed, Cartesia Sonic 3.6 API.
Sources
Related posts
More in Comparisons
- Cartesia Ink STT hours per plan vs Sume STT $0.60 per hour
Cartesia plans include about 9 to 741 hours of Ink speech-to-text. That works out near $0.40 to $0.54 an hour; Sume STT is $0.60 an hour. Concurrency compared.
- Cheapest 720p AI video per second: Veo, xAI, Sume
At 720p, Veo 3.1 lists $0.05 to $0.40 per second, xAI lists grok-imagine-video at $0.05, and Sume lists Grok Imagine Video 1.5 at $0.0125. Read 2026-10-01.
- Claude directory drops MCPB: remote vs local server, and Sume
Claude's directory no longer accepts MCPB desktop extensions. A remote HTTPS server needs no package; here is what that means for Sume.
- Claude Message Batches 100,000 requests vs a Sume bulk run of 100
A Claude Message Batch holds 100,000 requests or 256 MB. A Sume bulk run queues 1 to 100 Format runs with a concurrency window of 1 to 16. Plan accordingly.
Written by Sume