OpenAI custom voices: consent phrase, 30 s sample, 20 voice cap
OpenAI custom voices need a recorded consent phrase, a sample of 30 seconds or less and sales enablement. What that means next to cloning a voice in Sume.

To create a custom voice with OpenAI you record two clips from the same person: a consent recording that reads one of OpenAI's fixed phrases, and a sample of 30 seconds or less. Access is limited to eligible customers who talk to OpenAI sales, and an organization can hold at most 20 voices, per OpenAI's custom voices guide (read 2026-10-02).
Sume works differently. You clone once in the app and then send the voice id to text to speech; there is no API route that creates a clone, and no consent recording is collected in the dialog.
What OpenAI requires
The guide says the consent recording must contain exactly one of 16 supported phrases and that any divergence from the script fails. The consent recording and the sample must come from the same person. Once created, the voice is used by passing its id in the voice parameter of the Text to Speech or Realtime API.
| Item | What the guide says |
|---|---|
| Access | Limited to eligible customers; contact sales |
| Consent recording | Exactly one of 16 supported phrases; divergence fails |
| Sample audio | 30 seconds or less |
| Sample formats | mpeg, wav, ogg, aac, flac, webm or mp4 |
| Voice cap | At most 20 voices per organization |
| GPT-Live voices | At least five seconds of actual speech and 15 transcribed text tokens |
What Sume does today
Sume's Voices library in the app has a create dialog that takes a name, a gender, a language and an uploaded or microphone-recorded clip. That dialog has no consent-recording step, and the public API has no route that creates a clone.
After a clone is ready it is a library voice. Text to speech then takes that voice by id, so the clone step and the speaking step are separate. If you need a verified consent recording stored next to the voice, Sume does not provide one; keep the signed release and the recording yourself.
Practical differences
Three differences matter when you plan a voice project.
- OpenAI ties a voice to a recorded phrase; Sume ties it to the account that uploaded the clip.
- OpenAI caps voices at 20 per organization; Sume's docs do not state a voice count, so check the Voices library before you plan a large cast.
- OpenAI voices work through OpenAI's own speech endpoints; Sume voices work through Sume's text to speech jobs, where you can also join clips with Timeline audio.
A safe workflow with either
Get a written release that names the use, the voice owner and the duration. Record the consent separately from the sample. Store both next to the voice id. Then test with a short script before a long one. For law and platform rules around AI voices, see the posts on the No Fakes Act and Tennessee's ELVIS Act.
Sources
Related posts
More in Comparisons
- OpenRouter key credit limit and 402 vs a Sume per-run spend cap
OpenRouter caps a key's credits and reports limit_remaining. A Sume Format run carries its own per-run generation cap. Where each guardrail sits and what fails.
- OpenRouter models fallback array and 3-entry limit vs Sume
OpenRouter's models array tries the next model on downtime, rate limits or moderation; fallbacks allows 3. Sume's allow_fallbacks has no effect.
- OpenRouter provider.sort and max_price vs Sume's inert sort
OpenRouter's provider.sort picks price, throughput or latency and turns off load balancing. On Sume's image route, sort is accepted and changes nothing.
- Perso AI dubbing: 10 speakers, 2-speaker lip sync, vs Sume
Perso says it detects up to 10 speakers and lip-syncs two. Sume's avatar video uses one avatar per final video. What to use for a multi-speaker dub.
Written by Sume