Voice cloning API: clone once in the app, speak via the API
Sume has no API route that creates a voice clone. You clone once in the app, then send the voi_ id to the text to speech API on every call.

A voice cloning API has two parts: a call that turns an audio sample into a reusable voice, and a call that speaks text in that voice. On Sume only the second part is an API. You clone a voice once in the Sume app (Assets → Voices, from an audio file you upload or a recording), copy its voi_ id, and send that id as voice.id on POST https://api.sume.com/v1/tts-1.0/generate from your backend.
The API facts come from the TTS 1.0 schema in the Sume API reference, the OpenAPI document behind the API reference docs; the app steps come from the Sume app's current code. Both were read on 2026-09-29. The step-by-step clone walkthrough, with its upload limits, is in AI voiceover in your own voice.
Is there an API endpoint that creates the clone?
No. The public API has no route that creates a voice from a sample or from a description. Both happen in the app: the Voices page says “Clone a voice or invent one from a prompt”, and its create dialog offers “Upload or record audio to clone, or describe a person and we generate the voice.” So plan the clone as a one-time manual step per voice, and automate everything after it.
| Step | Where it happens | What you get |
|---|---|---|
| Clone a voice from an upload or a recording | Sume app, Assets → Voices | A voice in your Voices library |
| Invent a voice from a description | Sume app, Assets → Voices | A voice in your Voices library |
| Get the voice's id | The voice's Copy ID action | A Voices library id: voi_ plus 32 hex characters |
| Speak text in the voice | POST /v1/tts-1.0/generate with voice.id | A job whose result links the audio file |
How does my backend speak in the cloned voice?
Store the voi_ id in your config or database next to the voice's name, and send it on each request. Keep the API key in a server-side environment variable: Sume's authentication docs tell browser and mobile clients to call your backend, which attaches the key.
transcript: the text, up to 20,000 characters; spaces and punctuation count toward usage.voice.id: thevoi_id. Library ids are resolved to the stored voice before the job is queued.language: set it for every non-English transcript; when omitted it defaults to English.Idempotency-Keyheader: reuse the same key when you retry a submit, so the retry returns the original job instead of billing a second one.- The submit returns a job; poll
status_urluntil it is terminal, or pass awebhook_urlfor the terminal callback, then readresult_url.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: welcome-line-001" \
-d '{
"transcript": "Thanks for calling. Your order is on its way.",
"voice": { "mode": "id", "id": "voi_…" },
"language": "en",
"mode": "async"
}'What happens if the voice id is wrong?
The request fails fast. voice.id must be a TTS voice UUID or a Voices library id (voi_ plus 32 hex characters), not a voice name from another TTS product; any other shape is rejected with 400 and invalid_voice_id before a job is queued or credits are reserved. Copy the id verbatim from the app.
If you send an avatar reference (avatar_id or avatar_handle) together with voice.id, the two must match, or the request fails with 400.
What does speech in a cloned voice cost?
Speech in a cloned voice is billed as text to speech: $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, counted on the transcript's characters. The Audio section of the API rate card has three rows, music generation, speech transcription and text to speech; it lists no price for the in-app clone step itself. The output defaults to MP3 at 44,100 Hz and 128 kbps; wav and raw are the other containers.
Whose voice can I clone?
Sume's Terms of Service put this on you: “You represent that you have all rights and permissions needed for the content you submit, including permission to use any person's likeness or voice.” The acceptable-use list also says not to submit voices “that you do not have permission to use.” That is what the terms say; it is not legal advice.
Paid plans include commercial use of generated outputs “as described on the pricing page,” per the same terms; text to speech commercial use covers that question.
Sources
Related posts
More in Models
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
- Music generation API: the Sume Music Router with Lyria 3.5
Sume's Music Router turns a text prompt into a track via POST /v1/music-router/generate. sume/music-auto picks the engine, Lyria 3.5 today.
Written by Sume