HeyGen voice clone API: 20+ minutes, paid slot, no Sume clone
HeyGen's professional voice clone needs 1-10 recordings totaling 20+ minutes and a paid slot. Sume's TTS tool selects voices and has no clone upload.

HeyGen's professional voice clone, added in September 2026, trains from 1-10 recordings of one speaker totaling 20+ minutes, and each cloned voice occupies a purchased slot. Sume's TTS tool does not train a voice: it takes a voice selector, and its description lists no voice-cloning or voice-upload widget.
HeyGen figures below are from its API changelog; Sume statements are from the TTS tool description in the hosted MCP server code and the MCP tools and gates page, read 2026-09-30.
What does HeyGen's professional clone require?
The changelog says you train a dedicated voice adapter on the HeyGen Voice model, using 1-10 recordings of the same speaker totaling 20 minutes or more. POST /v3/models/audio/voices creates the voice (or retrains one when voice_id is supplied), and you poll the voice until its status is ACTIVE.
| Item | HeyGen changelog |
|---|---|
| Training audio | 1-10 recordings, same speaker, 20+ minutes total |
| Slot | Each voice occupies a purchased professional voice clone slot |
| Trainings | Five per slot per monthly billing period |
| Synthesis | 0.6 API credits per generated minute |
| Over the slot limit | The voice returns voice_expired until a slot is added |
| Short recording | Instant Clone remains available to all API plans |
What does Sume's TTS tool take instead?
The tts_create tool description asks for exactly one transcript input plus a voice selector: voice.id, or avatar_id / avatar_handle, in which case Sume resolves that avatar's ready TTS voice. It states that there is no voice-cloning or voice-upload widget, so a recording of your speaker is not an input to that tool.
Spend works per character: the tool is paid per character, requires an idempotency_key, and supports dry_run to preview cost and max_spend_usd to cap it.
Which one fits my project?
If the requirement is that the voice be a specific real person trained from their own recordings, the HeyGen route described above is built for that, and its slot and credit terms apply. If you need narration in a chosen catalog voice, or the voice already attached to a Sume avatar, selecting by id is enough and nothing needs training. The Sume docs describe no voice-cloning route; see voice cloning API for the wider picture.
What should I check before committing to a clone?
Confirm you hold consent from the speaker for the recordings you upload, and read the vendor's own terms; this post is not legal advice. Then work out the slot count you need, because HeyGen's changelog ties each voice to one slot. For avatar videos where the voice comes with the avatar, see Generate avatar video for how scripts become speech.
Sources
Related posts
More in Pricing
- Lyria 3 Clip Preview 30-second price vs Sume Music fixed price
Google lists Lyria 3 Clip Preview (30s) at a per-song price. Sume Music charges one fixed price per accepted generation, whatever length the prompt asks for.
- MiniMax H3 Max reference images: 2 free, then per image
MiniMax includes the first 2 images free on H3 Max (5 on H3) and charges per extra image. On Sume, check the catalog for accepted reference types.
- MiniMax H3 Max input video: billed per second, by resolution
MiniMax bills H3 Max input video by duration, at a rate set by output resolution. Sume lists video references for H3 Max and reserves list x 1.25.
- MiniMax H3 Max price per second: promo ends Sept 30, Sume bills list
fal lists MiniMax H3 Max at a 50% promo through September 30. Sume's rate is the standard list times 1.25, with 5 s and 15 s clip totals by resolution.
Written by Sume