HeyGen instant voice clone audio in Sume? Only Sume-hosted audio
HeyGen's new instant voice clone can output TTS audio, but Sume's talking-still routes take only Sume-hosted audio up to 10 MB. What it blocks and what works.

You cannot send a HeyGen instant-clone audio file straight to Sume's talking-still routes: audio_url there must be on Sume's media host, and the docs do not give a public upload path for arbitrary audio. The workable path is to generate the voice with Sume's own TTS and feed that file to the still-plus-audio route.
HeyGen's changelog, read 2026-10-11, lists a new "mode": "instant" on POST /v3/models/audio/voices that creates a clone from a single recording, active within seconds. It also says POST /v3/models/audio/tts accepts instant voices with an expressiveness_boost parameter, and that the feature is free during the preview period. The Sume side is from the OpenAPI schema for VEED Fabric 1.0, the models overview and Media inputs.
What Sume's still-plus-audio routes require
Two Sume routes turn a still and audio into a talking clip: VEED Fabric 1.0 at POST /v1/veed/fabric-1.0 and MiniMax H3 Max Lip Sync at POST /v1/minimax/h3-max/lip-sync. Both take the same body: exactly one visual source (image_url, or a ready avatar by handle or id), an audio_url, and a measured duration_seconds. For Fabric, the schema describes audio_url as a public HTTPS URL on the Sume media host, typically a TTS segment, with non-Sume hosts rejected and a maximum of 10 MB; duration_seconds runs from 1 to 300.
The lip-sync alternative has a tighter window. MiniMax H3 Max Lip Sync takes the same body but accepts audio of 5 to 14.8 seconds, and its output length follows the audio. So a 4-second line fits neither route's guard rails the same way, and a 40-second script has to be sliced before it reaches the H3 route. Fabric accepts a duration_seconds up to 300, which is why long narration usually goes there.
The lip-sync alternative has a tighter window. MiniMax H3 Max Lip Sync takes the same body but accepts audio of 5 to 14.8 seconds, and its output length follows the audio. So a 4-second line is below the H3 floor but acceptable to Fabric, and a 40-second script has to be sliced before it reaches the H3 route. Fabric accepts a duration_seconds up to 300, which is why long narration usually goes there.
Where an outside audio file would come from
The table shows the paths I could find in the docs for getting audio onto the Sume host.
| Source of the audio | Accepted by Fabric audio_url? | Why |
|---|---|---|
| Sume TTS output | Yes | It is hosted on the Sume media host |
| A Sume timeline or detached audio file | Yes | Also a Sume-hosted artifact |
| HeyGen instant-voice TTS file | Not directly | Hosted by HeyGen, not Sume |
| A TikTok or Instagram video's audio | Via media import | Import is limited to those two platforms |
| Any other uploaded file | No public route found | Signed upload URLs are not in the public contract |
What works instead
Generate the narration with Sume TTS, then pass the returned artifact URL as audio_url. Keep the audio under 10 MB and measure its length for duration_seconds. If you want a script-driven avatar and do not need your own recording, skip the audio step entirely and use Avatar 1.0, which speaks a script you send.
For a clone of your own voice, Sume's voice cloning is in the app, not in the API, so an automated pipeline cannot create one. The existing voice cloning roundup covers what to use.
Whichever route you use, measure the audio length yourself. The duration_seconds field is used to reserve credits at admission, and Sume's docs describe it as the measured length of the audio, so rounding it far below the true length is not a safe shortcut.
Check before you plan around a preview
HeyGen describes the instant clone as free during a preview period, and preview terms can change, so do not build a cost model on it without rechecking its page. On the Sume side, the media-import route fetches TikTok and Instagram links only and bills a fixed estimate per accepted import, so it is not a general upload path either.
If your workflow absolutely needs an externally recorded voice on a talking still, today's honest answer is that the Sume public API cannot take it. Re-read the Media inputs page when you plan, since the supported inputs are listed there.
Sources
Related posts
More in Comparisons
- Is AI video 'Full HD' native or upscaled? Kandinsky 6.0 vs Sume rows
Kandinsky 6.0 renders 864x480 and adds Full HD with a plug-in. Sume's H3 Max 1080p is a latent refinement from 768p. What that means for price and detail.
- Is Kandinsky 6.0 Video on Sume? No: query the catalog by need
Sume lists no Kandinsky 6.0 id. Map each Kandinsky feature to a Sume request field, then filter GET /v1/videos/models for sound and a 5 s duration.
- Kandinsky 6.0 lip-syncs in one pass; on Sume a talking face is Fabric
Kandinsky 6.0 Video makes speech and lip-sync inside the clip. Sume's docs send on-camera speech to Fabric or H3 Max Lip Sync with your audio, not a video id.
- Kandinsky 6.0 Pro: 292 s a clip on an H100, 100 clips take 8.1 hours
Kandinsky's published timing is 292 seconds per 5-second Pro HD clip on an H100. That is 8.1 hours for 100 clips, set against hosted per-clip prices on Sume.
Written by Sume