HeyGen instant voice clone audio in Sume? Only Sume-hosted audio

HeyGen's new instant voice clone can output TTS audio, but Sume's talking-still routes take only Sume-hosted audio up to 10 MB. What it blocks and what works.

4 min readSume
All posts

You cannot send a HeyGen instant-clone audio file straight to Sume's talking-still routes: audio_url there must be on Sume's media host, and the docs do not give a public upload path for arbitrary audio. The workable path is to generate the voice with Sume's own TTS and feed that file to the still-plus-audio route.

HeyGen's changelog, read 2026-10-11, lists a new "mode": "instant" on POST /v3/models/audio/voices that creates a clone from a single recording, active within seconds. It also says POST /v3/models/audio/tts accepts instant voices with an expressiveness_boost parameter, and that the feature is free during the preview period. The Sume side is from the OpenAPI schema for VEED Fabric 1.0, the models overview and Media inputs.

What Sume's still-plus-audio routes require

Two Sume routes turn a still and audio into a talking clip: VEED Fabric 1.0 at POST /v1/veed/fabric-1.0 and MiniMax H3 Max Lip Sync at POST /v1/minimax/h3-max/lip-sync. Both take the same body: exactly one visual source (image_url, or a ready avatar by handle or id), an audio_url, and a measured duration_seconds. For Fabric, the schema describes audio_url as a public HTTPS URL on the Sume media host, typically a TTS segment, with non-Sume hosts rejected and a maximum of 10 MB; duration_seconds runs from 1 to 300.

The lip-sync alternative has a tighter window. MiniMax H3 Max Lip Sync takes the same body but accepts audio of 5 to 14.8 seconds, and its output length follows the audio. So a 4-second line fits neither route's guard rails the same way, and a 40-second script has to be sliced before it reaches the H3 route. Fabric accepts a duration_seconds up to 300, which is why long narration usually goes there.

The lip-sync alternative has a tighter window. MiniMax H3 Max Lip Sync takes the same body but accepts audio of 5 to 14.8 seconds, and its output length follows the audio. So a 4-second line is below the H3 floor but acceptable to Fabric, and a 40-second script has to be sliced before it reaches the H3 route. Fabric accepts a duration_seconds up to 300, which is why long narration usually goes there.

Where an outside audio file would come from

The table shows the paths I could find in the docs for getting audio onto the Sume host.

Ways to get audio to a Sume talking-still route (Sume docs and OpenAPI; HeyGen changelog read 2026-10-11)
Source of the audioAccepted by Fabric audio_url?Why
Sume TTS outputYesIt is hosted on the Sume media host
A Sume timeline or detached audio fileYesAlso a Sume-hosted artifact
HeyGen instant-voice TTS fileNot directlyHosted by HeyGen, not Sume
A TikTok or Instagram video's audioVia media importImport is limited to those two platforms
Any other uploaded fileNo public route foundSigned upload URLs are not in the public contract

What works instead

Generate the narration with Sume TTS, then pass the returned artifact URL as audio_url. Keep the audio under 10 MB and measure its length for duration_seconds. If you want a script-driven avatar and do not need your own recording, skip the audio step entirely and use Avatar 1.0, which speaks a script you send.

For a clone of your own voice, Sume's voice cloning is in the app, not in the API, so an automated pipeline cannot create one. The existing voice cloning roundup covers what to use.

Whichever route you use, measure the audio length yourself. The duration_seconds field is used to reserve credits at admission, and Sume's docs describe it as the measured length of the audio, so rounding it far below the true length is not a safe shortcut.

Check before you plan around a preview

HeyGen describes the instant clone as free during a preview period, and preview terms can change, so do not build a cost model on it without rechecking its page. On the Sume side, the media-import route fetches TikTok and Instagram links only and bills a fixed estimate per accepted import, so it is not a general upload path either.

If your workflow absolutely needs an externally recorded voice on a talking still, today's honest answer is that the Sume public API cannot take it. Re-read the Media inputs page when you plan, since the supported inputs are listed there.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume