HeyGen video translation API: modes, languages, cost
HeyGen's video translation API is POST /v3/video-translations: a video URL plus target languages, in Speed or Precision mode, billed per minute.

HeyGen's video translation API is POST /v3/video-translations: you send a public video URL and a list of output_languages, and HeyGen translates the speech, by default clones the original speaker's voice, and lip-syncs the mouth to the new audio. It runs in Speed mode by default or in Precision mode, and translate_audio_only: true stops before lip sync and returns the new audio track.
Every HeyGen fact below comes from HeyGen's own developer docs and help center, read on 2026-09-28 and listed under Sources. The last section covers what Sume's API can and cannot do for the same job.
How do I call the HeyGen video translation API?
HeyGen's quick start has three steps. Authentication is an x-api-key header.
- List valid target languages with
GET /v3/video-translations/languages. - Submit
POST /v3/video-translationswithvideo({ "type": "url", "url": … }or anasset_idfromPOST /v3/assets),output_languages,modeandtitle. The URL must be publicly accessible. - Several languages in one request return one ID per language in
video_translation_ids. - Poll
GET /v3/video-translations/{video_translation_id}until the status iscompletedorfailed, or passcallback_urlto get a webhook instead. - A completed translation carries
video_url, plussrt_caption_urlandvtt_caption_urlwhen those caption files exist. HeyGen says to check that each caption field is present before reading it.
curl --request POST \
--url 'https://api.heygen.com/v3/video-translations' \
--header "x-api-key: $HEYGEN_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"video": { "type": "url", "url": "https://example.com/video.mp4" },
"output_languages": ["Spanish", "French"],
"mode": "speed",
"title": "My Translated Video"
}'What is the difference between Speed and Precision?
Both modes run one lip sync engine. HeyGen says Precision "uses avatar inference and multiple models to re-render the speaker's mouth movements" and takes longer to process; it recommends Precision for faces with significant movement, side angles or occlusions, and Speed for quick drafts and batch jobs. Audio-only translation skips lip sync, and in Precision mode it skips avatar inference.
| Option | How to request it | API rate | Credits per min |
|---|---|---|---|
| Speed, audio only (no lip sync) | "translate_audio_only": true | $0.57 | 11 |
| Speed, lip sync | "mode": "speed" (the default) | $0.81 | 16 |
| Precision, lip sync | "mode": "precision" | $1.50 | 30 |
Which languages does HeyGen video translation support?
HeyGen's developer page says it translates videos "into 175+ languages". Its help center keeps a list of languages for audio-only translation and lip sync, many split by region, such as French (Belgium, Canada, France, Switzerland). For API calls, the valid values are whatever GET /v3/video-translations/languages returns; the batch guide points there too.
brand_glossary_idapplies a glossary:do_not_translate_termsstay as written andforced_translationsget their fixed replacement in every target language.speaker_numsets the number of speakers (auto by default), anddisable_music_trackstrips background music.- A preset stock voice instead of the cloned speaker is an Enterprise feature, turned on by request.
How do batch translations and billing work?
POST /v3/video-translations/batches takes up to 100 items and answers 202 Accepted with one batch_id. A payload with several output_languages expands to one item per language, and the 100-item cap counts the expanded items. Each item is processed on its own, so one bad source does not fail the rest, and an Idempotency-Key header makes retries safe (batches guide).
Billing uses HeyGen's pay-as-you-go API credits, bought in US dollars by any user, including free users, with no plan needed. Translations are charged on the source duration even when enable_dynamic_duration lets the output length vary. HeyGen API pricing covers the other rows of HeyGen's price table.
Can Sume translate an existing video?
Not in one call. Sume's lip sync routes, VEED Fabric 1.0 (POST /v1/veed/fabric-1.0) and MiniMax H3 Max Lip Sync, take a still image plus audio and return a talking clip (Models), so neither redraws the mouth in existing footage. The pieces are separate on the API reference: STT 1.0 transcribes audio of up to 10 minutes, and TTS 1.0 speaks a transcript in the language you set. Avatar 1.0 speaks English only in current code. How to dub an avatar video with lip sync and Translate a video's voiceover by API show those paths.
Sources
- HeyGen: Video Translation - Speed (read 2026-09-28)
- HeyGen: Video Translation - Precision (read 2026-09-28)
- HeyGen: Video Translations batches (read 2026-09-28)
- HeyGen Help: Video Translation languages we support (read 2026-09-28)
- HeyGen API pricing explained (read 2026-09-28)
- Models
- API reference
- Sume API reference
Related posts
More in Sume Avatar 1.0
- HeyGen vs Argil: avatar APIs, avatar creation and pricing
HeyGen and Argil both turn a script into an avatar video by API. HeyGen bills pay-as-you-go dollars per minute; Argil bills plan credits per minute.
- HeyGen vs Creatify: avatar API, ad tools and pricing
HeyGen's API centers on avatars, voice and translation, paid with pay-as-you-go credits. Creatify's centers on video ads, sold as monthly API plans.
- HeyGen vs Synthesia: API, pricing, avatars and limits
HeyGen and Synthesia both make scripted avatar videos and sell a live avatar. They differ in API billing, video length, and how custom avatars are made.
- How do AI avatars work? Face, voice, and video, step by step
An AI avatar video is assembled from parts: a face, a voice to match it, and motion that makes the face say your script. How each step works on Sume.
Written by Sume