HeyGen video translation API: modes, languages, cost

HeyGen's video translation API is POST /v3/video-translations: a video URL plus target languages, in Speed or Precision mode, billed per minute.

5 min readSume
All posts

HeyGen's video translation API is POST /v3/video-translations: you send a public video URL and a list of output_languages, and HeyGen translates the speech, by default clones the original speaker's voice, and lip-syncs the mouth to the new audio. It runs in Speed mode by default or in Precision mode, and translate_audio_only: true stops before lip sync and returns the new audio track.

Every HeyGen fact below comes from HeyGen's own developer docs and help center, read on 2026-09-28 and listed under Sources. The last section covers what Sume's API can and cannot do for the same job.

How do I call the HeyGen video translation API?

HeyGen's quick start has three steps. Authentication is an x-api-key header.

  • List valid target languages with GET /v3/video-translations/languages.
  • Submit POST /v3/video-translations with video ({ "type": "url", "url": … } or an asset_id from POST /v3/assets), output_languages, mode and title. The URL must be publicly accessible.
  • Several languages in one request return one ID per language in video_translation_ids.
  • Poll GET /v3/video-translations/{video_translation_id} until the status is completed or failed, or pass callback_url to get a webhook instead.
  • A completed translation carries video_url, plus srt_caption_url and vtt_caption_url when those caption files exist. HeyGen says to check that each caption field is present before reading it.
curl --request POST \
  --url 'https://api.heygen.com/v3/video-translations' \
  --header "x-api-key: $HEYGEN_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "video": { "type": "url", "url": "https://example.com/video.mp4" },
    "output_languages": ["Spanish", "French"],
    "mode": "speed",
    "title": "My Translated Video"
  }'

What is the difference between Speed and Precision?

Both modes run one lip sync engine. HeyGen says Precision "uses avatar inference and multiple models to re-render the speaker's mouth movements" and takes longer to process; it recommends Precision for faces with significant movement, side angles or occlusions, and Speed for quick drafts and batch jobs. Audio-only translation skips lip sync, and in Precision mode it skips avatar inference.

From HeyGen's Speed and Precision guides and API pricing explained, read 2026-09-28. Rates apply to API use only.
OptionHow to request itAPI rateCredits per min
Speed, audio only (no lip sync)"translate_audio_only": true$0.5711
Speed, lip sync"mode": "speed" (the default)$0.8116
Precision, lip sync"mode": "precision"$1.5030

Which languages does HeyGen video translation support?

HeyGen's developer page says it translates videos "into 175+ languages". Its help center keeps a list of languages for audio-only translation and lip sync, many split by region, such as French (Belgium, Canada, France, Switzerland). For API calls, the valid values are whatever GET /v3/video-translations/languages returns; the batch guide points there too.

  • brand_glossary_id applies a glossary: do_not_translate_terms stay as written and forced_translations get their fixed replacement in every target language.
  • speaker_num sets the number of speakers (auto by default), and disable_music_track strips background music.
  • A preset stock voice instead of the cloned speaker is an Enterprise feature, turned on by request.

How do batch translations and billing work?

POST /v3/video-translations/batches takes up to 100 items and answers 202 Accepted with one batch_id. A payload with several output_languages expands to one item per language, and the 100-item cap counts the expanded items. Each item is processed on its own, so one bad source does not fail the rest, and an Idempotency-Key header makes retries safe (batches guide).

Billing uses HeyGen's pay-as-you-go API credits, bought in US dollars by any user, including free users, with no plan needed. Translations are charged on the source duration even when enable_dynamic_duration lets the output length vary. HeyGen API pricing covers the other rows of HeyGen's price table.

Can Sume translate an existing video?

Not in one call. Sume's lip sync routes, VEED Fabric 1.0 (POST /v1/veed/fabric-1.0) and MiniMax H3 Max Lip Sync, take a still image plus audio and return a talking clip (Models), so neither redraws the mouth in existing footage. The pieces are separate on the API reference: STT 1.0 transcribes audio of up to 10 minutes, and TTS 1.0 speaks a transcript in the language you set. Avatar 1.0 speaks English only in current code. How to dub an avatar video with lip sync and Translate a video's voiceover by API show those paths.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume