HeyGen text to speech API: endpoint, limits and price

HeyGen's API has a text to speech endpoint, POST /v3/voices/speech: up to 5,000 characters per call, catalog voices, word timestamps, $0.12 a minute.

5 min readSume
All posts

Yes: HeyGen's API has a text to speech endpoint that returns audio without making a video. POST /v3/voices/speech takes 1–5,000 characters and a voice_id from HeyGen's Starfish-compatible voices, and returns an audio_url, the audio's duration and optional word timestamps. HeyGen's API pricing lists this Starfish speech at $0.12 per minute. Professional voice clones use a separate HeyGen Voice endpoint.

Everything about HeyGen below comes from its developer docs and its API pricing article, read on 2026-09-28 and linked in each caption. The Sume section at the end comes from the TTS 1.0 schema in the Sume API reference.

What does a HeyGen text to speech request take?

Send the request with your HeyGen API key in the X-Api-Key header. A successful call answers 200 with the audio_url in the response body.

From HeyGen's Generate Speech reference and Text to Speech guide, read 2026-09-28.
FieldWhat HeyGen's docs say
text (required)1–5,000 characters; pause tags take the time in seconds, not milliseconds
voice_id (required)A stock, designed or instant-clone voice that the Starfish engine supports; list them with GET /v3/voices?engine=starfish
input_typetext (default) or ssml
speed0.5–2.0, default 1
language / localeBase code such as en, auto-detected when omitted; a BCP-47 locale such as pt-BR overrides it
Responseaudio_url and duration, plus optional request_id and word_timestamps (each word with start and end in seconds)

How do HeyGen's professional voice clones speak?

A professional clone runs on HeyGen Voice, not Starfish, so it has its own endpoints and terms, per the HeyGen Voice and HeyGen Voice Speech pages:

  • It is a paid feature: each professional voice takes a purchased voice clone slot, trained from 1–10 recordings of one speaker totaling at least 20 minutes.
  • POST /v3/models/audio/tts waits and returns one mono 44.1 kHz WAV; POST /v3/models/audio/tts/stream streams ordered audio parts as Server-Sent Events.
  • text is 1–5,000 characters and language is required. Each endpoint allows 30 requests per minute per workspace member.
  • An instant clone, made from a single recording in minutes, runs on Starfish and works with POST /v3/voices/speech like any catalog voice.

How much does HeyGen text to speech cost?

HeyGen's API pricing article lists Speech on Starfish at $0.12 per minute (2 credits). HeyGen Voice synthesis costs 0.6 API credits per generated minute, on top of the voice slot. How HeyGen's pay-as-you-go API credits work is in HeyGen API pricing.

Can I use the speech in a HeyGen avatar video?

Yes. HeyGen's Audio to Video guide says the audio_url from POST /v3/voices/speech can be passed straight into audio_url on POST /v3/videos, or you can skip the audio step and send script plus voice_id. The separate TTS call helps when you reuse one narration across several videos. Audio to avatar AI covers the general pattern.

What if I only need speech from a TTS API?

If you only need audio, a general TTS API also works. Sume's TTS 1.0 is one option, with different limits:

  • POST /v1/tts-1.0/generate takes a transcript of up to 20,000 characters, with the voice set by voice.id or by an avatar's voice through avatar_id or avatar_handle.
  • Set language for every non-English transcript; omitted, it defaults to English. timestamps.words: true returns word timings.
  • It runs as an asynchronous job with polling or a webhook, not a stream.
  • It costs $0.0475 per 1,000 characters, spaces and punctuation included, plus a 5.5% agent fee by default. Text to speech API has the full request.

Sources

Related posts

More in Models

All Models posts

Written by Sume