HeyGen text to speech API: endpoint, limits and price
HeyGen's API has a text to speech endpoint, POST /v3/voices/speech: up to 5,000 characters per call, catalog voices, word timestamps, $0.12 a minute.

Yes: HeyGen's API has a text to speech endpoint that returns audio without making a video. POST /v3/voices/speech takes 1–5,000 characters and a voice_id from HeyGen's Starfish-compatible voices, and returns an audio_url, the audio's duration and optional word timestamps. HeyGen's API pricing lists this Starfish speech at $0.12 per minute. Professional voice clones use a separate HeyGen Voice endpoint.
Everything about HeyGen below comes from its developer docs and its API pricing article, read on 2026-09-28 and linked in each caption. The Sume section at the end comes from the TTS 1.0 schema in the Sume API reference.
What does a HeyGen text to speech request take?
Send the request with your HeyGen API key in the X-Api-Key header. A successful call answers 200 with the audio_url in the response body.
| Field | What HeyGen's docs say |
|---|---|
text (required) | 1–5,000 characters; pause tags take the time in seconds, not milliseconds |
voice_id (required) | A stock, designed or instant-clone voice that the Starfish engine supports; list them with GET /v3/voices?engine=starfish |
input_type | text (default) or ssml |
speed | 0.5–2.0, default 1 |
language / locale | Base code such as en, auto-detected when omitted; a BCP-47 locale such as pt-BR overrides it |
| Response | audio_url and duration, plus optional request_id and word_timestamps (each word with start and end in seconds) |
How do HeyGen's professional voice clones speak?
A professional clone runs on HeyGen Voice, not Starfish, so it has its own endpoints and terms, per the HeyGen Voice and HeyGen Voice Speech pages:
- It is a paid feature: each professional voice takes a purchased voice clone slot, trained from 1–10 recordings of one speaker totaling at least 20 minutes.
POST /v3/models/audio/ttswaits and returns one mono 44.1 kHz WAV;POST /v3/models/audio/tts/streamstreams ordered audio parts as Server-Sent Events.textis 1–5,000 characters andlanguageis required. Each endpoint allows 30 requests per minute per workspace member.- An instant clone, made from a single recording in minutes, runs on Starfish and works with
POST /v3/voices/speechlike any catalog voice.
How much does HeyGen text to speech cost?
HeyGen's API pricing article lists Speech on Starfish at $0.12 per minute (2 credits). HeyGen Voice synthesis costs 0.6 API credits per generated minute, on top of the voice slot. How HeyGen's pay-as-you-go API credits work is in HeyGen API pricing.
Can I use the speech in a HeyGen avatar video?
Yes. HeyGen's Audio to Video guide says the audio_url from POST /v3/voices/speech can be passed straight into audio_url on POST /v3/videos, or you can skip the audio step and send script plus voice_id. The separate TTS call helps when you reuse one narration across several videos. Audio to avatar AI covers the general pattern.
What if I only need speech from a TTS API?
If you only need audio, a general TTS API also works. Sume's TTS 1.0 is one option, with different limits:
POST /v1/tts-1.0/generatetakes atranscriptof up to 20,000 characters, with the voice set byvoice.idor by an avatar's voice throughavatar_idoravatar_handle.- Set
languagefor every non-English transcript; omitted, it defaults to English.timestamps.words: truereturns word timings. - It runs as an asynchronous job with polling or a webhook, not a stream.
- It costs $0.0475 per 1,000 characters, spaces and punctuation included, plus a 5.5% agent fee by default. Text to speech API has the full request.
Sources
- HeyGen Generate Speech reference (read 2026-09-28)
- HeyGen Text to Speech guide (read 2026-09-28)
- HeyGen Voices overview (read 2026-09-28)
- HeyGen Voice model (read 2026-09-28)
- HeyGen Voice Speech (read 2026-09-28)
- HeyGen Audio to Video (read 2026-09-28)
- HeyGen API pricing explained (read 2026-09-28)
- Sume API reference
- API reference
- Sume API pricing
Related posts
More in Models
- Higgsfield alternatives: studio, API and MCP compared
Higgsfield alternatives by the part you use: its studio apps, its API or its MCP link. fal, Kling, Replicate, Runway, Seedance and Sume, dated.
- Higgsfield vs Kling AI: studio or model maker?
Kling makes the Kling models and sells its own apps and API. Higgsfield is a multi-model studio that also runs Kling 3.0. Access and billing compared.
- Higgsfield AI vs Runway: apps, API, models and pricing
Higgsfield and Runway both sell a creative app and a video API billed in credits. How their models, API billing and job flow compare, from their pages.
- How does AI video generation work? Inputs, jobs, and limits
AI video generation turns a prompt, and optionally an image, into every frame of a short clip in one async job. What goes in, and why clips stay short.
Written by Sume