AI voice generator for games: voice every line as a file

An AI voice generator for games turns each dialogue line into an audio file in a character's voice. Voices, engine formats, batch runs and cost.

5 min readSume
All posts

An AI voice generator for games turns each line of your dialogue script into an audio file in a character's voice, so NPC barks, quest lines and cutscene dialogue can be voiced in a batch and re-voiced when the script changes. It is an offline asset step: you generate the files ahead of time and ship them with the game.

On Sume each line is one text to speech request to POST /v1/tts-1.0/generate, billed at $0.0475 per 1,000 characters plus a 5.5% agent fee by default. The fields come from the TTS schema in the Sume API reference (see API reference), and batch behavior from Jobs and results and Generation admission, read on 2026-09-29.

How do I give each character its own voice?

Map every character to one voice and send it on each of that character's lines. A request selects the voice with voice.id (a TTS voice UUID or a Voices library id, voi_ plus 32 hex) or with a ready avatar's avatar_id / avatar_handle. New character voices are made in the Sume app under Assets → Voices, where you can clone from a recording or describe a person and have a voice generated; there is no API route for that step. AI voice generator with a prompt covers describing one.

If a voice is cloned from a real actor, the Sume Terms of Service say you represent that you have permission to use any person's voice you submit. That is what the terms say, not legal advice.

What audio format can I get for my game engine?

Set output_format to what your engine imports. The default is MP3 at 44,100 Hz and 128 kbps; for uncompressed game assets, ask for WAV:

From the TTS 1.0 output_format schema in the Sume API reference, read 2026-09-29.
FieldValues
containermp3, wav, raw
sample_rate (Hz)8000, 16000, 22050, 24000, 44100, 48000
encoding (wav/raw)pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw
bit_rate (mp3)32000, 64000, 96000, 128000, 192000

How do I voice hundreds of lines without duplicates?

Loop over your dialogue table and send one request per line. Build the Idempotency-Key from the line id and its script revision: a retry after a timeout then returns the original job instead of billing a second one. Reuse a key only with the same body; a changed line gets a new key. metadata stores your line id on the job.

Jobs past your workspace's concurrency limit are accepted as queued while queue capacity remains; when the queue is full, submits fail with 429 queue_full. In current code a same-key retry replays that refusal, so after the queue drains, resubmit that line with a new key. Use mode: "webhook" with a webhook_url, or poll the job's status, then download each file and name it by line id.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: blacksmith-greet-012-r3" \
  -d '{
    "transcript": "Back again? That blade will not sharpen itself.",
    "voice": { "mode": "id", "id": "9f3e2b1c-6a4d-4c8e-9f1a-2b3c4d5e6f70" },
    "generation_config": { "speed": 0.95, "emotion": "gruff" },
    "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 48000 },
    "metadata": { "line_id": "blacksmith-greet-012" },
    "mode": "async"
  }'

Can I direct the performance of each line?

Partly. generation_config takes speed (0.6 to 1.5), volume (0.5 to 2.0) and emotion, a free-form guide of up to 64 characters. The schema documents no list of emotion values, so treat it as a hint and listen to the result. Set language for any non-English line; omitted, it defaults to English. For a scene with several characters joined into one file, see Text to speech with multiple voices.

How much does AI game voice acting cost, and what can't it do?

Speech is billed per character at $0.0475 per 1,000 characters; spaces and punctuation count, and each retake is billed again. The line counts and lengths below are example assumptions.

  • No live voice: TTS 1.0 is an async job with polling or a webhook, not streaming, so it can't voice player-driven dialogue at runtime.
  • One request takes up to 20,000 characters, and audio longer than 1,200 seconds fails with tts_duration_exceeded.
Computed from the Text to speech row on API pricing ($0.0475 per 1,000 characters), read 2026-09-29. Before the 5.5% agent fee.
ExampleLinesCharactersPrice
Short game: 3 characters20012,000$0.57
Indie RPG: 20 characters2,000160,000$7.60
Larger game: 60 characters10,000900,000$42.75

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume