Voice actor or AI voice for a 30-second ad? What an API covers

Hire a voice actor or use AI text to speech for a 30-second ad: what Sume's TTS API does, what it cannot do, and the $0.315 audio bill for three takes.

4 min readSume
All posts

Use an AI voice when the 30-second ad is a draft, a variant or one of many cuts, and hire a voice actor when a human performance, a specific recognizable voice or a legal sign-off on a talent contract is the point. On Sume the AI route is text to speech through POST /v1/tts-1.0/generate, and a 450-character script costs $0.03 per take at the rate in the catalog.

This post compares the two routes on what each can do, not on taste. The Sume figures come from the OpenAPI file and the Timeline and Music Router docs listed in the sources. We state no human rates because we did not read a vendor page for them; get a quote for those.

What a 30-second script means in characters

A spoken ad runs at roughly 150 words a minute, so 30 seconds is about 75 words, or 400 to 450 characters with spaces. Sume bills TTS per transcript character, and spaces and punctuation count, so the script you paste is the meter. The request cap is 20,000 characters, which is far above an ad, and synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded.

Pace is a request field, not a studio session. generation_config.speed takes 0.6 to 1.5, volume takes 0.5 to 2.0, and emotion is a free-text guide up to 64 characters. If the take lands at 33 seconds, you change the script or the speed and run it again.

Side by side

Rates and limits below are from the Sume OpenAPI file and catalog (read 2026-10-11); the human column describes the work, not a price.

Voice actor versus Sume TTS for one 30-second ad (Sume values read 2026-10-11)
QuestionVoice actorSume TTS
Cost of one extra takeDepends on the contract; ask$0.03 for 450 characters
Script change after approvalA new sessionEdit the text, submit again
Specific, recognizable human voiceYes, with a talent agreementOnly voices on your workspace or the Voices library
Own-voice cloneRecord itCloning is app-only, not in the API
Pronunciation fixSay it againpronunciation_dict_id, or respell in the script
Word timings for captionsTranscribe afterwardstimestamps.words: true on the job
LanguageWhoever you hirelanguage field, set for every non-English script

What the API route actually looks like

You select a voice with a top-level avatar_id or avatar_handle (use an avatar whose voice.status is ready) or with voice.id, which must be a voice UUID or a voi_ library id. A voice name from another vendor is rejected with a 400 before any credit is reserved. Pass timestamps.words: true if you plan to burn captions from the timings, and ask for wav if the file will feed a lip-sync or a later join, because the default is mp3 at 44.1 kHz and 128 kbps.

The job is async by default. mode: sync waits at most 30 seconds, then returns the current state and you poll the status URL. The result carries a Sume-hosted audio artifact. That artifact can go straight into a Timeline render as the spine.

import asyncio, os
import httpx

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}", "Idempotency-Key": "ad-take-001"}

async def main():
    body = {
        "transcript": "Meet the lamp that follows your day. Warm at dawn, calm at dusk. Order today.",
        "avatar_handle": "@your_avatar",
        "generation_config": {"speed": 1.0},
        "timestamps": {"words": True},
        "output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100},
        "mode": "async",
    }
    async with httpx.AsyncClient(timeout=60) as c:
        r = await c.post(f"{API}/v1/tts-1.0/generate", json=body, headers=H)
        print(r.status_code, r.json()["data"]["request_id"])

asyncio.run(main())

The audio bill for a finished 30-second ad

Take three takes of the 450-character script, one Music Router bed and a one-minute Timeline render (a 30-second output bills as one started minute). The arithmetic uses the catalog prices: TTS $0.0475 per 1,000 characters rounds to $0.03 for 450 characters, Music is a flat $0.125 per generation, and Timeline renders at $0.10 per output minute.

  • The bed prompt should end with an instruction for no vocals, so it does not fight the narration, as the Music Router examples do.
  • Duck the bed under the voice with soundtrack.duck_db in the Timeline render, which needs a real voice spine.
Three-take audio bill for one 30-second ad (Sume catalog, read 2026-10-11)
LineQuantityPriceSubtotal
TTS take, 450 characters3$0.03$0.09
Music Router bed1$0.125$0.125
Timeline render, 1 started minute1$0.10$0.10
Total$0.315

When to hire a person anyway

Pick a human when the ad needs a performance the script cannot carry, such as dry timing or a read that must sound like a particular spokesperson, or when a brand has a talent agreement that names the voice. Pick TTS when you will test several angles, localize into more languages, or change the script after review. A common split is TTS for the first ten variants and a booked session for the one that wins.

Sume does not hold a human voice for you: voice cloning is app-only, so an API pipeline uses the voices that already exist on your workspace. Check the voice list before you promise a client a specific sound, and read the AI voiceover for ads page for the cut-by-cut workflow, or the preflight checklist before you render.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume