A/B test two voices on one ad script: cost of two takes
Voice both takes of a 1,000-character ad with two Sume voices for $0.095. Compare length and sound, then render only the winner for $0.10 a minute.

To A/B test two voices on one ad, voice the identical script with each and listen before you render any video. On Sume a 1,000-character script costs $0.0475 per voice, so two takes cost $0.095 and ten candidate voices cost $0.475 (API reference). Render only the winner. A one-minute render is another $0.10.
Hold everything else still
A voice test is only fair if nothing else changes. Use the same transcript, language, generation_config and output_format for both takes. Change the voice selector alone: avatar_handle for each candidate. Key each job to the voice, so a rerun returns the same take instead of billing a new one.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, json=body, timeout=60,
headers={**H, "Idempotency-Key": key})
r.raise_for_status()
job = r.json()["request_id"]
while True:
s = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30).json()
if s.get("terminal"):
break
time.sleep(s.get("next_poll_after_seconds") or 3)
res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
res.raise_for_status()
return res.json()
SCRIPT = open("ad.txt").read().strip()
for handle in ["voice-a", "voice-b"]:
res = run("/v1/tts-1.0/generate", {
"transcript": SCRIPT,
"avatar_handle": handle,
"output_format": {"container": "wav", "encoding": "pcm_s16le",
"sample_rate": 44100},
}, f"ab-{handle}-v1")
print(handle, res.get("duration_seconds"), res.get("audio_url"))What to compare
The job gives you the length, and your ears do the rest. Two voices reading the same 1,000 characters will not run to the same duration, and that matters when the ad has a fixed slot. A voice that runs several seconds longer may force a cut in the copy.
- Length:
duration_secondsagainst the slot. - Pace and clarity: play both on a phone speaker, since most short video is heard on one.
- Name and number lines: look at how each voice says the offer and the price.
- Language fit: for non-English ads, make sure each voice is meant for the language. Mismatches raise
tts_voice_language_warning.
Why test more voices now
New models make more voices to choose from, and Microsoft's MAI-Voice-2.1 adds voice cloning with consent guardrails on top (Microsoft AI, read 2026-10-04). The test cost here is mostly your own listening time. Judge on a blind basis: ask a teammate to rank the files without knowing which voice made which.
After the winner
Freeze the winner's settings in a JSON preset, as in the preset guide, then render the final video once. Run the same test again when a new voice ships, using the same script, so the comparison stays fair.
Budget for a bigger audition
Five voices on a 1,000-character script cost $0.2375, and each voice also gets its own duration_seconds to compare against the slot. Add the final render at $0.10 per output minute and a full test of five candidates plus one finished one-minute video is $0.3375. Keep the losing takes: a rejected voice for an ad can be the right one for the next campaign.
Sources
Related posts
More in Use cases
- AB 853 from Jan 2027: keep originals, not just uploads
California AB 853 adds platform provenance duties from Jan 1, 2027. Store the original Sume outputs you generate so you can answer provenance questions later.
- Agency approval loop for AI video: 480p proof, then the final on Sume
Show clients a cheap 480p Seedance 2.5 proof, get sign-off, then pay for the 1080p final. Cost table for a 30 s clip: $8.06 proof, $42.65 final, plus rounds.
- AI action figure box from a selfie: a gpt-image-2.5 reference edit
Make the AI action-figure-in-a-box image from one selfie: send it as an input reference to gpt-image-2.5, then letter the name on the box in code.
- AI animated short film pipeline: shots, Timeline, and a score
Higgsfield published a guide to AI animated shorts. Here is the job as a Sume API pipeline: shots from /v1/videos, a Timeline cut and a Music Router score.
Written by Sume