A/B test two voices on one ad script: cost of two takes

Voice both takes of a 1,000-character ad with two Sume voices for $0.095. Compare length and sound, then render only the winner for $0.10 a minute.

4 min readSume
All posts

To A/B test two voices on one ad, voice the identical script with each and listen before you render any video. On Sume a 1,000-character script costs $0.0475 per voice, so two takes cost $0.095 and ten candidate voices cost $0.475 (API reference). Render only the winner. A one-minute render is another $0.10.

Hold everything else still

A voice test is only fair if nothing else changes. Use the same transcript, language, generation_config and output_format for both takes. Change the voice selector alone: avatar_handle for each candidate. Key each job to the voice, so a rerun returns the same take instead of billing a new one.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(path, body, key):
    r = requests.post(API + path, json=body, timeout=60,
                      headers={**H, "Idempotency-Key": key})
    r.raise_for_status()
    job = r.json()["request_id"]
    while True:
        s = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30).json()
        if s.get("terminal"):
            break
        time.sleep(s.get("next_poll_after_seconds") or 3)
    res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
    res.raise_for_status()
    return res.json()

SCRIPT = open("ad.txt").read().strip()
for handle in ["voice-a", "voice-b"]:
    res = run("/v1/tts-1.0/generate", {
        "transcript": SCRIPT,
        "avatar_handle": handle,
        "output_format": {"container": "wav", "encoding": "pcm_s16le",
                          "sample_rate": 44100},
    }, f"ab-{handle}-v1")
    print(handle, res.get("duration_seconds"), res.get("audio_url"))

What to compare

The job gives you the length, and your ears do the rest. Two voices reading the same 1,000 characters will not run to the same duration, and that matters when the ad has a fixed slot. A voice that runs several seconds longer may force a cut in the copy.

  • Length: duration_seconds against the slot.
  • Pace and clarity: play both on a phone speaker, since most short video is heard on one.
  • Name and number lines: look at how each voice says the offer and the price.
  • Language fit: for non-English ads, make sure each voice is meant for the language. Mismatches raise tts_voice_language_warning.

Why test more voices now

New models make more voices to choose from, and Microsoft's MAI-Voice-2.1 adds voice cloning with consent guardrails on top (Microsoft AI, read 2026-10-04). The test cost here is mostly your own listening time. Judge on a blind basis: ask a teammate to rank the files without knowing which voice made which.

After the winner

Freeze the winner's settings in a JSON preset, as in the preset guide, then render the final video once. Run the same test again when a new voice ships, using the same script, so the comparison stays fair.

Budget for a bigger audition

Five voices on a 1,000-character script cost $0.2375, and each voice also gets its own duration_seconds to compare against the slot. Add the final render at $0.10 per output minute and a full test of five candidates plus one finished one-minute video is $0.3375. Keep the losing takes: a rejected voice for an ad can be the right one for the next campaign.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume