A/B test three voices on one line: parallel Sume TTS jobs in Python

Submit the same 300-character line to three voices at once with one Idempotency-Key each, poll with backoff and compare the files. The test costs $0.04275.

5 min readSume
All posts

What it costs and how it runs

Three voices on one 300-character line is 900 characters in total, which is 0.9 x $0.0475 = $0.04275 on Sume TTS. Submit the three jobs at the same time with asyncio.gather, give each its own Idempotency-Key, poll each with backoff, and listen to the three files side by side.

Pick the voice by ear on a real line from your script, not on the voice name. Voices with the same label can sound very different on your text.

Test design

Change one thing at a time. Here the only variable is the voice id; the line, language, speed and volume stay the same.

Three-voice test plan and cost at $0.0475 per 1,000 characters (read 2026-10-04)
RunVoiceCharactersCost
Avoi_A300$0.01425
Bvoi_B300$0.01425
Cvoi_C300$0.01425
Total900$0.04275

The script

Replace the three ids with ones from the voices_list MCP tool or your Assets Voices page. The keys carry the voice name, so a retry of one run cannot double bill.

import asyncio, os, httpx
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
LINE = "Our new plan starts today, and the first month is on us."
VOICES = ["voi_A", "voi_B", "voi_C"]

async def one(c, voice):
    body = {"transcript": LINE, "voice": {"id": voice}, "language": "en"}
    r = await c.post(f"{API}/v1/tts-1.0/generate", json=body,
                     headers={**H, "Idempotency-Key": f"ab-line1-{voice}"})
    r.raise_for_status()
    job, delay = r.json()["request_id"], 2
    while True:
        s = (await c.get(f"{API}/v1/jobs/{job}/status", headers=H)).json()
        if s["status"] in ("failed", "canceled"): raise RuntimeError(s)
        if s.get("result_ready"): break
        await asyncio.sleep(delay); delay = min(delay * 2, 15)
    res = (await c.get(f"{API}/v1/jobs/{job}/result", headers=H)).json()
    return voice, (res.get("result") or res)["artifacts"][0]["url"]

async def main():
    async with httpx.AsyncClient(timeout=60) as c:
        for voice, url in await asyncio.gather(*(one(c, v) for v in VOICES)):
            print(voice, url)

asyncio.run(main())

Judging the files

Listen on the device your audience will use, with the same loudness. If one voice is quieter, that is a generation_config.volume difference you can fix later (0.5 to 2), so do not reject it for level alone. Ask two people who have not seen the script which one they would trust to read it, and write down why before you decide.

  • Keep the winning voice id in one config file.
  • Freeze speed, volume and language with it so every episode matches.
  • Do not resubmit a paid job that looks slow; poll its status.
  • Delete nothing; keep the other two files as a record of the test.

Why parallel is safe here

The three jobs are independent, so running them together saves wall-clock time without changing the cost. The key names matter: ab-line1-voi_A is stable, so if your script crashes after submitting and you run it again, Sume returns the existing job instead of creating a second paid one. Change the key only when you change the text or the settings, and use a new prefix for each new line you test.

Run the script for three or four lines from different parts of your script: an opening, a number-heavy sentence and a closing call to action. A voice that wins on the opening may stumble on figures. At 300 characters a line, four lines across three voices is 3,600 characters, or $0.171, still less than a coffee.

When you have a winner, write down the voice id, the speed, the volume and the language in one place. Those four values are the whole recipe for sounding the same next month. Do the same for any second voice you keep for dialogue, so a future episode can be rebuilt from the notes alone.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume