Blind-test Sonic 3.6 against 3.5 on your own script

A vendor's blind-test percentage is not yours. Render the same lines with two catalog versions through the TTS Router, shuffle them, and let listeners vote.

6 min readSume
All posts

How do you check whether a newer TTS version really sounds better? Render the same lines with both versions, shuffle the files, and have listeners pick without labels. Sume's TTS Router takes an explicit model, so one script and one voice can produce both takes.

Cartesia's blog (read 2026-10-04) lists Sonic-3.6, announced 2026-08-27, with listeners preferring it in up to 93 percent of blind head-to-head tests across fifteen locales. That is the vendor's test set. Your script, voice and language are the only ones that count for your product.

List the catalog first

GET /v1/tts-router/models is the source of truth. At the time of reading, the contract names sonic-3.6, sonic-3.5, sonic-3, sonic-latest (an alias for sonic-3.6) and sonic-preview, a beta that does not work with pro voice clones and fails with voice_model_mismatch. Pin the two concrete ids you compare, not the alias, so the test stays reproducible when the alias moves.

Render both takes

The router requires model. Use one voice id and the same lines for both, with an idempotency key that includes the model so retries do not mix them.

import os, requests

API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
VOICE = os.environ["SUME_VOICE_ID"]
lines = ["Thanks for calling. How can I help?", "Your order ships Friday."]
for model in ["sonic-3.6", "sonic-3.5"]:
    for i, text in enumerate(lines):
        r = requests.post(API + "/v1/tts-router/generate",
            headers={**H, "Idempotency-Key": f"ab-{model}-{i}"},
            json={"model": model, "transcript": text,
                  "voice": {"mode": "id", "id": VOICE},
                  "output_format": {"container": "wav"}})
        r.raise_for_status()
        print(model, i, r.json()["data"]["job"]["id"])

Run the vote

Download each finished file, give them random names, and keep the mapping in a file listeners never see. Ask each listener, per line, which one they prefer or if they cannot tell. Twenty listeners and ten lines gives you two hundred votes, enough to see a large difference and not enough to claim a small one.

Count ties. A high share of cannot-tell means the upgrade is not audible for your use, and you can choose by cost or stability. Record the winning model_id from the job so the production request pins it, as described in Jobs and results.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume