Blind-test Sonic 3.6 against 3.5 on your own script
A vendor's blind-test percentage is not yours. Render the same lines with two catalog versions through the TTS Router, shuffle them, and let listeners vote.

How do you check whether a newer TTS version really sounds better? Render the same lines with both versions, shuffle the files, and have listeners pick without labels. Sume's TTS Router takes an explicit model, so one script and one voice can produce both takes.
Cartesia's blog (read 2026-10-04) lists Sonic-3.6, announced 2026-08-27, with listeners preferring it in up to 93 percent of blind head-to-head tests across fifteen locales. That is the vendor's test set. Your script, voice and language are the only ones that count for your product.
List the catalog first
GET /v1/tts-router/models is the source of truth. At the time of reading, the contract names sonic-3.6, sonic-3.5, sonic-3, sonic-latest (an alias for sonic-3.6) and sonic-preview, a beta that does not work with pro voice clones and fails with voice_model_mismatch. Pin the two concrete ids you compare, not the alias, so the test stays reproducible when the alias moves.
Render both takes
The router requires model. Use one voice id and the same lines for both, with an idempotency key that includes the model so retries do not mix them.
import os, requests
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
VOICE = os.environ["SUME_VOICE_ID"]
lines = ["Thanks for calling. How can I help?", "Your order ships Friday."]
for model in ["sonic-3.6", "sonic-3.5"]:
for i, text in enumerate(lines):
r = requests.post(API + "/v1/tts-router/generate",
headers={**H, "Idempotency-Key": f"ab-{model}-{i}"},
json={"model": model, "transcript": text,
"voice": {"mode": "id", "id": VOICE},
"output_format": {"container": "wav"}})
r.raise_for_status()
print(model, i, r.json()["data"]["job"]["id"])
Run the vote
Download each finished file, give them random names, and keep the mapping in a file listeners never see. Ask each listener, per line, which one they prefer or if they cannot tell. Twenty listeners and ten lines gives you two hundred votes, enough to see a large difference and not enough to claim a small one.
Count ties. A high share of cannot-tell means the upgrade is not audible for your use, and you can choose by cost or stability. Record the winning model_id from the job so the production request pins it, as described in Jobs and results.
Sources
Related posts
More in Developers
- Browser voice app that starts Sume jobs: keep the key on your server
Voice apps run in the browser over WebRTC, but Sume keys belong on a server. A route handler that holds the key, allowlists models, reuses idempotency keys.
- Sume bulk queue 404 format_run_queue_not_found: three causes
A Sume bulk queue poll returned 404 format_run_queue_not_found. The id is wrong or the queue is another owner's. How to tell which, and what to do next.
- Bulk run 409 idempotency_conflict: find the first queue by queue_id
A Sume bulk create that reuses a key with a different payload returns 409 with details.queue_id. How to read the original queue and decide what to resend.
- C2PA 2.2: file types that can carry credentials vs Sume outputs
C2PA 2.2 manifests can be embedded in JPEG, PNG, WebP, SVG, MP4, MOV and more. How that list lines up with the formats Sume image and video jobs return.
Written by Sume