30 voice-over variants for an ad test cost $0.171 on Sume TTS

Six hooks, five voices, 120 characters each: 3,600 characters is $0.171 at Sume's $0.0475 per 1,000. MAI-Voice-2.1-Flash lists the same text at $0.054.

5 min readSume
All posts

Thirty voice-over variants for an A/B test, built from six hooks read by five voices at 120 characters each, total 3,600 characters and cost $0.171 on Sume TTS (3.6 x $0.0475). Voice is a small line in such a test. The expensive part is the video under it and the media spend that tests it, so cheap audio lets you test more hooks, not fewer.

Compare the rates

Sume bills by character on a flat rate. Microsoft lists MAI-Voice-2.1 at $22 per million characters and the Flash model at $15 per million (Microsoft AI, read 2026-10-09). The same 3,600 characters cost different amounts on each.

Cost of 3,600 characters of speech; Sume catalog and Microsoft AI announcement read 2026-10-09.
ServiceRate3,600 charactersNotes
Sume TTS 1.0$0.0475 per 1,000 characters$0.171Async job; avatar or voice.id selector
MAI-Voice-2.1$22 per 1M characters$0.0792Public preview, no SLA
MAI-Voice-2.1-Flash$15 per 1M characters$0.054Public preview, no SLA

Run the grid without paying twice

Build a list of 30 pairs, hook by voice. Give each request a deterministic Idempotency-Key such as hook-3-voice-2, so a retry after a timeout returns the original job instead of a second charge. Keep every other setting constant: speed, volume, language and output format. Only then is a difference in results a difference in voice or hook.

A voice selector is an avatar_id, an avatar_handle or a voice.id, as the OpenAPI reference lists. Choose five that are ready, listen once to each on the same sentence, and drop any that does not suit the brand before you spend on video.

Where the cost really sits

Thirty finished ads need thirty renders. A render is $0.10 per output minute, so a 15-second ad is billed as 1 minute, $0.10, and thirty of them are $3.00. The voice, at $0.171, is under 6 percent of that. If you attach the same voice to many clips, you can render one video per hook and let the voices differ only in the audio. Start with the hooks and drop the weak voices early.

Reading the results

Change one thing at a time. If hook 3 beats hook 1 with every voice, the hook matters more than the voice and you can stop testing voices. If one voice wins across all six hooks, keep it and spend the next round on hooks alone. With thirty variants and a small audience per cell, results are noisy, so treat a clear gap as a signal and a small one as a tie. The audio cost of a second round of the same size is another $0.171.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume