30 voice-over variants for an ad test cost $0.171 on Sume TTS
Six hooks, five voices, 120 characters each: 3,600 characters is $0.171 at Sume's $0.0475 per 1,000. MAI-Voice-2.1-Flash lists the same text at $0.054.

Thirty voice-over variants for an A/B test, built from six hooks read by five voices at 120 characters each, total 3,600 characters and cost $0.171 on Sume TTS (3.6 x $0.0475). Voice is a small line in such a test. The expensive part is the video under it and the media spend that tests it, so cheap audio lets you test more hooks, not fewer.
Compare the rates
Sume bills by character on a flat rate. Microsoft lists MAI-Voice-2.1 at $22 per million characters and the Flash model at $15 per million (Microsoft AI, read 2026-10-09). The same 3,600 characters cost different amounts on each.
| Service | Rate | 3,600 characters | Notes |
|---|---|---|---|
| Sume TTS 1.0 | $0.0475 per 1,000 characters | $0.171 | Async job; avatar or voice.id selector |
| MAI-Voice-2.1 | $22 per 1M characters | $0.0792 | Public preview, no SLA |
| MAI-Voice-2.1-Flash | $15 per 1M characters | $0.054 | Public preview, no SLA |
Run the grid without paying twice
Build a list of 30 pairs, hook by voice. Give each request a deterministic Idempotency-Key such as hook-3-voice-2, so a retry after a timeout returns the original job instead of a second charge. Keep every other setting constant: speed, volume, language and output format. Only then is a difference in results a difference in voice or hook.
A voice selector is an avatar_id, an avatar_handle or a voice.id, as the OpenAPI reference lists. Choose five that are ready, listen once to each on the same sentence, and drop any that does not suit the brand before you spend on video.
Where the cost really sits
Thirty finished ads need thirty renders. A render is $0.10 per output minute, so a 15-second ad is billed as 1 minute, $0.10, and thirty of them are $3.00. The voice, at $0.171, is under 6 percent of that. If you attach the same voice to many clips, you can render one video per hook and let the voices differ only in the audio. Start with the hooks and drop the weak voices early.
Reading the results
Change one thing at a time. If hook 3 beats hook 1 with every voice, the hook matters more than the voice and you can stop testing voices. If one voice wins across all six hooks, keep it and spend the next round on hooks alone. With thirty variants and a small audience per cell, results are noisy, so treat a clear gap as a signal and a small one as a tie. The audio cost of a second round of the same size is another $0.171.
Sources
Related posts
More in Use cases
- Three 175-second Shorts from a 28-minute recording for $0.36
Cut three 175-second YouTube Shorts from one 28-minute recording on Sume: two detach ranges $0.02, transcript $0.28, three trims $0.06, total $0.36.
- Three hook variants for one avatar ad: $3.31 on Standard
Test three 6-second hooks on Standard for $3.312 in total, then render the winner as a 30-second video on Max for $16.50. Sume bills avatar video by the second.
- 3-minute Shorts and monetization: plan the cut and the price
YouTube ties Short monetization to the Shorts Feed and blocks claimed Shorts over a minute. Cut a 3-minute Short for $0.02 to $0.30.
- Make a three-minute YouTube Short with a Format and check duration_ms
YouTube counts square or vertical uploads up to three minutes as Shorts. Call a Sume Format, then test duration_ms and aspect ratio on the receipt first.
Written by Sume