MAI-Voice-2.1 or Sume TTS for a presenter ad read: 3 prices

For a 700-character ad read, MAI-Voice-2.1 is about 1.5 cents, Flash 1 cent and Sume TTS 4 cents. Fabric needs the audio as a public URL.

5 min readSume
All posts

On price alone, MAI-Voice-2.1 is cheaper than Sume TTS. For a 700-character ad read, MAI-Voice-2.1 at $22 per million characters is about 1.5 cents, MAI-Voice-2.1-Flash at $15 per million is about 1 cent, and Sume TTS is 4 cents. But Sume does not list MAI-Voice in its catalog, so the choice is really between buying the voice elsewhere and bringing the file in, or generating it inside the same job pipeline as your picture and video.

The prices

The MAI figures come from a third-party release tracker (read 2026-10-08), which lists MAI-Voice-2.1 as text-to-speech in 23 languages and 26 locales at $22 per million characters, and the Flash variant at $15 per million with a vendor-stated 150 ms end-to-end latency. Microsoft's news page (read 2026-10-08) names MAI-Voice-2.1; I give no other MAI specifications. For Sume, the catalog lists Sonic 3.6 at $0.000038 per character, and Sume bills list x 1.25, rounded up to the cent per job. The 700-character length is my assumption for a short ad read.

Cost of a 700-character read, as of 2026-10-08
OptionRateCalculationCost
MAI-Voice-2.1$22 per million characters700 x 22 / 1,000,000$0.0154
MAI-Voice-2.1-Flash$15 per million characters700 x 15 / 1,000,000$0.0105
Sume TTS (sonic-3.6)$0.000038 per character list700 x 0.000038 x 1.25 = 0.03325, rounded up$0.04

What the cents do not cover

  • Sume rounds each TTS job up to a whole cent, so many tiny reads cost more than one long one. A MAI read has no such rounding in the tracker figures.
  • Sume accepts up to 20,000 characters per TTS request.
  • A finished Sume TTS job records model_id, voice, language, output_format, generation_config and speed, so you can reproduce the next line.
  • The 150 ms figure is for Flash as stated by the vendor. Sume TTS runs as a job, not a streaming call.

The talking-face constraint

If the voice will drive a face, the Sume path is a still plus an audio clip on VEED Fabric 1.0, with a public HTTPS audio_url and a measured duration_seconds. Video models do not lip-sync to a generated voice. A MAI-Voice file works if you host it at a public HTTPS URL. Sume documents that media URLs must be public HTTPS, so avoid private or localhost addresses. Measure the file length yourself, since Fabric needs the duration in seconds.

Which to pick

Pick MAI-Voice if the lowest per-character cost, or a language on its list, decides the matter, and you are happy to host the file. Pick Sume TTS if you want the read, the still and the talking clip in one wallet, one job log and one reproducible record. At 700 characters the difference is about 2.5 cents, so for a single ad it rarely decides the question. Check the job result fields before you rely on them.

Check it before you run it

Every figure above is a catalog list price times 1.25, rounded up to the cent, as of 2026-10-08. Catalogs change, so before a large batch, read the current model entry in the docs and recompute the one line that matters for your case. Write the arithmetic next to the job in your own notes: list rate, seconds or characters, multiplier, rounding. If the result differs from the wallet charge by more than a cent, the catalog entry has changed, and the docs page is the place to find out why.

Run one small job first. Submit a single request with the pinned model id and the settings in the tables, poll the returned polling_url until it finishes, and compare the charge with your estimate. Then scale up. Using a pinned id for the test matters, because sume/auto never names the family, so you cannot tie its charge to the row you priced.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume