2,000 product-name audio clips: 300,000 characters, MAI vs Sume

A catalog of 2,000 spoken product lines at 150 characters costs $4.50 on MAI Flash, $6.60 on MAI-Voice-2.1 and $14.25 on Sume, which cuts one file per line.

5 min readSume
All posts

Voicing 2,000 product lines of about 150 characters is 300,000 characters: $4.50 on MAI-Voice-2.1-Flash, $6.60 on MAI-Voice-2.1, and $14.25 on Sume TTS at list. The difference is $9.75 against Flash. What Sume adds for that money is one job that can hand back a separate, sample-exact audio file for every line.

The bill

The shape of the work matters as much as the price. A catalog needs one clip per SKU, not one long track, and the way you cut the job decides how many requests you make. The table below counts only characters, so the file handling that follows it is where the real difference shows.

2,000 product lines at 150 characters: character cost (read 2026-10-05)
OptionCharactersList cost
MAI-Voice-2.1-Flash ($15 per 1M characters)300,000$4.50
MAI-Voice-2.1 ($22 per 1M characters)300,000$6.60
Sume TTS ($0.0475 per 1,000 characters)300,000$14.25

Cutting the catalog into jobs

Sume TTS takes up to 20,000 characters in one request, so 133 lines of 150 characters (19,950 characters) fit in a job. With 2,000 lines that is 16 jobs. Ask for timestamps.words: true and segmentation: {mode: "sentence", emit_audio: true}, write each product line as a sentence that ends in a period, and each sentence comes back as its own slice with an audio_url.

Slices need a wav or raw container. With mp3 you get the segment timings and no per-segment file, which defeats the purpose here. Each job must also stay inside the 1,200-second audio limit, or it fails with tts_duration_exceeded, so spot-check one full job's duration before you queue all 16.

Where the file count lands

On Flash the natural unit is one call per line, because a call tops out at 45 seconds of audio. That is 2,000 requests, each needing its own handling and file naming. The price is lower, but the output is 2,000 separate responses to store, where Sume gives you 16 job results with the files already hosted.

Name each slice by its sentence index. The segment index is zero-based and stable, so a lookup table from SKU to job id and index is all you need to rebuild the set after a copy change.

What the table leaves out

Microsoft's figures are the launch prices from 2026-10-01 (Microsoft AI, read 2026-10-05): $22 per 1M characters for MAI-Voice-2.1 and $15 per 1M characters for MAI-Voice-2.1-Flash, which makes up to 45 seconds of audio per call. Sume's figure is its list rate per transcript character, and on Sume spaces and punctuation count toward usage, so count characters of the final text.

The cheaper per-character price does not make the Microsoft route a drop-in swap. A Sume TTS request is an async job whose result carries a hosted audio_url, duration, optional word timings and sentence slices, and that hosted file is what Timeline audio and renders accept. Count the extra work of storing and joining files outside Sume before you call the gap a saving.

Check language coverage first. Microsoft lists 23 languages and 26 locales for MAI-Voice-2.1, and Sume takes a language field that you must set for every non-English transcript. A cheaper model that does not speak your target language is not an option, so settle that before you compare totals.

Pilot first

Run one job of 20 lines first and read the result before the batch. Confirm that each segment text matches its SKU, that the slices play cleanly, and that the duration is far below 1,200 seconds. Use dry_run on the MCP tool to preview the cost and max_spend_usd to cap it, and give every job its own idempotency_key so a retry never bills a line twice.

If a few names are read wrongly, fix them in a pronunciation dictionary and re-run only the affected jobs. Re-running one job of 133 lines costs about $0.95 at list.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume