Recipe video narration: 30 dishes a month on MAI Flash vs Sume
Eight steps at 220 characters is 1,760 characters a recipe. Thirty recipes a month cost $0.79 on MAI Flash, $1.16 on MAI-Voice-2.1 and $2.51 on Sume.

A recipe video with eight 220-character steps is 1,760 characters of narration, and thirty of them a month cost $0.79 on MAI-Voice-2.1-Flash, $1.16 on MAI-Voice-2.1 and $2.51 on Sume TTS 1.0. At list rates that is 52,800 characters: $1.16 on MAI-Voice-2.1 ($22 per 1M characters), $0.792 on MAI-Voice-2.1-Flash ($15 per 1M), and $2.51 on Sume TTS 1.0 ($47.50 per 1M, or $0.0475 per 1,000 characters).
The inputs are assumptions for this scenario: eight steps of 220 characters per recipe and thirty recipes a month. Microsoft's prices are from its announcement page and Sume's from the API reference and pricing page, all read 2026-10-05. The totals are rate times characters, before any per-job rounding.
What the bill looks like
Short step-by-step narration is where a per-character price is easiest to forecast: the character count is fixed once the recipe card is written.
| Model | Rate per 1M characters | One run | Total over 30 recipes |
|---|---|---|---|
| MAI-Voice-2.1 | $22.00 | $0.039 | $1.16 |
| MAI-Voice-2.1-Flash | $15.00 | $0.026 | $0.792 |
| Sume TTS 1.0 | $47.50 | $0.084 | $2.51 |
Why per-step jobs help a recipe edit
Each step in a recipe video is usually cut to its own shot, so the voice needs to line up with the picture rather than run as one block. Sume can return word timings and sentence-level audio slices from a single job, which gives you one clip per step without a separate splitting tool.
- Set
timestamps.wordsto true andsegmentation.modetosentenceto get gaplesssegments[]with start and end times. - Slices need a WAV or raw container; with MP3 you get the timings but no per-segment audio URL.
- If one step changes, regenerate just that sentence as its own job: the price follows the characters, so a 220-character fix is about one cent.
- Microsoft's model page lists emotion control and instant voice matching for MAI-Voice-2.1; Sume has an optional
generation_config.emotionstring and picks voices by avatar handle or voice id.
Run it on Sume
This sends one recipe as a single WAV job with sentence segmentation, then prints the start time of each step. Run it once per recipe and keep the job id with the recipe record.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def speak(text, key, **extra):
body = {"transcript": text, "language": "en", **extra}
body.setdefault("avatar_handle", os.environ.get("SUME_AVATAR_HANDLE", ""))
r = requests.post(f"{API}/v1/tts-1.0/generate", timeout=60,
headers={**H, "Idempotency-Key": key}, json=body)
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(s["next_poll_after_seconds"] or 2)
if s["sume_status"] != "completed":
raise RuntimeError(s["sume_status"])
res = requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
return res
steps = ["Heat the oil in a wide pan.", "Add the onions and stir for three minutes."]
text = " ".join(steps)
res = speak(text, "recipe-0042", timestamps={"words": True},
segmentation={"mode": "sentence"},
output_format={"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100})
for seg in res["segments"]:
print(seg["index"], seg["start"], seg["audio_url"])
When MAI is the better pick
MAI-Voice-2.1-Flash wins on price for a short-step format like this, and Microsoft markets Flash for latency-sensitive uses. A recorded recipe video is not latency sensitive, so the lower rate is the main reason to choose it. Choose Sume when the step clips feed straight into a timeline and caption render in the same account.
Sources
Related posts
More in Pricing
- Reconcile Black Friday video spend per SKU from the run usage receipt
Each finished Format run has a usage object with billable, debited, held and refunded amounts. Sum them per SKU from your own ledger, not the wallet total.
- Sume timeline bills whole minutes: a 61-second Reel costs twice a 60
Timeline 1.0 charges $0.10 per output minute, rounded up. A 60 s Reel bills one minute, a 61 s Reel two, and a 181 s Reel four. Cut to the edge of the minute.
- Restaurant menu audio: 60 dishes at 90 characters, MAI vs Sume cost
Reading a 60-dish menu aloud is 5,400 characters: about 8 cents on MAI-Voice-2.1-Flash, 12 cents on MAI-Voice-2.1 and 26 cents on Sume TTS 1.0.
- Review a draft: Sume caption job $0.20, or transcript $0.01/min
A standalone caption job is $0.20 flat for clips up to 60 seconds. A video-inspect transcript adds $0.01 per audio minute. Which one to use on a draft.
Written by Sume