Recipe video narration: 30 dishes a month on MAI Flash vs Sume

Eight steps at 220 characters is 1,760 characters a recipe. Thirty recipes a month cost $0.79 on MAI Flash, $1.16 on MAI-Voice-2.1 and $2.51 on Sume.

5 min readSume
All posts

A recipe video with eight 220-character steps is 1,760 characters of narration, and thirty of them a month cost $0.79 on MAI-Voice-2.1-Flash, $1.16 on MAI-Voice-2.1 and $2.51 on Sume TTS 1.0. At list rates that is 52,800 characters: $1.16 on MAI-Voice-2.1 ($22 per 1M characters), $0.792 on MAI-Voice-2.1-Flash ($15 per 1M), and $2.51 on Sume TTS 1.0 ($47.50 per 1M, or $0.0475 per 1,000 characters).

The inputs are assumptions for this scenario: eight steps of 220 characters per recipe and thirty recipes a month. Microsoft's prices are from its announcement page and Sume's from the API reference and pricing page, all read 2026-10-05. The totals are rate times characters, before any per-job rounding.

What the bill looks like

Short step-by-step narration is where a per-character price is easiest to forecast: the character count is fixed once the recipe card is written.

Cost of 1,760 characters, and of 52,800 characters over 30 recipes, at list rates (read 2026-10-05)
ModelRate per 1M charactersOne runTotal over 30 recipes
MAI-Voice-2.1$22.00$0.039$1.16
MAI-Voice-2.1-Flash$15.00$0.026$0.792
Sume TTS 1.0$47.50$0.084$2.51

Why per-step jobs help a recipe edit

Each step in a recipe video is usually cut to its own shot, so the voice needs to line up with the picture rather than run as one block. Sume can return word timings and sentence-level audio slices from a single job, which gives you one clip per step without a separate splitting tool.

  • Set timestamps.words to true and segmentation.mode to sentence to get gapless segments[] with start and end times.
  • Slices need a WAV or raw container; with MP3 you get the timings but no per-segment audio URL.
  • If one step changes, regenerate just that sentence as its own job: the price follows the characters, so a 220-character fix is about one cent.
  • Microsoft's model page lists emotion control and instant voice matching for MAI-Voice-2.1; Sume has an optional generation_config.emotion string and picks voices by avatar handle or voice id.

Run it on Sume

This sends one recipe as a single WAV job with sentence segmentation, then prints the start time of each step. Run it once per recipe and keep the job id with the recipe record.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def speak(text, key, **extra):
    body = {"transcript": text, "language": "en", **extra}
    body.setdefault("avatar_handle", os.environ.get("SUME_AVATAR_HANDLE", ""))
    r = requests.post(f"{API}/v1/tts-1.0/generate", timeout=60,
                      headers={**H, "Idempotency-Key": key}, json=body)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(s["next_poll_after_seconds"] or 2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    res = requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
    return res

steps = ["Heat the oil in a wide pan.", "Add the onions and stir for three minutes."]
text = " ".join(steps)
res = speak(text, "recipe-0042", timestamps={"words": True},
            segmentation={"mode": "sentence"},
            output_format={"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100})
for seg in res["segments"]:
    print(seg["index"], seg["start"], seg["audio_url"])

When MAI is the better pick

MAI-Voice-2.1-Flash wins on price for a short-step format like this, and Microsoft markets Flash for latency-sensitive uses. A recorded recipe video is not latency sensitive, so the lower rate is the main reason to choose it. Choose Sume when the step clips feed straight into a timeline and caption render in the same account.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume