Quiz audio: 100 questions at 140 characters, one job, sliced per line
A hundred 140-character quiz questions are 14,000 characters: $0.21 on MAI Flash, $0.31 on MAI-Voice-2.1 and $0.67 on Sume, which can slice one job by sentence.

A quiz with 100 questions of 140 characters is 14,000 characters of speech, and it costs $0.21 on MAI-Voice-2.1-Flash, $0.31 on MAI-Voice-2.1 and $0.67 on Sume TTS 1.0, where one request can also be cut into one clip per question. At list rates that is 14,000 characters: $0.308 on MAI-Voice-2.1 ($22 per 1M characters), $0.210 on MAI-Voice-2.1-Flash ($15 per 1M), and $0.665 on Sume TTS 1.0 ($47.50 per 1M, or $0.0475 per 1,000 characters).
The inputs are assumptions for this scenario: 100 questions of 140 characters each, run once; the 20,000-character Sume request limit is not reached. Microsoft's prices are from its announcement page and Sume's from the API reference and pricing page, all read 2026-10-05. The totals are rate times characters, before any per-job rounding.
What the bill looks like
Per-clip pricing hides the real choice here, which is one job for the whole round or one job per question.
| Model | Rate per 1M characters | One run | Total over one quiz round |
|---|---|---|---|
| MAI-Voice-2.1 | $22.00 | $0.308 | $0.308 |
| MAI-Voice-2.1-Flash | $15.00 | $0.210 | $0.210 |
| Sume TTS 1.0 | $47.50 | $0.665 | $0.665 |
One job or a hundred
One job keeps the narrator's delivery consistent across the round and needs one idempotency key. A hundred jobs let you regenerate a single bad question without touching the others. Sume supports the first with a way back to the second: ask for sentence segmentation and you get one clip per question from the same job.
- Terminate every question with a full stop or question mark; sentences are grouped on terminal punctuation.
- Use a WAV or raw container:
emit_audioslices are only produced for those, and each slice has its ownaudio_url. - Segments are gapless, so
segment[i].endequalssegment[i+1].start, with a default 70 ms lead after the last word. - Microsoft's pages do not state a per-request limit for MAI-Voice-2.1; Flash's announcement lists 45 seconds of audio, which a 100-question round would exceed.
Run it on Sume
This joins the questions into one transcript, requests slices, and prints the first five clip URLs. Each question stays one numbered segment.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def speak(text, key, **extra):
body = {"transcript": text, "language": "en", **extra}
body.setdefault("avatar_handle", os.environ.get("SUME_AVATAR_HANDLE", ""))
r = requests.post(f"{API}/v1/tts-1.0/generate", timeout=60,
headers={**H, "Idempotency-Key": key}, json=body)
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(s["next_poll_after_seconds"] or 2)
if s["sume_status"] != "completed":
raise RuntimeError(s["sume_status"])
res = requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
return res
questions = [f"Question {i}: what is the capital of country number {i}?" for i in range(1, 101)]
res = speak(" ".join(questions)[:20000], "quiz-round-7", timestamps={"words": True},
segmentation={"mode": "sentence", "emit_audio": True},
output_format={"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100})
for seg in res["segments"][:5]:
print(seg["index"], round(seg["start"], 2), seg["audio_url"])
When MAI is the better pick
If you already run on Microsoft Foundry and want the lowest per-character price, MAI-Voice-2.1-Flash is $0.46 cheaper on this round. Sume is the better choice if the next step is cutting each question into a video and the per-sentence slices save you a splitting stage.
Sources
Related posts
More in Pricing
- Quote a 30-second AI video clip: Seedance 2.5 per clip and per minute
How to quote a 30-second Seedance 2.5 clip on Sume: price at 480p, 720p and 1080p, a per-minute equivalent, a retake allowance and the rounding rule.
- Re-render your top old videos first under a fixed $3 budget
Rank clips from a retired video pipeline by views per dollar of Sume re-render cost, then fill a fixed budget greedily. Python, arithmetic shown.
- Re-render a gpt-image-1 library before Oct 23: cost on 2.5
gpt-image-1 shuts down Oct 23, 2026. Re-rendering 1,000 images on GPT Image 2.5 at 1024x1024 costs $7.35 at low, $16.46 at medium and $65.85 at high on Sume.
- Recast a 2-minute video: four 30 s jobs, $45 at 768p on Sume
One Recast job takes 5-30 s of source, no shot over 15 s. A 2-minute video is four jobs: $45.00 at 768p, $67.50 at 1080p. How to split and cost it.
Written by Sume