Picture-book narration: 24 pages at 200 characters, MAI vs Sume cost

A 24-page picture book at 200 characters a page is 4,800 characters: 7 cents on MAI-Voice-2.1-Flash, 11 cents on MAI-Voice-2.1 and 23 cents on Sume TTS.

5 min readSume
All posts

A 24-page picture book with 200 characters a page is 4,800 characters of narration, costing about 7 cents on MAI-Voice-2.1-Flash, 11 cents on MAI-Voice-2.1 and 23 cents on Sume TTS 1.0; twelve books cost $0.86, $1.27 and $2.74. At list rates that is 57,600 characters: $1.27 on MAI-Voice-2.1 ($22 per 1M characters), $0.864 on MAI-Voice-2.1-Flash ($15 per 1M), and $2.74 on Sume TTS 1.0 ($47.50 per 1M, or $0.0475 per 1,000 characters).

The inputs are assumptions for this scenario: 24 pages, 200 characters per page, twelve books in a series. Microsoft's prices are from its announcement page and Sume's from the API reference and pricing page, all read 2026-10-05. The totals are rate times characters, before any per-job rounding.

What the bill looks like

Twelve books is the total column. At these sizes the narrator quality and page timing matter more than the bill.

Cost of 4,800 characters, and of 57,600 characters over 12 books, at list rates (read 2026-10-05)
ModelRate per 1M charactersOne runTotal over 12 books
MAI-Voice-2.1$22.00$0.106$1.27
MAI-Voice-2.1-Flash$15.00$0.072$0.864
Sume TTS 1.0$47.50$0.228$2.74

Page-turn timing is the feature

A read-along app needs to know when each page's audio starts and ends so the page turns on cue. That is a timing output, not a price.

  • Request timestamps.words and segmentation.mode: sentence; the gapless segments[] give each sentence a start and end you can map to a page.
  • Generate each page as its own job when pages must be replaceable; with a WAV container each job returns its own file.
  • Use generation_config.speed below 1.0, down to 0.6, for early readers.
  • Check the voice and language together: one language per request, matching the voice.

Run it on Sume

This generates one WAV per page and prints its duration from the last word's end time.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def speak(text, key, **extra):
    body = {"transcript": text, "language": "en", **extra}
    body.setdefault("avatar_handle", os.environ.get("SUME_AVATAR_HANDLE", ""))
    r = requests.post(f"{API}/v1/tts-1.0/generate", timeout=60,
                      headers={**H, "Idempotency-Key": key}, json=body)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(s["next_poll_after_seconds"] or 2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    res = requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
    return res

pages = ["Mina found a red kite in the attic.", "It was too big for her small hands."]
for i, text in enumerate(pages, 1):
    res = speak(text, f"book-03-page-{i:02d}", timestamps={"words": True},
                generation_config={"speed": 0.9},
                output_format={"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100})
    print(i, res["words"][-1]["end"], res["artifacts"][0]["url"])

When MAI is the better pick

MAI-Voice-2.1's expressive range is the draw for storytelling, and Microsoft lists granular emotion control. Sume offers a free-text emotion string in generation_config. Listen to a sample page on both before committing to a series.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume