LLM-written script to Sume TTS: speak only the approved draft

Draft with Gemini 3.8 Flash or any LLM, approve, then send hash-keyed paragraphs to Sume TTS so unchanged lines never bill twice. $0.0475 per 1,000 characters.

5 min readSume
All posts

The cheapest way to voice a script that an LLM wrote is to keep the drafting loop out of the paid path: revise as often as you like in text, and send only the approved paragraphs to Sume TTS, each keyed by a hash of its text. Unchanged paragraphs then return the original job instead of being spoken, and billed, a second time. Sume TTS costs $0.0475 per 1,000 characters (read 2026-10-03).

Google's model page lists gemini-3.8-flash as stable, described as its "most intelligent Flash model" and positioned for long-horizon software engineering, agents and enterprise workflows, with 3.7 and 3.6 Flash labelled previous-generation (read 2026-10-03). The page I read gives no token limits or prices, and none of this article depends on them: any model that returns text can write the script, and Sume only sees the final characters.

Why hash the paragraph

Sume TTS 1.0 accepts a transcript of up to 20,000 characters, and a submit with an Idempotency-Key that repeats an earlier request returns the original job. If the key is a hash of the exact paragraph plus the voice settings, then an unchanged paragraph is a replay, and only edited paragraphs create new work.

Reusing a key for a different payload is a conflict: Sume answers 409 idempotency_conflict. That is the behaviour you want here, because a hash key can only collide with the same text. Keep the voice, language and speed in the hash input, so changing the narrator regenerates everything on purpose.

Drafting loop cost, a 450-word script of about 2,700 characters (read 2026-10-03)
ApproachCharacters spokenTTS cost
Speak only the approved draft2,700$0.128
Speak each of 3 drafts plus the final10,800$0.513
Hash-keyed chunks, 1 of 2 chunks edited after approval2,700 + about 1,350$0.192

A hash-keyed submitter

The script splits on blank lines, packs paragraphs into chunks of about 2,000 characters, and submits each chunk with a key derived from its content, so an edit re-speaks only the chunk that contains it. Keep requests small. A failure then costs a few cents, and a retake touches one chunk.

import hashlib, os, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
VOICE = {"id": os.environ["NARRATOR_VOICE_ID"]}

def chunks(script, limit=2000):
    out, cur = [], ""
    for para in [p.strip() for p in script.split("\n\n") if p.strip()]:
        if cur and len(cur) + len(para) + 2 > limit:
            out.append(cur)
            cur = ""
        cur = f"{cur}\n\n{para}" if cur else para
    return out + ([cur] if cur else [])

def speak(script):
    ids = []
    for text in chunks(script):
        key = hashlib.sha256((text + VOICE["id"] + "normal").encode()).hexdigest()[:32]
        r = requests.post(f"{API}/v1/tts-1.0/generate",
                          headers={**H, "Idempotency-Key": key},
                          json={"transcript": text, "voice": VOICE, "speed": "normal"})
        r.raise_for_status()
        ids.append(r.json()["request_id"])
    return ids

print(speak(open("approved_script.txt").read()))

Joining the pieces

Each job returns its own audio. Join them with Timeline audio, operation: "concat", which takes up to 20 parts and costs $0.01 flat per job. Its default output is wav, which the docs call sample-exact, and they recommend keeping wav when the file will be joined again; mp3 re-adds priming padding at every edge. All parts must share one channel layout, or the job fails with audio_parts_channel_mismatch, so do not mix mono and stereo narration.

The result carries segments[] with the start of each part, which you re-base your video slot starts against. If the script needs more than 20 chunks, join in two levels.

What to do before you press send

Freeze the script. Put the approved text in a file under version control and run the submitter from that file only, so nobody pastes a half-edited draft into a live call.

Choose the chunk size by the cost of a retake. At 2,000 characters a chunk costs about $0.095, so a bad read costs under a dime. Pick paragraph boundaries that match natural pauses so the join does not land mid-sentence.

Log the key beside the job id for every chunk. When a reviewer says line four sounds flat, you can see at a glance which paragraph produced it, rewrite only that paragraph, and know that the other chunks will replay unchanged. The same log doubles as an audit trail of what was spoken, which matters when a script is approved by someone who is not the person pressing send.

Check the numbers on your own script: count characters, multiply by $0.0000475, and compare with the table. A 6,000 character script is $0.285 once and $1.14 if you speak four drafts.

What Sume does and does not do

Sume speaks the text you send and replays a matching idempotent request. It does not write the script, it does not know which draft you approved, and it does not edit audio at sentence level; replace a retaken chunk and re-run the concat. Gemini 3.8 Flash appears here only as an example of an LLM that can draft text.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume