LLM-written script to Sume TTS: speak only the approved draft
Draft with Gemini 3.8 Flash or any LLM, approve, then send hash-keyed paragraphs to Sume TTS so unchanged lines never bill twice. $0.0475 per 1,000 characters.

The cheapest way to voice a script that an LLM wrote is to keep the drafting loop out of the paid path: revise as often as you like in text, and send only the approved paragraphs to Sume TTS, each keyed by a hash of its text. Unchanged paragraphs then return the original job instead of being spoken, and billed, a second time. Sume TTS costs $0.0475 per 1,000 characters (read 2026-10-03).
Google's model page lists gemini-3.8-flash as stable, described as its "most intelligent Flash model" and positioned for long-horizon software engineering, agents and enterprise workflows, with 3.7 and 3.6 Flash labelled previous-generation (read 2026-10-03). The page I read gives no token limits or prices, and none of this article depends on them: any model that returns text can write the script, and Sume only sees the final characters.
Why hash the paragraph
Sume TTS 1.0 accepts a transcript of up to 20,000 characters, and a submit with an Idempotency-Key that repeats an earlier request returns the original job. If the key is a hash of the exact paragraph plus the voice settings, then an unchanged paragraph is a replay, and only edited paragraphs create new work.
Reusing a key for a different payload is a conflict: Sume answers 409 idempotency_conflict. That is the behaviour you want here, because a hash key can only collide with the same text. Keep the voice, language and speed in the hash input, so changing the narrator regenerates everything on purpose.
| Approach | Characters spoken | TTS cost |
|---|---|---|
| Speak only the approved draft | 2,700 | $0.128 |
| Speak each of 3 drafts plus the final | 10,800 | $0.513 |
| Hash-keyed chunks, 1 of 2 chunks edited after approval | 2,700 + about 1,350 | $0.192 |
A hash-keyed submitter
The script splits on blank lines, packs paragraphs into chunks of about 2,000 characters, and submits each chunk with a key derived from its content, so an edit re-speaks only the chunk that contains it. Keep requests small. A failure then costs a few cents, and a retake touches one chunk.
import hashlib, os, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
VOICE = {"id": os.environ["NARRATOR_VOICE_ID"]}
def chunks(script, limit=2000):
out, cur = [], ""
for para in [p.strip() for p in script.split("\n\n") if p.strip()]:
if cur and len(cur) + len(para) + 2 > limit:
out.append(cur)
cur = ""
cur = f"{cur}\n\n{para}" if cur else para
return out + ([cur] if cur else [])
def speak(script):
ids = []
for text in chunks(script):
key = hashlib.sha256((text + VOICE["id"] + "normal").encode()).hexdigest()[:32]
r = requests.post(f"{API}/v1/tts-1.0/generate",
headers={**H, "Idempotency-Key": key},
json={"transcript": text, "voice": VOICE, "speed": "normal"})
r.raise_for_status()
ids.append(r.json()["request_id"])
return ids
print(speak(open("approved_script.txt").read()))Joining the pieces
Each job returns its own audio. Join them with Timeline audio, operation: "concat", which takes up to 20 parts and costs $0.01 flat per job. Its default output is wav, which the docs call sample-exact, and they recommend keeping wav when the file will be joined again; mp3 re-adds priming padding at every edge. All parts must share one channel layout, or the job fails with audio_parts_channel_mismatch, so do not mix mono and stereo narration.
The result carries segments[] with the start of each part, which you re-base your video slot starts against. If the script needs more than 20 chunks, join in two levels.
What to do before you press send
Freeze the script. Put the approved text in a file under version control and run the submitter from that file only, so nobody pastes a half-edited draft into a live call.
Choose the chunk size by the cost of a retake. At 2,000 characters a chunk costs about $0.095, so a bad read costs under a dime. Pick paragraph boundaries that match natural pauses so the join does not land mid-sentence.
Log the key beside the job id for every chunk. When a reviewer says line four sounds flat, you can see at a glance which paragraph produced it, rewrite only that paragraph, and know that the other chunks will replay unchanged. The same log doubles as an audit trail of what was spoken, which matters when a script is approved by someone who is not the person pressing send.
Check the numbers on your own script: count characters, multiply by $0.0000475, and compare with the table. A 6,000 character script is $0.285 once and $1.14 if you speak four drafts.
What Sume does and does not do
Sume speaks the text you send and replays a matching idempotent request. It does not write the script, it does not know which draft you approved, and it does not edit audio at sentence level; replace a retaken chunk and re-run the concat. Gemini 3.8 Flash appears here only as an example of an LLM that can draft text.
Sources
Related posts
More in Use cases
- Localize a webinar Short: add captions first, dub only if it earns it
Localizing a webinar Short for new markets: start with translated captions at $0.20 a language, and only dub where captions don't get the views.
- Luma event cover image: square, 800 px minimum, text off the corners
Luma recommends a square 1:1 event cover at 800 by 800 px or more, static over GIF. How to generate one with Sume that survives rounded corners and small cards.
- Lunar New Year 2027 greeting video: a plan from 126 days out
Lunar New Year falls on Saturday 6 February 2027. Build a 15-second greeting from six stills, a music bed and caption cues for about $0.62 in Sume costs.
- Mailchimp video merge tag: the thumbnail links out, so pick it well
Mailchimp video merge tags show a thumbnail that links to a URL, so the still is the ad. Choose the frame with Sume video frames: 1 to 24 stills, source size.
Written by Sume