Shadowing Audio for 30 Sentences With script_run: Cost and Setup

Make 30 sentence-length audio clips for a language lesson with one script_run program and tts_create: the cost for 80-character lines and the limits.

4 min readSume
All posts

The short answer

Thirty sentences of about 80 characters each is 2,400 characters of speech, which costs $0.114 on Sume's TTS rate of $0.0475 per 1,000 characters (read 2026-10-08). One script_run program can create all 30 tts_create jobs in a loop, so you ask once and the platform runs the calls for you.

Why script_run for this

Sume's MCP docs describe script_run as a short JavaScript program that calls hosted tools in a loop or in parallel and returns one value. The page names this exact shape: use it when a turn needs three or more independent calls of the same kind, such as one tts_create per sentence. Inside the script, await sume.call(name, arguments) runs any listed tool with the same gates as a direct call.

Each paid call still needs its own idempotency_key. Derive it from the lesson id and sentence number, for example lesson4-s07, so a retry returns the same job. The run has limits you set: timeout_seconds from 5 to 55, max_calls and max_paid_calls. Set max_paid_calls to 30 and nothing past the plan can be spent.

Budget for a 30-sentence lesson (Sume catalog and docs, read 2026-10-08)
ItemValueResult
Characters per sentence802,400 total
TTS price$0.0475 per 1,000 characters$0.114
Max single request20,000 characters$0.95 estimate
Calls in the script30 paid tts_createmax_paid_calls: 30
Speech-to-text check, 3 minutes$0.01 per audio minute$0.03

Using the clips

Each job returns an audio artifact on media.sume.com. To make one practice file with gaps for the learner to repeat, join the clips with timeline audio (concat or split, $0.01 per job) or use them as audio.parts[] in a Timeline 1.0 render, which takes up to 20 gapless slices per spine. For a lesson of 30 sentences, split into two spines of 15.

The completed TTS job records the engine, the voice id, the language, the output format and the synthesis settings. Docs say to read those from the job to make the next line sound the same. For a course, that is how lesson 5 keeps the voice of lesson 1.

  • Keep one idea per sentence; very long lines cost more and are harder to shadow.
  • Use tts_source_verify_spine to compare selected TTS jobs with the accepted script before you assemble.
  • Check pronunciation of names and numbers by ear; a synthetic voice can misread them.

When not to use it

If you have only two or three lines, call tts_create directly. A script adds overhead and is for volume. If your lesson needs a human voice for assessment, TTS can model the target sound but cannot replace a teacher's feedback.

A last practical point is review cost. Listen to all 30 clips once before you ship; if four need a new take, that is four more calls and about $0.0152 at the same length. The risk is not the price but a mispronounced word that a learner copies. Keep a list of terms to spell out phonetically in the script for next time.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume