Shadowing Audio for 30 Sentences With script_run: Cost and Setup
Make 30 sentence-length audio clips for a language lesson with one script_run program and tts_create: the cost for 80-character lines and the limits.

The short answer
Thirty sentences of about 80 characters each is 2,400 characters of speech, which costs $0.114 on Sume's TTS rate of $0.0475 per 1,000 characters (read 2026-10-08). One script_run program can create all 30 tts_create jobs in a loop, so you ask once and the platform runs the calls for you.
Why script_run for this
Sume's MCP docs describe script_run as a short JavaScript program that calls hosted tools in a loop or in parallel and returns one value. The page names this exact shape: use it when a turn needs three or more independent calls of the same kind, such as one tts_create per sentence. Inside the script, await sume.call(name, arguments) runs any listed tool with the same gates as a direct call.
Each paid call still needs its own idempotency_key. Derive it from the lesson id and sentence number, for example lesson4-s07, so a retry returns the same job. The run has limits you set: timeout_seconds from 5 to 55, max_calls and max_paid_calls. Set max_paid_calls to 30 and nothing past the plan can be spent.
| Item | Value | Result |
|---|---|---|
| Characters per sentence | 80 | 2,400 total |
| TTS price | $0.0475 per 1,000 characters | $0.114 |
| Max single request | 20,000 characters | $0.95 estimate |
| Calls in the script | 30 paid tts_create | max_paid_calls: 30 |
| Speech-to-text check, 3 minutes | $0.01 per audio minute | $0.03 |
Using the clips
Each job returns an audio artifact on media.sume.com. To make one practice file with gaps for the learner to repeat, join the clips with timeline audio (concat or split, $0.01 per job) or use them as audio.parts[] in a Timeline 1.0 render, which takes up to 20 gapless slices per spine. For a lesson of 30 sentences, split into two spines of 15.
The completed TTS job records the engine, the voice id, the language, the output format and the synthesis settings. Docs say to read those from the job to make the next line sound the same. For a course, that is how lesson 5 keeps the voice of lesson 1.
- Keep one idea per sentence; very long lines cost more and are harder to shadow.
- Use
tts_source_verify_spineto compare selected TTS jobs with the accepted script before you assemble. - Check pronunciation of names and numbers by ear; a synthetic voice can misread them.
When not to use it
If you have only two or three lines, call tts_create directly. A script adds overhead and is for volume. If your lesson needs a human voice for assessment, TTS can model the target sound but cannot replace a teacher's feedback.
A last practical point is review cost. Listen to all 30 clips once before you ship; if four need a new take, that is four more calls and about $0.0152 at the same length. The risk is not the price but a mispronounced word that a learner copies. Keep a list of terms to spell out phonetically in the script for next time.
Sources
Related posts
More in Use cases
- Launch kit: 4 clips, 12 images, a voiceover and a track for $9.98
Four 8-second 1080p clips, twelve 2K images, a 1,200-character voiceover and a track cost $9.982 with Wan 3.0, or $20.4716 with Seedance 2.5 clips.
- LinkedIn allows 25 video uploads in 24 hours: plan a 60-variant test
LinkedIn Campaign Manager caps video uploads at 25 per 24 hours, desktop only. A 60-clip test needs three days. Cut and price the variants on Sume first.
- LinkedIn landscape video ad maxes at 1920x1080; Sume allows up to 2160
LinkedIn's landscape video range is 640x360 to 1920x1080, but Sume timeline accepts sizes to 2160. Pick 1920x1080 on purpose and verify with video inspect.
- Localize a Claude Motion explainer: four caption languages for $0.80
Caption one Motion MP4 in English, Spanish, French and Korean with four Sume jobs at $0.20 each, $0.80 in all, no speech-to-text. Korean needs its own style.
Written by Sume