How do I narrate a DIY tutorial step by step with a TTS API?
Narrate an 8-step DIY tutorial with one TTS job per step: 1,570 characters, $0.10 on Sume. Why per-step jobs make a fixed step a 1-cent redo.

To narrate a DIY tutorial step by step, write one short script per step, send each as its own TTS job, and keep the files in order. Eight steps of 160 to 240 characters come to 1,570 characters, and on Sume TTS 1.0 the eight jobs cost $0.10 because each job rounds up to a whole cent, against $0.08 if the same text were one job. The extra cents buy you the ability to redo step five without touching the other seven.
Tutorials are corrected more than any other kind of narration. A measurement is wrong, a tool is renamed, a viewer asks for a slower step. Per-step files turn each of those into a small, cheap retake.
Write each step to stand alone
A step script should name the action, the object and the check: 'Press the clip into the groove until it clicks.' Keep it between 150 and 250 characters, which Cartesia's pricing page puts at roughly 12 to 20 seconds of speech at 750 to 800 credits a minute. Anything longer is two steps.
Open every file with a very short cue such as 'Step three.' It gives a viewer who scrubs through the video a reliable landmark, and it lets you reorder steps without re-recording, because the step number is spoken in each file and a reordered script carries it with it.
Generate each step with its own key
Give each job an idempotency key that includes the step and a version, such as tutorial-step-05-v2. A retry with the same key reuses the result; a corrected script gets a new version and a new job. In JavaScript the loop is short, and the polling follows next_poll_after_seconds from the status envelope.
const base = "https://api.sume.com/v1";
const h = { Authorization: `Bearer ${process.env.SUME_API_KEY}`, "Content-Type": "application/json" };
const steps = ["Step one. Unplug the lamp and lay it on a towel.", "Step two. Remove the three screws under the base."];
for (const [i, transcript] of steps.entries()) {
const res = await fetch(`${base}/tts-1.0/generate`, {
method: "POST",
headers: { ...h, "Idempotency-Key": `tutorial-step-${i + 1}-v1` },
body: JSON.stringify({ transcript, voice: { id: process.env.VOICE_ID }, language: "en" }),
});
const job = await res.json();
console.log(i + 1, res.status, job.job_id ?? job.data?.job_id);
}Line the steps up with the video
If you edit the video elsewhere, download each step's audio and place it at the start of its clip. If you assemble it on Sume, import the files and use Timeline 1.0 with an audio.parts[] list, which joins up to 20 parts gapless without re-synthesis; set each video slot's start to the running total of the step lengths. A silent screen recording or a sequence of photos works as the picture, and a short bed of music at a low gain keeps long pauses from sounding empty.
Handling the awkward words
Tutorials are full of part names, brand names and units that a voice may read badly. Spell a number the way it is said, write 'millimetres' in full if the abbreviation is read as letters, and keep a pronunciation dictionary and attach its id to each request, so a fix made once applies to all eight steps.
Use generation_config.speed between 0.6 and 1.5 for tricky steps: 0.9 for a fiddly step, 1.0 elsewhere. The emotion field is free text up to 64 characters, so 'calm, encouraging' is plenty. Neither changes the price.
Review before you publish
Play the files in order with the video muted and ask whether someone could do the job from the audio alone. Then play them with the picture and check that each cue lands on the right frame. A step that runs longer than its clip is a script problem: trim the words, do not speed the voice past 1.2.
Keep the scripts in version control next to the audio so a later change to a product has an obvious home. The idempotency key list doubles as a changelog.
What it costs next to the alternatives
For a tutorial this size the money is not the issue: every provider below stays well under a dollar. The table uses the same 1,600 characters for each and shows Sume's per-job rounding as a real number.
- Retake a step: one job, one cent, no change to the rest.
- Change the voice for the whole video: eight jobs, and the cost above again.
- Add a language: translate each step and repeat the eight jobs with
languageset.
| Provider and model | Listed rate | Cost |
|---|---|---|
| Sume TTS 1.0, eight jobs | $0.0475 per 1,000 characters, whole cents per job | $0.10 |
| Sume TTS 1.0, one job | same rate | $0.08 |
| ElevenLabs Flash/Turbo | $0.04 per 1,000 characters | $0.063 |
| OpenAI tts-1 | $15 per 1M characters | $0.024 |
Sources
Related posts
More in Developers
- Do I pay for a failed AI avatar video job? Refunds on Sume
Sume reserves the avatar video price at submit, captures it on completion, and releases or refunds it where a job fails. What it means for retries.
- Does PNG, JPEG or WebP change the price of an AI image on Sume?
No. On Sume's Image API, output_format picks the file type, not the price: per-image cards and GPT Image 2.5 token math ignore it. Which models list which.
- duration vs duration_seconds on each Sume video route
/v1/videos takes duration; motion control and lip-sync take duration_seconds; recast and edit read the source clip. One table of what each does.
- Fade in and out on a Timeline render: output fade seconds 0 to 5
Set output.fade_in_seconds and fade_out_seconds (0 to 5 s, sum within the render length). The music bed has its own fade_out_seconds, up to 10.
Written by Sume