How do I narrate a DIY tutorial step by step with a TTS API?

Narrate an 8-step DIY tutorial with one TTS job per step: 1,570 characters, $0.10 on Sume. Why per-step jobs make a fixed step a 1-cent redo.

4 min readSume
All posts

To narrate a DIY tutorial step by step, write one short script per step, send each as its own TTS job, and keep the files in order. Eight steps of 160 to 240 characters come to 1,570 characters, and on Sume TTS 1.0 the eight jobs cost $0.10 because each job rounds up to a whole cent, against $0.08 if the same text were one job. The extra cents buy you the ability to redo step five without touching the other seven.

Tutorials are corrected more than any other kind of narration. A measurement is wrong, a tool is renamed, a viewer asks for a slower step. Per-step files turn each of those into a small, cheap retake.

Write each step to stand alone

A step script should name the action, the object and the check: 'Press the clip into the groove until it clicks.' Keep it between 150 and 250 characters, which Cartesia's pricing page puts at roughly 12 to 20 seconds of speech at 750 to 800 credits a minute. Anything longer is two steps.

Open every file with a very short cue such as 'Step three.' It gives a viewer who scrubs through the video a reliable landmark, and it lets you reorder steps without re-recording, because the step number is spoken in each file and a reordered script carries it with it.

Generate each step with its own key

Give each job an idempotency key that includes the step and a version, such as tutorial-step-05-v2. A retry with the same key reuses the result; a corrected script gets a new version and a new job. In JavaScript the loop is short, and the polling follows next_poll_after_seconds from the status envelope.

const base = "https://api.sume.com/v1";
const h = { Authorization: `Bearer ${process.env.SUME_API_KEY}`, "Content-Type": "application/json" };
const steps = ["Step one. Unplug the lamp and lay it on a towel.", "Step two. Remove the three screws under the base."];

for (const [i, transcript] of steps.entries()) {
  const res = await fetch(`${base}/tts-1.0/generate`, {
    method: "POST",
    headers: { ...h, "Idempotency-Key": `tutorial-step-${i + 1}-v1` },
    body: JSON.stringify({ transcript, voice: { id: process.env.VOICE_ID }, language: "en" }),
  });
  const job = await res.json();
  console.log(i + 1, res.status, job.job_id ?? job.data?.job_id);
}

Line the steps up with the video

If you edit the video elsewhere, download each step's audio and place it at the start of its clip. If you assemble it on Sume, import the files and use Timeline 1.0 with an audio.parts[] list, which joins up to 20 parts gapless without re-synthesis; set each video slot's start to the running total of the step lengths. A silent screen recording or a sequence of photos works as the picture, and a short bed of music at a low gain keeps long pauses from sounding empty.

Handling the awkward words

Tutorials are full of part names, brand names and units that a voice may read badly. Spell a number the way it is said, write 'millimetres' in full if the abbreviation is read as letters, and keep a pronunciation dictionary and attach its id to each request, so a fix made once applies to all eight steps.

Use generation_config.speed between 0.6 and 1.5 for tricky steps: 0.9 for a fiddly step, 1.0 elsewhere. The emotion field is free text up to 64 characters, so 'calm, encouraging' is plenty. Neither changes the price.

Review before you publish

Play the files in order with the video muted and ask whether someone could do the job from the audio alone. Then play them with the picture and check that each cue lands on the right frame. A step that runs longer than its clip is a script problem: trim the words, do not speed the voice past 1.2.

Keep the scripts in version control next to the audio so a later change to a product has an obvious home. The idempotency key list doubles as a changelog.

What it costs next to the alternatives

For a tutorial this size the money is not the issue: every provider below stays well under a dollar. The table uses the same 1,600 characters for each and shows Sume's per-job rounding as a real number.

  • Retake a step: one job, one cent, no change to the rest.
  • Change the voice for the whole video: eight jobs, and the cost above again.
  • Add a language: translate each step and repeat the eight jobs with language set.
Cost of 1,570 characters of tutorial narration, vendor list rates read 2026-10-07 and Sume TTS 1.0 catalog rate.
Provider and modelListed rateCost
Sume TTS 1.0, eight jobs$0.0475 per 1,000 characters, whole cents per job$0.10
Sume TTS 1.0, one jobsame rate$0.08
ElevenLabs Flash/Turbo$0.04 per 1,000 characters$0.063
OpenAI tts-1$15 per 1M characters$0.024

Sources

Related posts

More in Developers

All Developers posts

Written by Sume