AI voiceover for a PowerPoint: one TTS job per slide, 20 for $0.57

Narrate a slide deck by sending each slide's notes as its own Sume TTS job. 20 slides at 600 characters each cost $0.57, and you keep one audio file per slide.

4 min readSume
All posts

The short answer

Send the speaker notes of each slide as a separate text-to-speech job, then drop each audio file onto its slide. On Sume that is one POST /v1/tts-1.0/generate per slide at $0.0475 per 1,000 characters. A 20-slide deck with about 600 characters of notes per slide is 12,000 characters, which is 12 x $0.0475 = $0.57.

One job per slide beats one long job for a deck. If slide 7 changes, you regenerate 600 characters, not the whole talk, and you never have to find where slide 7 starts inside a single long file.

What the cost looks like

Spaces and punctuation count as characters, so paste the notes exactly as you will have them read. Price is linear, so you can scale it in your head.

Deck narration cost on Sume TTS at $0.0475 per 1,000 characters (read 2026-10-04)
SlidesCharacters per slideTotal charactersCost
106006,000$0.285
2060012,000$0.57
4060024,000$1.14

One request per slide

Pick a voice once, reuse its id for every slide, and set language explicitly if the deck is not English. Writes need an Idempotency-Key, so use the slide number in it; a retry of slide 7 then cannot bill twice.

Requests are jobs. By default the response is a 202 with status and result URLs. Poll GET /v1/jobs/{id}/status, then read the audio from GET /v1/jobs/{id}/result. The file is in result.artifacts[] as audio/mpeg on media.sume.com. Details are in the jobs and results guide.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: deck-q4-slide-07" \
  -d '{"transcript":"Revenue grew in every region this quarter.",
       "voice":{"id":"voi_YOUR_VOICE_ID"},"language":"en"}'

Keep the deck sounding like one person

Every slide must use the same voice id and the same generation_config (speed 0.6 to 1.5, volume 0.5 to 2). Freeze those values in one place so slide 20 does not drift from slide 1. If one slide sounds louder, change its volume and regenerate only that slide.

  • Write numbers and acronyms the way they should be spoken before you submit.
  • Keep each slide under the 20,000-character request limit, which is far above a slide's notes.
  • Name files by slide number so your deck tool inserts them in order.
  • Store the job id next to each slide so you can fetch the file again.

What Sume does not do here

Sume produces the audio files. It does not edit your .pptx or time slide transitions. Use the audio length in the job result when you set auto-advance in your slide software.

A workflow you can repeat each quarter

Treat the deck like source code. Keep the notes in one text file per slide, with the slide number in the file name. When a number changes, edit the text file, bump the version in the idempotency key (deck-q4-slide-07-v2), and submit just that slide. Your list of job ids then doubles as a change log: you can see which slides were regenerated and when.

Before you commit to a full run, generate slide 1 and slide 2 only and listen on the device your audience will use. A voice that sounds right in headphones can be too fast in a meeting room. Adjust speed once, then reuse the setting for the other 18 slides. The cost of the test is under three cents.

If the deck is translated, treat each language as its own set of 20 jobs with its own voice and language value. The cost simply doubles for two languages: 24,000 characters is $1.14 at the same rate.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume