Turn a voice memo into a narrated short: transcribe, edit, re-voice
Speak your idea into a phone, transcribe it with Sume STT, tidy the text, then re-voice and caption it. The memo becomes the script, not the audio.

A voice memo makes a good script and a poor soundtrack. The route is to transcribe the memo, edit the text into the lines you want spoken, voice it with TTS, and caption the result. Sume's STT 1.0 gives you the first draft: text plus words[] timings, from a public HTTPS audio URL of up to 10 minutes.
Step 1: transcribe
Send the memo's audio_url to POST /v1/stt-1.0/transcribe. Omit language_code for auto-detect, or pass a hint such as en. Pass duration_seconds so the reservation fits, and add segmentation: {"mode": "sentence"} to get one line per sentence. That is the shape you want for editing.
Step 2: edit the text
Cut the filler, fix the numbers and names, and keep one idea per sentence. Do not feed the raw transcript back into TTS, because it keeps every hesitation. Spell out anything the voice might misread, such as an order code or an email address, in the form it should be spoken.
Step 3: re-voice
Send the edited lines to POST /v1/tts-1.0/generate. transcript can be up to 20000 characters. Use a voice whose language tag matches the language you send.
Step 4: caption with the script
Render the clip, then call POST /v1/video-captions with the same edited text as script_text. A caption job reads script_text and aligns it to the speech, so the burned words match your script and not the recognizer's guess. Without script_text, the job runs speech-to-text itself.
What each step needs, read 2026-10-06:
| Step | Endpoint | Input |
|---|---|---|
| Transcribe | POST /v1/stt-1.0/transcribe | Public HTTPS audio_url, up to 600 s |
| Re-voice | POST /v1/tts-1.0/generate | Transcript up to 20000 characters |
| Caption | POST /v1/video-captions | Public video_url, optional script_text |
Why not keep the memo's own audio
A memo is usually recorded in a room, a car or a street. Noise, wind and a phone microphone are hard to hide, and a voice that stumbles reads as a rough cut. A fresh read from the edited text is cleaner and shorter.
If the memo is the point, because it is a testimony or an interview, keep the audio and transcribe it only for captions. In that case send the caption job the clip itself, and let the recognizer or your own script_text supply the words.
Keep the original memo if you want a record of what was said. If you use another person's voice in the memo, get their agreement before you publish.
Sources
Related posts
More in Use cases
- YouTube watch series button: viewers may land on episode 7 first
A viewer who finds one Short in their feed sees a watch series button, so any episode can be the first they see. Write and render episodes that stand alone.
- What is a consistent story for a YouTube show? AI series checklist
YouTube asks shows for the same characters or hosts, one long story, or one topic. A checklist for keeping an AI-made series consistent across episodes.
- Where a YouTube Shorts series can appear: search, Shows tab and more
YouTube's help lists search, Recommended shows, Continue watching and a Shorts icon in the Shows tab. What each means for how you plan a series.
- Which Sume video model makes a 15-second 9:16 ad? A catalog filter
List the models in GET /v1/videos/models that accept 15 seconds at 9:16 with a short Python filter, instead of trusting a stale table.
Written by Sume