Turn a voice memo into a narrated short: transcribe, edit, re-voice

Speak your idea into a phone, transcribe it with Sume STT, tidy the text, then re-voice and caption it. The memo becomes the script, not the audio.

4 min readSume
All posts

A voice memo makes a good script and a poor soundtrack. The route is to transcribe the memo, edit the text into the lines you want spoken, voice it with TTS, and caption the result. Sume's STT 1.0 gives you the first draft: text plus words[] timings, from a public HTTPS audio URL of up to 10 minutes.

Step 1: transcribe

Send the memo's audio_url to POST /v1/stt-1.0/transcribe. Omit language_code for auto-detect, or pass a hint such as en. Pass duration_seconds so the reservation fits, and add segmentation: {"mode": "sentence"} to get one line per sentence. That is the shape you want for editing.

Step 2: edit the text

Cut the filler, fix the numbers and names, and keep one idea per sentence. Do not feed the raw transcript back into TTS, because it keeps every hesitation. Spell out anything the voice might misread, such as an order code or an email address, in the form it should be spoken.

Step 3: re-voice

Send the edited lines to POST /v1/tts-1.0/generate. transcript can be up to 20000 characters. Use a voice whose language tag matches the language you send.

Step 4: caption with the script

Render the clip, then call POST /v1/video-captions with the same edited text as script_text. A caption job reads script_text and aligns it to the speech, so the burned words match your script and not the recognizer's guess. Without script_text, the job runs speech-to-text itself.

What each step needs, read 2026-10-06:

Voice memo to short, read 2026-10-06
StepEndpointInput
TranscribePOST /v1/stt-1.0/transcribePublic HTTPS audio_url, up to 600 s
Re-voicePOST /v1/tts-1.0/generateTranscript up to 20000 characters
CaptionPOST /v1/video-captionsPublic video_url, optional script_text

Why not keep the memo's own audio

A memo is usually recorded in a room, a car or a street. Noise, wind and a phone microphone are hard to hide, and a voice that stumbles reads as a rough cut. A fresh read from the edited text is cleaner and shorter.

If the memo is the point, because it is a testimony or an interview, keep the audio and transcribe it only for captions. In that case send the caption job the clip itself, and let the recognizer or your own script_text supply the words.

Keep the original memo if you want a record of what was said. If you use another person's voice in the memo, get their agreement before you publish.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume