Re-voice a recording: Sume STT then TTS as two jobs
Transcribe a recording with Sume STT, fix the text, then speak it with Sume TTS. Two jobs, two prices, and where the human edit goes.

To re-voice a recording, run it through Sume STT to get the text, edit that text, then send it to Sume TTS. That is two separate jobs with two separate prices: $0.01 per audio minute for the transcript and $0.0475 per 1,000 characters for the new speech. Nothing in the API copies the original performance; you are writing a fresh read from the text, and the edit step is yours.
This suits corrected narration, a cleaner take of a scratch track, or a new language after you translate the text yourself.
What does it cost?
Prices come from the Sume OpenAPI spec and the public price book. A 5-minute recording whose transcript is 4,000 characters costs $0.05 to transcribe and $0.19 to speak, a derived total of $0.24.
| Step | Quantity | Rate | Cost |
|---|---|---|---|
| STT | 5 audio minutes | $0.01 per minute | $0.05 |
| TTS | 4,000 characters | $0.0475 per 1,000 | $0.19 |
| Total | $0.24 |
Where does the human edit go?
Between the jobs. Read the transcript, fix names and numbers, and remove filler you do not want spoken. Do not feed raw STT output straight into TTS and ship it; the transcript may carry mistakes that become audible.
- Keep the new text under 20,000 characters per request.
- Pass a
languagethat matches the voice, or expect a language mismatch check.
What does the chain look like in code?
This reads the edited text from a file you control, so the human step stays outside the script:
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"
stt = requests.post(B + "/v1/stt-1.0/transcribe", headers=H, timeout=60,
json={"audio_url": os.environ["AUDIO_URL"], "language_code": "en"})
print("transcribe job", stt.json()["data"]["job"]["id"])
# fetch /result, edit the text into edited.txt, then:
text = open("edited.txt", encoding="utf-8").read()
tts = requests.post(B + "/v1/tts-1.0/generate", headers=H, timeout=60,
json={"transcript": text, "voice": {"id": os.environ["VOICE_ID"]},
"language": "en"})
print("speech job", tts.json()["data"]["job"]["id"])
Where does this scale?
For a whole back catalog, see detaching audio and transcribing a podcast archive. For scripts that exceed one request, see 100,000 characters across five jobs.
Sources
Related posts
More in Use cases
- Real estate agent intro clip: headshot, logo and listing photos
Send a headshot, your logo and listing photos as reference_image_urls to Gemini Omni Flash 1.1 and address each as <IMAGE_REF_n> for a 3 to 10 second intro.
- Listing clip from photos with Seedance 2.5, and what to disclose
Nine listing photos can feed a Seedance 2.5 reference-to-video request on Sume. Price a 30 s clip, plan the camera path, label it AI-generated.
- Real estate price-reduction clip from one listing photo: $0.83 each
A 5-second vertical price-drop clip from the hero photo with the new price burned in as cues: $0.825 per listing, $9.90 for 12, on Sume from docs.
- Rehearsal dinner slideshow: 24 photos and Timeline's 8-fade cap
A 24-photo slideshow with fades fails if every cut is a fade: Timeline refuses more than 8 chained transitions. Insert hard cuts; 96 seconds costs $0.20.
Written by Sume