Re-voice a recording: Sume STT then TTS as two jobs

Transcribe a recording with Sume STT, fix the text, then speak it with Sume TTS. Two jobs, two prices, and where the human edit goes.

5 min readSume
All posts

To re-voice a recording, run it through Sume STT to get the text, edit that text, then send it to Sume TTS. That is two separate jobs with two separate prices: $0.01 per audio minute for the transcript and $0.0475 per 1,000 characters for the new speech. Nothing in the API copies the original performance; you are writing a fresh read from the text, and the edit step is yours.

This suits corrected narration, a cleaner take of a scratch track, or a new language after you translate the text yourself.

What does it cost?

Prices come from the Sume OpenAPI spec and the public price book. A 5-minute recording whose transcript is 4,000 characters costs $0.05 to transcribe and $0.19 to speak, a derived total of $0.24.

Worked example for a 5 minute recording, Sume list prices read 2026-10-04
StepQuantityRateCost
STT5 audio minutes$0.01 per minute$0.05
TTS4,000 characters$0.0475 per 1,000$0.19
Total$0.24

Where does the human edit go?

Between the jobs. Read the transcript, fix names and numbers, and remove filler you do not want spoken. Do not feed raw STT output straight into TTS and ship it; the transcript may carry mistakes that become audible.

  • Keep the new text under 20,000 characters per request.
  • Pass a language that matches the voice, or expect a language mismatch check.

What does the chain look like in code?

This reads the edited text from a file you control, so the human step stays outside the script:

import os, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"

stt = requests.post(B + "/v1/stt-1.0/transcribe", headers=H, timeout=60,
    json={"audio_url": os.environ["AUDIO_URL"], "language_code": "en"})
print("transcribe job", stt.json()["data"]["job"]["id"])
# fetch /result, edit the text into edited.txt, then:
text = open("edited.txt", encoding="utf-8").read()
tts = requests.post(B + "/v1/tts-1.0/generate", headers=H, timeout=60,
    json={"transcript": text, "voice": {"id": os.environ["VOICE_ID"]},
          "language": "en"})
print("speech job", tts.json()["data"]["job"]["id"])

Where does this scale?

For a whole back catalog, see detaching audio and transcribing a podcast archive. For scripts that exceed one request, see 100,000 characters across five jobs.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume