What to log from a TTS job result so you can recreate a voiceover

A finished Sume TTS 1.0 job echoes model_id, voice, language, output_format and generation_config. Save them with the audio URL to rebuild the same take.

4 min readSume
All posts

The voiceover that your client approved in October needs a single changed line in March. If you only kept the MP3, you must guess the voice, the speed and the emotion. A completed Sume TTS 1.0 job records them for you, so what remains is to save the right fields.

The fields the result already echoes

According to the jobs and results page, a completed text_to_speech job returns model_id, voice as {mode, id}, language, output_format, generation_config and speed. Each is null when your request did not send it. The docs say to read these values from the job "to make the next line sound the same."

TTS 1.0 result fields worth storing (docs.sume.com jobs guide, read 2026-10-05)
FieldWhy keep it
audio_urlThe file. Archive it, since links do not live forever
model_idTTS 1.0 always uses sonic-3.6, so a model change shows up here
voice.idSame voice for the next line
languageEchoed from your request
generation_configVolume, speed and emotion actually applied
speedThe deprecated enum, null if you used generation_config

What the result does not hold

Your original transcript is not in the result, and neither is your pronunciation dictionary id or your idempotency key. Store those yourself. A good row in your own table is the key, the transcript text, the transcript hash, and the fields above as JSON.

Code

Save the result next to the text. SQLite is enough.

import json, sqlite3, hashlib
db = sqlite3.connect("voiceovers.db")
db.execute("create table if not exists vo (key text primary key, sha text, text text, result text)")

def log(key, text, result):
    keep = {k: result.get(k) for k in ("audio_url", "model_id", "voice", "language",
            "output_format", "generation_config", "speed")}
    sha = hashlib.sha256(text.encode("utf-8")).hexdigest()
    db.execute("insert or replace into vo values (?,?,?,?)", (key, sha, text, json.dumps(keep)))
    db.commit()

log("promo-014", "Free shipping ends Sunday.", {"audio_url": "https://media.sume.com/x.mp3",
    "model_id": "sonic-3.6", "language": "en", "generation_config": {"emotion": "calm"}})

Rebuilding a take

Read the row, send the stored voice.id, language and generation_config with the new text, and keep the same output_format. Costs are the same as before: $0.0475 per 1,000 characters, rounded up to a whole cent for each request. Pin nothing else. The model is chosen by Sume, so a future engine change appears in model_id, and you can see it before you ship the new line.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume