Repeat a TTS take: read generation_config and speed from the job
A completed Sume TTS job records its engine, voice, language, output_format, generation_config and speed. Read them back to make the next line sound the same.

A completed Sume text-to-speech job records how its audio was made: model_id, voice, language, output_format, generation_config and speed. Read those fields from the job result and send them again with the next sentence, and the new line is requested with the same engine, voice and delivery settings. A setting you never sent comes back as null.
Which fields come back?
Jobs and results lists them: model_id is the engine, voice is { "mode": "id", "id": "..." }, then language, output_format, and the two settings the audio was synthesized with, generation_config and speed. The TTS worker also returns character_count, the audio URL, and, when you asked for timings, words and duration_seconds.
One consequence: you do not need to keep your own copy of the request. The job is the receipt.
| Field | What it holds | If you did not send it |
|---|---|---|
model_id | The engine that made the audio | Set by the surface (TTS 1.0 or the router) |
voice | { "mode": "id", "id": ... } | Always present |
language | Language the voice spoke in | null |
generation_config | volume, speed, emotion as sent | null |
speed | The deprecated slow, normal or fast value | null |
How do I reuse them?
Copy the values into the next request. In a script, read the first job's JSON, then build the second body from voice.id, language, output_format and generation_config. For a router job, copy model_id into model. If generation_config is null, leave it out of the next body rather than sending null.
{
"model_id": "sonic-3.6",
"voice": { "mode": "id", "id": "<voice id from the job>" },
"language": "en",
"output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" },
"generation_config": { "speed": 0.9, "emotion": "calm" },
"speed": null
}What does this not guarantee?
It repeats the request, not the waveform. Synthesis is a model call, and a model update behind an alias such as sonic-3.6 can change the sound of the same inputs; the snapshot post covers that. Keep the audio file when exact bytes matter.
The settings are copied onto the result because the original request object is not exposed on a public job.
Sources
Related posts
More in Developers
- Reconcile a no-code run with GET /v1/jobs and idempotency_key
When a Zap, scenario or flow loses its job ids, list jobs with GET /v1/jobs and join on idempotency_key. Pages are newest first, 100 at most, cursor-paged.
- Retry an avatar video request without paying twice
A timed-out avatar video submit can be retried safely with the same Idempotency-Key. What Sume returns, what causes a 409, and how to build the key.
- Ruby Net::HTTP: submit and poll a Sume job with one key
A Ruby recipe using only Net::HTTP: submit a Sume image job with an Idempotency-Key, set open and read timeouts, poll status_url, and return the artifact URLs.
- Feed scraped product copy to a Format run: input, not instruction
Putting a supplier's text into a Format's instruction lets it steer the run. Sume's input field is treated as data, with a 64-key and 2 MiB limit.
Written by Sume