Repeat a TTS take: read generation_config and speed from the job

A completed Sume TTS job records its engine, voice, language, output_format, generation_config and speed. Read them back to make the next line sound the same.

4 min readSume
All posts

A completed Sume text-to-speech job records how its audio was made: model_id, voice, language, output_format, generation_config and speed. Read those fields from the job result and send them again with the next sentence, and the new line is requested with the same engine, voice and delivery settings. A setting you never sent comes back as null.

Which fields come back?

Jobs and results lists them: model_id is the engine, voice is { "mode": "id", "id": "..." }, then language, output_format, and the two settings the audio was synthesized with, generation_config and speed. The TTS worker also returns character_count, the audio URL, and, when you asked for timings, words and duration_seconds.

One consequence: you do not need to keep your own copy of the request. The job is the receipt.

Fields on a completed TTS job result, from Jobs and results and the TTS worker code, read 2026-10-02.
FieldWhat it holdsIf you did not send it
model_idThe engine that made the audioSet by the surface (TTS 1.0 or the router)
voice{ "mode": "id", "id": ... }Always present
languageLanguage the voice spoke innull
generation_configvolume, speed, emotion as sentnull
speedThe deprecated slow, normal or fast valuenull

How do I reuse them?

Copy the values into the next request. In a script, read the first job's JSON, then build the second body from voice.id, language, output_format and generation_config. For a router job, copy model_id into model. If generation_config is null, leave it out of the next body rather than sending null.

{
  "model_id": "sonic-3.6",
  "voice": { "mode": "id", "id": "<voice id from the job>" },
  "language": "en",
  "output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" },
  "generation_config": { "speed": 0.9, "emotion": "calm" },
  "speed": null
}

What does this not guarantee?

It repeats the request, not the waveform. Synthesis is a model call, and a model update behind an alias such as sonic-3.6 can change the sound of the same inputs; the snapshot post covers that. Keep the audio file when exact bytes matter.

The settings are copied onto the result because the original request object is not exposed on a public job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume