Read back a Sume TTS job to keep the next line in the same voice

A completed Sume text_to_speech job records model_id, voice, language, output_format, generation_config and speed. Reuse them so line two matches line one.

5 min readSume
All posts

What the job record keeps

A completed text-to-speech job on Sume records how the audio was made: model_id (the engine), voice as an object with mode id and an id, language, output_format, and the synthesis settings generation_config and speed. Each setting is null when the request did not send it. Read these values from the job to make the next line sound the same.

Fields to copy

Recorded TTS fields (Sume docs, read 2026-10-08)
FieldWhat it holdsReuse on the next request
model_idThe engine usedSend the same concrete model id
voiceMode id plus the voice idSend the same voice selector
languageLanguage usedSend the same value
output_formatAudio formatKeep it identical for clean joins
generation_config and speedSynthesis settings, null if unsentCopy only the non-null ones

Why null matters

A null means you did not choose and the provider default applied. If you copy a null as an explicit value you may change the result. Copy only fields that are non-null, and leave the rest unsent. If you used an alias such as sonic-latest, pin the concrete model_id from the record for follow-up lines.

Pro voice clones are incompatible with sonic-preview; those requests fail with voice_model_mismatch.

A series routine

Store the first job's record with your project, send its values on every later request, and join the takes with Timeline audio at $0.01 per job. A 10-line series of 900 characters costs 10 x ceil(4.275) = 50 cents of voice plus 1 cent of join. See Jobs and results and Timeline audio.

Where to read it

Fetch the job result from GET /v1/jobs/:id/result after the job completes, and read the settings from the recorded fields. The same job envelope pattern is used across Sume models, so the polling code you have for video or music works here too.

Store the record with your project, not only the audio URL. The audio is the output; the record is the recipe.

Drift checks

If a later line sounds different, diff its record against the first. The usual causes are an alias that moved, a changed speed, a different output format or a voice id swapped by mistake. Fixing the cause costs nothing; regenerating the line costs the per-character price, for example 900 characters is 5 cents.

Sources

More in Developers

All Developers posts

Written by Sume