Read back a Sume TTS job to keep the next line in the same voice
A completed Sume text_to_speech job records model_id, voice, language, output_format, generation_config and speed. Reuse them so line two matches line one.

What the job record keeps
A completed text-to-speech job on Sume records how the audio was made: model_id (the engine), voice as an object with mode id and an id, language, output_format, and the synthesis settings generation_config and speed. Each setting is null when the request did not send it. Read these values from the job to make the next line sound the same.
Fields to copy
| Field | What it holds | Reuse on the next request |
|---|---|---|
| model_id | The engine used | Send the same concrete model id |
| voice | Mode id plus the voice id | Send the same voice selector |
| language | Language used | Send the same value |
| output_format | Audio format | Keep it identical for clean joins |
| generation_config and speed | Synthesis settings, null if unsent | Copy only the non-null ones |
Why null matters
A null means you did not choose and the provider default applied. If you copy a null as an explicit value you may change the result. Copy only fields that are non-null, and leave the rest unsent. If you used an alias such as sonic-latest, pin the concrete model_id from the record for follow-up lines.
Pro voice clones are incompatible with sonic-preview; those requests fail with voice_model_mismatch.
A series routine
Store the first job's record with your project, send its values on every later request, and join the takes with Timeline audio at $0.01 per job. A 10-line series of 900 characters costs 10 x ceil(4.275) = 50 cents of voice plus 1 cent of join. See Jobs and results and Timeline audio.
Where to read it
Fetch the job result from GET /v1/jobs/:id/result after the job completes, and read the settings from the recorded fields. The same job envelope pattern is used across Sume models, so the polling code you have for video or music works here too.
Store the record with your project, not only the audio URL. The audio is the output; the record is the recipe.
Drift checks
If a later line sounds different, diff its record against the first. The usual causes are an alias that moved, a changed speed, a different output format or a voice id swapped by mistake. Fixing the cause costs nothing; regenerating the line costs the per-character price, for example 900 characters is 5 cents.
Sources
More in Developers
- Worker crashed mid-poll: list Sume jobs and join on idempotency_key
After a restart, GET /v1/jobs?status=processing lists your in-flight video jobs. Each row carries the idempotency_key you sent, so match your records on it.
- Recast prompt limit is 2,000 characters: what to write, what it costs
The H3 Max Recast prompt is optional and capped at 2,000 characters. See what to put in it for a person swap and why the length never changes the price on Sume.
- Refresh 1,200 SKU video ads before Black Friday: waves by plan
On an empty workspace, 1,200 jobs take 300 submission waves on Free, 67 on Pro and 14 on Scale, using Sume's wave_size_hint of 75% of accepted capacity.
- Reserve, capture, refund: a 20-clip batch with 3 failures, 2 cancels
Worked example of how a Sume balance moves across a 20-job batch when 3 jobs fail and 2 are canceled while queued: what is held, captured and released.
Written by Sume