Text to dialogue API: continuity between requests on Sume
Sume's TTS has no previous_text field. Keep a line seamless with one longer transcript, or join separate takes with a gapless Timeline audio concat.

Sume's text-to-speech request has no field for the text before or after a line, so continuity between separate requests is not something you pass in. Put the lines that must flow into one transcript, or render takes separately and join them with Timeline audio concat.
The ElevenLabs changelog entry for Sept 28 added previous_text and future_text (max 100 characters each) and previous_request_ids / next_request_ids (max 3) to Text-to-Dialogue. Sume's OpenAPI contract, read 2026-09-30, contains none of those names.
What can I do in one Sume request?
POST /v1/tts-1.0/generate accepts a transcript of up to 20000 characters. Optional timestamps.words with segmentation.mode: "sentence" returns word timings and gapless per-sentence segments[] (segment[i].end == segment[i+1].start); audio slices additionally need emit_audio and a wav or raw container. Only "sentence" is supported in v1. If the whole exchange fits in one transcript, one request keeps its delivery in one render.
How do I join separate takes without a gap?
Import the audio, then call POST /v1/timeline-1.0/audio with operation: "concat" and up to 20 ordered parts[]. The docs describe the join as sample-domain: no re-TTS and no silence at the seams. The call needs an Idempotency-Key, and the job is polled through /v1/jobs/:id/status and /result.
| Need | ElevenLabs Text-to-Dialogue | Sume |
|---|---|---|
| Context from neighboring text | previous_text / future_text | No such field; use one transcript |
| Link to earlier requests | previous_request_ids / next_request_ids | No such field |
| Join finished takes | Not stated in the changelog entry | concat, sample-domain, no added silence |
What can go wrong when joining?
Parts must share one channel layout, or the job fails with audio_parts_channel_mismatch. Keep wav output (the default) when the file will be joined again: mp3 re-adds priming padding at every edge. Joining preserves each take as rendered, so a delivery mismatch between two takes is not smoothed; it is only made gapless. Details are in Timeline audio.
Which one should I pick?
Use one transcript when the lines fit in 20000 characters and one voice setup. Use concat when takes are produced at different times or need individual retries. For per-line voices, see text to speech with multiple voices.
Sources
Related posts
More in Developers
- TikTok Research API documentation: 30-day dates vs Sume date_posted
The Research API takes start_date and end_date up to 30 days apart. Sume's tiktok_search_keyword takes one fixed date_posted value instead.
- TikTok Research API access: is_random order vs Sume order
The Research API returns videos by decreasing ID unless is_random is true. Sume's order=play_count sorts a bounded 24-item sample, not a full ranking.
- TikTok Research API rate limit: max_count 100 vs Sume's 24
The Research API returns up to 100 videos per call (max_count). Sume's tiktok_search_keyword returns at most 24 items over 3 feed pages per call.
- Join voiceover clips inside one timeline render with audio.parts
Timeline 1.0 takes up to 20 gapless audio.parts slices with source_in and duration. Skip the timeline-audio job when the join is only for one render.
Written by Sume