Join two generated tracks without a gap: timeline audio concat
One Lyria generation too short? Join up to 20 Sume-hosted audio files into one gapless wav with timeline audio concat and read segment offsets.

How do you join two AI-generated music tracks into one file? Use POST /v1/timeline-1.0/audio with operation: "concat" and a parts[] list. The Timeline audio docs say the join is sample-domain, so there is no silence at the seams and no re-synthesis, and the result is a durable media.sume.com file with the offsets for each part.
Parts are ordered, 1 to 20 of them, and every URL must already be audio in your workspace on media.sume.com. Import files first with POST /v1/media-imports.
The request
Each part is { url, source_in?, duration? }. Do not send url or ranges at the top level; that returns audio_concat_takes_no_url or audio_concat_takes_no_ranges. Idempotency-Key is required.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: concat-tracks-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/track1.mp3" },
{ "url": "https://media.sume.com/artifacts/artf_demo/track2.mp3", "source_in": 0.1, "duration": 28 }
]
}'What comes back
A kind: timeline_audio result with one audio_url, a duration_seconds and segments[] holding index, start and duration_seconds. Those offsets are what you re-base video[].start against on a Timeline 1.0 render.
| Rule | Value |
|---|---|
| Parts | 1 to 20, ordered |
| Channel layout | Must match, else audio_parts_channel_mismatch |
| Output format | wav default, or mp3 |
| Produced length | Up to 1800 seconds |
| Public rate | $0.01 flat per job, confirm in GET /v1/catalog |
wav or mp3
The default is sample-exact pcm_s16le wav. mp3 is smaller but re-adds priming padding at every edge, so keep wav for any file you will join again. Beds from the router come back as audio/mpeg, so the first join of two of them is already an mp3 input; decide the output format with that in mind.
When not to mint a file
If the join is only needed inside one render, put the parts on audio.parts[] of the Timeline 1.0 request and skip this job. Mint a file only when you want to reuse it.
Sources
Related posts
More in Developers
- Container snapshot restore: resume a Sume job from its job_id
Cloudflare Containers can snapshot a running container. Save the Sume job_id in it, and after a restore read the job's status_url instead of resubmitting.
- Copilot CLI dynamic workflows: let them call the Sume CLI
GitHub added dynamic workflows to Copilot CLI on Oct 1, 2026. Give it the Sume CLI skill pack and the hosted MCP endpoint, and keep paid calls behind a gate.
- createSumeClient timeout is 10 minutes per request: tune it for polls
The Sume SDK client waits up to 10 minutes on each HTTP request, so one hung status poll can stall waitForJob for 10 minutes. Use a short-timeout client.
- CrewAI conversational flows: confirm cost before a Sume render
In a CrewAI chat flow, have the step call Sume with dry_run first and ask the user to confirm the cost.
Written by Sume