Splice an AI ending onto a Short and keep the original audio

Gemini Omni clips come with their own sound. To keep your Short's voice and music, detach both audio tracks and join them as audio parts in one Timeline render.

5 min readSume
All posts

To put an AI-generated ending after your Short without losing its audio, run audio detach on the original, decide which sound the ending should carry, and join the audio as audio.parts[] in one Timeline 1.0 render with the two videos as video[] slots. The parts are joined in the sample domain, so there is no re-synthesis and no silence at the seam.

This uses audio detach, Timeline 1.0 and the Video Router docs, read on 2026-10-03. The ending idea itself, as reported for YouTube's Extend with AI, is a new segment of typically 8 seconds or less.

Why does the audio need handling at all?

Gemini Omni Flash 1.1 on Sume always generates synced audio, and generate_audio: false is rejected. So the ending clip arrives with its own sound, which may not match the room tone or music of your Short. If you simply concatenated the two MP4s you would get two different sound worlds.

Timeline avoids this by separating the picture from the sound. One audio spine runs the whole program, and the video slots sit on top of it by start and duration.

What are the steps?

Detach the audio of the original Short, detach the audio of the ending clip, and decide the spine.

  • Run POST /v1/audio-detach on each video. The default is a sample-exact wav at $0.01 per job; the source must have an audio track or you get detach_source_has_no_audio.
  • Build audio.parts[] as the original audio (with duration to the cut point), then the ending's audio. Up to 20 parts per render; all parts must share one channel layout or the worker returns audio_parts_channel_mismatch.
  • Set audio.duration_seconds to the sum of the parts. The limit is 1 to 1800 seconds.
  • Set video[0].start to 0 and the ending slot's start to the original's length. Starts must increase.
  • Optionally add a fade on the second slot: transition: {type: fade, duration: 0.25}, which must be at most 1 second.

What does the render request look like?

This example joins a 24-second Short with an 8-second ending. The ending's own audio is used at its start so the new sound lands with the new picture.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: splice-ending-001" \
  -d '{
    "audio": {
      "duration_seconds": 32,
      "parts": [
        {"url": "https://media.sume.com/artifacts/artf_demo/short.wav", "duration": 24},
        {"url": "https://media.sume.com/artifacts/artf_demo/ending.wav", "duration": 8}
      ]
    },
    "video": [
      {"source_url": "https://media.sume.com/artifacts/artf_demo/short.mp4", "start": 0, "duration": 24},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/ending.mp4", "start": 24, "duration": 8,
       "transition": {"type": "fade", "duration": 0.25}}
    ]
  }'

What does it cost?

Two detach jobs are $0.02, and the render is $0.10 per ceil output minute, so a 32-second result bills one minute. The detach and Timeline rates come from the public docs and should be confirmed live in GET /v1/catalog. You can ask POST /v1/timeline-1.0/plan first; it is unbilled and returns billable_minutes and estimated_cost_usd_micros.

What Sume does not do here

It does not match loudness between the two sources and has no normalizer, so listen to the seam. It does not remove the ending's native sound from the clip itself; Timeline simply plays your spine instead. If the ending has speech you want to keep, make the ending's part the spine for that stretch.

  • Check probe.has_audio with video inspect before detaching a clip that may be silent.
  • Keep wav until the final render; mp3 re-adds priming padding at every edge.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume