Live replay highlight reel in one render: detach once, no trims

Detach a live replay's audio once ($0.01), then build a 30-second highlight reel in one Timeline render using audio parts and source_in. No per-clip trim jobs.

5 min readSume
All posts

To build a highlight reel from a live-selling replay, detach the replay's audio once into a wav, then describe the reel in a single Timeline 1.0 document: audio.parts[] slices the wav, and each video[] slot points back into the same replay with source_in. One detach job ($0.01) plus one render ($0.10 per output minute) replaces a trim job per clip.

The catch is that you choose the moments. Sume does not rank the pitches, so this works best after you have a list of start times, for example from a transcript pass like the one in find product moments in a live-commerce replay.

Why use audio parts instead of trimming each clip?

Trimming is the right tool when you want each pitch as its own MP4. For a single reel, it adds a job per clip and then a render that stitches them. The Timeline 1.0 docs offer a shorter path: audio.parts[] takes up to 20 gapless slices, each with a url, an optional source_in and a duration, joined in the sample domain with no re-TTS.

The video side matches it. Each slot has source_url, an in-point source_in, an on-spine start and a duration. If the audio slice and the video slot use the same start time and length, the host's voice stays in sync with their lips for that pitch.

How do you detach the audio?

Call POST /v1/audio-detach with the replay's video_url. The default output is a sample-exact wav, which is what a timeline audio.url wants. The source can be up to 1800 seconds, but the output is capped at 900 seconds, so a whole track past 15 minutes needs a range.

For a 30-minute replay, detach a range that covers the pitches you want and subtract its start from each source_in into the wav. That offset arithmetic is yours to get right; the docs only state the 900-second cap and the range field. The video source_in values still use the replay's own clock.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: replay-audio-1" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/demo/replay.mp4",
    "format": "wav"
  }'

What does the one-render reel look like?

The script below builds three pitches of 10, 12 and 8 seconds into a 30-second reel and sends it to the unbilled plan endpoint. It assumes the detached wav and the replay share the same clock, which holds when the detach covers the whole track; every start time here is under 900 seconds, so a whole-track wav fits the output cap. Fades go on slots after the first, at 0.2 seconds with output.fps set to 30 so the fade is a whole number of frames.

If the plan returns what you expect, post the same document to POST /v1/timeline-1.0/render with an Idempotency-Key. Default output is 1080x1920, so a 16:9 live replay is cropped to fill the frame under the default fit of cover; set fit to contain or blur on a slot if you would rather keep the full picture.

import json, os, urllib.request

KEY = os.environ["SUME_API_KEY"]
BASE = "https://media.sume.com/artifacts/demo/"
PITCHES = [(212.4, 10), (455.0, 12), (810.0, 8)]  # (start s, length s)

parts, video, t = [], [], 0
for start, length in PITCHES:
    parts.append({"url": BASE + "replay.wav", "source_in": start,
                  "duration": length})
    slot = {"source_url": BASE + "replay.mp4", "source_in": start,
            "start": t, "duration": length}
    if video:
        slot["transition"] = {"type": "fade", "duration": 0.2}
    video.append(slot)
    t += length

doc = {"audio": {"parts": parts, "duration_seconds": t},
       "output": {"fps": 30}, "video": video}
req = urllib.request.Request(
    "https://api.sume.com/v1/timeline-1.0/plan", json.dumps(doc).encode(),
    {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"})
with urllib.request.urlopen(req) as r:
    print(json.load(r))

What does it cost, and what does it not do?

Rates below are public rates from the docs; confirm in GET /v1/catalog. The reel has no captions or product cards. Add captions to the finished file with video captions if you want them.

One limit to keep in mind: more than 12 slots makes the render chunk under the default strategy, and a plan covers up to 200 slots and 1800 seconds. A three-pitch reel is nowhere near that.

Highlight reel from one replay: jobs and rates (read 2026-10-02)
StepEndpointRateJobs for a 30-second reel
Detach audio oncePOST /v1/audio-detach$0.01 per job1 ($0.01)
Plan the timelinePOST /v1/timeline-1.0/planUnbilled1 ($0.00)
Render the reelPOST /v1/timeline-1.0/render$0.10 per output minute, rounded up1 ($0.10)
Per-clip trims, for comparisonPOST /v1/video-trim$0.02 per job3 ($0.06), plus a render still needed to join them

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume