Stitching AI Video Clips Into a 45-Second Ad With Sume Timeline
Most video models cap a clip at 15 seconds. Assemble several into one MP4 with Timeline at $0.10 per output minute, and what its warnings mean.

Most of Sume's video catalog caps a single clip at 15 seconds or less, only Seedance 2.5 and Wan 3.0 generate up to 30, and Veo 3.1 caps at 8 per Google's page. A 45-second ad is therefore several clips, and the assembly step decides whether it looks like one piece. Sume's Timeline 1.0 is the public assembly surface.
What Timeline takes
A Timeline request is one audio spine plus an ordered list of video slots, rendered to one MP4 on Sume's worker with ffmpeg. You send a declarative document, not filtergraphs. POST /v1/timeline-1.0/plan compiles it without billing, so you can check it first, and POST /v1/timeline-1.0/render renders it. There is no GET on the render route; poll GET /v1/jobs/{id}/status and read GET /v1/jobs/{id}/result.
The result carries video_url, duration_seconds, segment_count, billable_minutes and optional warnings.
Cost and defaults
The public rate is $0.10 per output minute, rounded up, and reserve is based on the audio spine's length. A 45-second ad is one billable minute, or $0.10. There is no model inference in assembly. The default output is 1080 by 1920 vertical, and an omitted frame rate follows the source clips, so mixed-frame-rate sources are met by repeating or dropping frames. The docs warn that this can judder.
Soft warnings, such as padded or looped short sources and snapped transitions, are not failures. Read them: a looped clip means a slot was longer than the footage you supplied.
Planning the cuts
Generate each shot at the length the slot needs, rather than trimming long clips, since you pay for every generated second. Match the frame rate across models where you can. Use the audio spine for the voice-over or music so the picture follows the sound, and put native-audio clips under it as scratch only.
For the shot choices that feed this, see which model fits the inputs you have.
Sources
Related posts
More in Media tools
- Sume STT words_truncated: the 20,000-word cap and what it means
Sume STT caps words[] at 20,000 entries and sets words_truncated and words_total if that is reached. A 600-second job should stay well under the cap.
- Swap the model in a clothing video with AI: H3 Max Recast
Recast swaps the person in a clip for a person from a photo, 5 to 30 seconds. What it keeps, what the docs leave open about the garment, and the Sume call.
- Transcript of a 10-minute client video before localizing it: cost
Sume video inspect transcribes at $0.01 per audio minute, with a 600-second hint cap, so 10 minutes is $0.10 plus compute. Set language_code, get sentences.
- TikTok ad caption budgets: 100 in TopView, 150 for Spark Pull, 4 lines
TikTok's TopView page caps ad captions at 100 characters, Spark Pull allows 150, and both display four lines. Limits by placement and short burned-in text.
Written by Sume