Online course trailer: avatar intro plus b-roll in one timeline render
Assemble a 30-second course trailer: an avatar intro, generated b-roll clips, and one audio spine rendered into a single MP4 with Sume's Timeline 1.0.
A course creator can build a 30-second trailer by making an avatar intro, generating two or three b-roll clips, and rendering them together with Sume's Timeline 1.0. Timeline takes one audio spine and ordered video slots and returns one MP4. The docs list its public rate as $0.10 per ceil(output minute), with no provider inference, since it only runs ffmpeg.
The generation jobs cost what their own model lines say. The assembly is the cheap part, and it is the part that stops a trailer from being four separate files.
Why assemble instead of generating one long clip
Trend roundups say multimodal workflows that connect text-to-video, image-to-video, voice, scripts and editing are the defining shift, and that 5 to 30 seconds is the practical range for polished clips (AI Video Generation Trends, read 2026-10-07). A trailer is that shape: a spoken hook, then pictures that illustrate it.
Separate jobs also let you redo one piece. If the second b-roll clip is wrong, you regenerate that clip, not the whole trailer.
The parts
Make the intro with the avatar talking-video route: a 9:16 script of about 10 seconds. Make b-roll with POST /v1/videos, where each clip is a visual of the course topic. Every file you place on the timeline must already be a media.sume.com artifact or asset of your workspace, so import anything external with POST /v1/media-imports first.
Use the avatar clip's audio as the spine, or write a separate voiceover. Timeline needs audio.duration_seconds between 1 and 1800 and an audio file unless you declare silence.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: course-trailer-001" \
-d '{
"audio": {
"url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
"duration_seconds": 30
},
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 10 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/broll-1.mp4", "start": 10, "duration": 10,
"transition": { "type": "fade", "duration": 0.25 } },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/broll-2.mp4", "start": 20, "duration": 10,
"transition": { "type": "fade", "duration": 0.25 } }
]
}'Plan first, then render
POST /v1/timeline-1.0/plan is an unbilled compile preflight. It checks the schema and the Sume-host URLs and returns the duration, segment count, billable minutes and an estimated cost, without creating a job. Run it on your document before you render.
The default output is 1080 by 1920. Note that a plan cannot predict short-source pad or loop warnings, which show up as soft warnings on the finished job.
| Item | Value | Note |
|---|---|---|
| Audio length | 1 to 1800 seconds | audio.duration_seconds is required |
| Video slots | 1 to 200 | Ordered, first start must be 0 |
| Transitions | fade, wipeleft, wiperight, slideup, slidedown, dissolve | Only on slots after the first |
| Default output | 1080 by 1920 MP4 | Set output width and height to change |
| Public rate | $0.10 per ceil(output minute) | Confirm in GET /v1/catalog |
What to check before you launch
Watch the whole trailer with sound off, then with sound on. Make sure the b-roll does not promise results the course cannot deliver, and that the avatar is disclosed as AI where your platform or market expects it.
Sources
Related posts
More in Use cases
- Outage status update video with an AI avatar: script and cost
A 20 to 30 second avatar clip can front a status-page update. What to script, what it costs by tier, and why the written update must go out first.
- How to make a pep talk audio track with music: voice, bed and render
A 90-second pep talk with a music bed is one TTS job, one Music Router job and one render: about $0.39 on Sume. Suno Speech beta does it in one pass.
- Personal trainer class promo: avatar hook, silent demo beat, CTA
Build an 11-second class promo for a fitness studio from three ordered scenes on one avatar, with a silence beat for the demo. Limits and request body included.
- Photo to talking selfie clip: put the spoken line in the Omni prompt
fal's Omni 1.1 example ends its prompt with a quoted spoken line. Here is that pattern on Sume from a still, the cost of 6 seconds, and a transcript check.
Written by Sume