AI B-roll generator: make cutaway clips from a prompt
An AI B-roll generator makes cutaway clips from a text prompt or a still. Match your edit's shape, pick a length the model allows, and skip the sound.

An AI B-roll generator makes short cutaway clips from a text prompt or a still image, so you can cover narration or an interview without filming or searching stock footage. Generate one shot per prompt, at your edit's aspect ratio (16:9 for a landscape edit, 9:16 for a vertical one) and at a length the model supports. Turn the clip's own sound off where the model allows, because B-roll plays under your audio.
Sume facts come from the Video generation and Video Router docs, read on 2026-09-28; per-model sound and length come from the catalog behind GET /v1/videos/models and the checks POST /v1/videos runs. Talking head vs B-roll explains the two kinds of footage.
How do I generate B-roll with AI?
- Write one shot per prompt: the subject, the action, the camera, and the light. Sume's docs suggest details about motion, camera angles, lighting, and scene composition.
- From text: send the prompt to
POST /v1/videos. Withmodel: "sume/auto", Sume picks the model; Auto defaults to 720p and 8 seconds, with 3–10 second clips at 16:9 or 9:16. - From a still: when the shot must show a real product or place, send its picture as the
first_frameinframe_images. Sume's own model guide makes B-roll as an Auto image, checked, then animated with Auto video. - Send the same
aspect_ratioandresolutionon every request, so every cutaway is requested at your edit's shape and size. - Pin a model where sound is optional and send
generate_audio: false:
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: broll-kneading-001" \
-d '{
"model": "kling-3",
"prompt": "Slow push-in on hands kneading dough on a floured wooden table, warm morning light, shallow depth of field",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 6,
"generate_audio": false
}'Which models can make B-roll from a text prompt?
Every model in the table except grok-imagine-video-1.5, which needs a first frame. All the others list both 16:9 and 9:16; Grok takes no aspect_ratio.
| Model | From text alone | Clip length | Sound |
|---|---|---|---|
seedance-2.5 | Yes | 4–30 s | Optional |
seedance-2, seedance-2-fast, seedance-2-mini | Yes | 4–15 s | Optional |
wan-3.0 | Yes | 2–30 s | Optional |
kling-3 | Yes | 4–15 s | Optional |
minimax-h3, minimax-h3-max | Yes | 5–15 s | Always on; false is refused |
gemini-omni-flash-1.1 | Yes | 3–10 s | Always on; false is refused |
grok-imagine-video-1.5 | No, it needs a first frame | 4–15 s | None |
Should B-roll clips have sound?
Usually not, because your voice-over or interview is the soundtrack. If you cut the edit in Sume, Timeline 1.0 builds one MP4 from an audio spine plus ordered video slots, and in current code it takes sound only from that spine and an optional soundtrack: each clip's own audio is dropped. In another editor, you can mute the clips from models where sound is always on. Faceless video API: voiceover, B-roll, music, and captions shows the Timeline edit.
How much does AI B-roll cost?
It is not free stock footage: each clip is generated for you and billed by the model, at the provider's list price × 1.25, reserved on submit, plus a 5.5% agent fee by default. Sound can change the rate: kling-3 is $0.14 per second without sound and $0.21 with it. Every model's rates are in pricing_skus on GET /v1/videos/models, and How much does an AI video cost? compares them.
How do I get the clips into my editor?
Poll GET /v1/videos/{id} until the status is completed, then download from GET /v1/videos/{id}/content with your API key. That route redirects to the file, so tell curl to follow it with -L. For a Sume edit instead, read each clip's media.sume.com URL from GET /v1/jobs/{id}/result, since Timeline takes only this workspace's Sume-hosted files.
What are the limits?
- One clip runs 2–30 seconds depending on the model; longer cutaways are two clips.
- Stills sent as frames must be public HTTPS URLs.
- Every frame after the first is generated, so check each clip for anything that looks wrong before it goes in the edit.
Sources
Related posts
More in Use cases
- AI birthday video from photos: animate, add music and text
Make an AI birthday video by animating a few photos into short clips, joining them over an original instrumental, and burning your message on screen.
- AI children's book illustrations: same character, every page
Settle the main character in one picture, reuse it as a reference on every page, and generate each page at the printed size with room for text.
- AI clothing model generator: put your garment on a model
An AI clothing model generator dresses a model in your garment: send a flat lay and a model photo to an image model, then check print and color.
- AI comic generator: consistent characters, panel by panel
Make an AI comic one panel at a time: reuse one reference image per character, give each panel its shape, and letter the speech bubbles yourself.
Written by Sume