AI B-roll generator: make cutaway clips from a prompt

An AI B-roll generator makes cutaway clips from a text prompt or a still. Match your edit's shape, pick a length the model allows, and skip the sound.

5 min readSume
All posts

An AI B-roll generator makes short cutaway clips from a text prompt or a still image, so you can cover narration or an interview without filming or searching stock footage. Generate one shot per prompt, at your edit's aspect ratio (16:9 for a landscape edit, 9:16 for a vertical one) and at a length the model supports. Turn the clip's own sound off where the model allows, because B-roll plays under your audio.

Sume facts come from the Video generation and Video Router docs, read on 2026-09-28; per-model sound and length come from the catalog behind GET /v1/videos/models and the checks POST /v1/videos runs. Talking head vs B-roll explains the two kinds of footage.

How do I generate B-roll with AI?

  • Write one shot per prompt: the subject, the action, the camera, and the light. Sume's docs suggest details about motion, camera angles, lighting, and scene composition.
  • From text: send the prompt to POST /v1/videos. With model: "sume/auto", Sume picks the model; Auto defaults to 720p and 8 seconds, with 3–10 second clips at 16:9 or 9:16.
  • From a still: when the shot must show a real product or place, send its picture as the first_frame in frame_images. Sume's own model guide makes B-roll as an Auto image, checked, then animated with Auto video.
  • Send the same aspect_ratio and resolution on every request, so every cutaway is requested at your edit's shape and size.
  • Pin a model where sound is optional and send generate_audio: false:
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: broll-kneading-001" \
  -d '{
    "model": "kling-3",
    "prompt": "Slow push-in on hands kneading dough on a floured wooden table, warm morning light, shallow depth of field",
    "aspect_ratio": "16:9",
    "resolution": "720p",
    "duration": 6,
    "generate_audio": false
  }'

Which models can make B-roll from a text prompt?

Every model in the table except grok-imagine-video-1.5, which needs a first frame. All the others list both 16:9 and 9:16; Grok takes no aspect_ratio.

From Video generation, Video Router, and the catalog behind GET /v1/videos/models, read 2026-09-28.
ModelFrom text aloneClip lengthSound
seedance-2.5Yes4–30 sOptional
seedance-2, seedance-2-fast, seedance-2-miniYes4–15 sOptional
wan-3.0Yes2–30 sOptional
kling-3Yes4–15 sOptional
minimax-h3, minimax-h3-maxYes5–15 sAlways on; false is refused
gemini-omni-flash-1.1Yes3–10 sAlways on; false is refused
grok-imagine-video-1.5No, it needs a first frame4–15 sNone

Should B-roll clips have sound?

Usually not, because your voice-over or interview is the soundtrack. If you cut the edit in Sume, Timeline 1.0 builds one MP4 from an audio spine plus ordered video slots, and in current code it takes sound only from that spine and an optional soundtrack: each clip's own audio is dropped. In another editor, you can mute the clips from models where sound is always on. Faceless video API: voiceover, B-roll, music, and captions shows the Timeline edit.

How much does AI B-roll cost?

It is not free stock footage: each clip is generated for you and billed by the model, at the provider's list price × 1.25, reserved on submit, plus a 5.5% agent fee by default. Sound can change the rate: kling-3 is $0.14 per second without sound and $0.21 with it. Every model's rates are in pricing_skus on GET /v1/videos/models, and How much does an AI video cost? compares them.

How do I get the clips into my editor?

Poll GET /v1/videos/{id} until the status is completed, then download from GET /v1/videos/{id}/content with your API key. That route redirects to the file, so tell curl to follow it with -L. For a Sume edit instead, read each clip's media.sume.com URL from GET /v1/jobs/{id}/result, since Timeline takes only this workspace's Sume-hosted files.

What are the limits?

  • One clip runs 2–30 seconds depending on the model; longer cutaways are two clips.
  • Stills sent as frames must be public HTTPS URLs.
  • Every frame after the first is generated, so check each clip for anything that looks wrong before it goes in the edit.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume