Direct response video ads: what they are and how to make them

A direct response video ad asks for one action now, such as buy or sign up, and is judged by that response. Its structure, and how to make variants with AI.

5 min readSume
All posts

A direct response video ad asks the viewer to do one thing now (buy, sign up, book, install) and is judged by how many people do it, not by how many see it. So the video is built around that action: a hook, the problem, proof or a demo, an offer, and one call to action, made in several variants so the ad platform can test them.

The ad structure below is general practice. The steps for building it with an AI presenter come from Sume's Generate avatar video docs, read on 2026-09-28.

What is direct response advertising?

Advertising that asks for a measurable response and tracks it. Brand advertising tries to be remembered; a direct response ad carries an offer, a deadline or reason to act, and a way to respond on the spot: a link, a code, a phone number, an install button. Because each response can be counted, advertisers test versions against each other and keep the ones that get more responses per dollar.

Direct response marketing is the same idea across channels (mail, email, search, TV); a video ad is one format of it. The long-form version is a video sales letter.

What does a direct response video ad contain?

Five beats, in this order, each doing one job:

  • Hook: the first seconds name the viewer or the problem, so the right people keep watching.
  • Problem: the situation in the viewer's words.
  • Proof or demo: the product doing the thing, shown rather than claimed.
  • Offer: what they get and on what terms, stated exactly.
  • Call to action: one action, said and shown on screen.

How do I make direct response video ads with AI?

Keep the offer and call to action fixed, and vary only what the test is about, usually the hook; otherwise you can't tell which change moved the response. With Sume's Avatar 1.0, POST /v1/avatar-1.0/talking-video can hold every beat in one clip from one reusable avatar, as ordered video_inputs scenes. Each variant is then the same request with a different first scene:

From Generate avatar video, read 2026-09-28.
Ad beatAcross variantsAvatar 1.0 request
HookVaries: the thing under testThe first video_inputs scene, voice.type: "text" with a short script
Demo or proofFixedA voice.type: "silence" scene with a required duration
Offer and call to actionFixedThe closing voice.type: "text" scene
Product on screenFixedTop-level product_image (omit it for none)
Whole adSame length and formatPlanned at 4–60 seconds; aspect_ratio 1:1, 3:4, 9:16, 4:3 or 16:9

How do I produce enough variants to test?

Render one clip per hook with the same avatar, demo, offer and close. The request body is in multi-scene avatar videos; Hook variations for UGC ads shows how to swap only the opening over a shared body, and Ad variations for video ads covers how many versions to plan and how to queue them in one bulk run.

What can't AI do for a direct response ad?

  • Pick the winner. That comes from your ad platform's results, not from the video tool.
  • Speak other languages from the avatar route: talking videos speak English only in the current code.
  • Render 4:5 from the avatar route: its aspect ratios are 1:1, 3:4, 9:16, 4:3 and 16:9.
  • Promise results. Keep claims in the ad to ones you can back up; the offer has to be real.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume