What is a video sales letter (VSL)? And making one with AI

A video sales letter (VSL) is a sales pitch delivered as one narrated video, from hook to offer. What goes in one, and how to build it with AI parts.

5 min readSume
All posts

A video sales letter (VSL) is a sales pitch delivered as a video: one written script, from hook to offer to call to action, read aloud over slides, footage or a presenter. It is a sales letter turned into narration, and it asks the viewer to take one action at the end, such as buying or booking a call.

You can build one from AI parts: a text-to-speech narration, presenter or B-roll clips, and one render that joins them. The Sume facts below come from the TTS schema in the Sume API reference and the Timeline 1.0 and Generate avatar video docs, read on 2026-09-28.

What goes into a video sales letter?

A VSL follows the order of a written sales letter, so the script does most of the work. A common outline:

  • Hook: one line that names who the video is for and what it promises.
  • Problem: the situation the viewer is in, in their words.
  • Story or proof: how the product came about, a demo, or results you can document.
  • Offer: what they get, the price, and any guarantee you actually give.
  • Call to action: one next step, repeated at the end.

How long is a video sales letter?

As long as the script, since the narration sets the pace; there is no fixed length. Write the script first, then time it. On Sume the hard limits are per part: one render can run up to 1,800 seconds (30 minutes), one TTS request takes up to 20,000 characters, and one avatar clip covers an estimated 4–60 seconds, so a presenter VSL is several clips. For a 30-second spot instead, see How to make a 30 second advertisement with AI.

How do I make a VSL with AI?

Make the sound first, then the pictures, then join them:

  • Narrate the script with POST /v1/tts-1.0/generate: send it as transcript with a voice selector. The whole letter can be one request, since audio up to 1,200 seconds is accepted.
  • Make the pictures: B-roll clips from a video model, or presenter clips from POST /v1/avatar-1.0/talking-video. Avatar speech is English only in the current code.
  • Render with POST /v1/timeline-1.0/render: the narration is audio, and the clips fill video[] slots at the times each line is spoken. Every URL must be a media.sume.com file in your workspace, such as the outputs of the steps above; Assemble a long-form video with the Timeline API covers the fields.
  • In the current code the render plays only its audio spine and an optional soundtrack; each clip's own audio is dropped. For a talking presenter, pull each clip's audio with POST /v1/audio-detach and join the parts (up to 20) as the spine. AI avatar video longer than 60 seconds walks through that.
  • Captions: in the current code the caption job refuses a source over 60 seconds or one without an audio stream, so a multi-minute VSL can't be captioned in one job; caption each presenter clip before the join.

What does an AI VSL cost on Sume?

Each part bills on its own, plus a 5.5% agent fee by default. B-roll clips from video models bill per second at their own rates; see API pricing.

From the TTS schema in the Sume API reference, Generate avatar video, Timeline 1.0 and API pricing, read 2026-09-28.
PartLimitPrice
Narration (TTS 1.0)Up to 20,000 characters per request; audio over 1,200 s fails$0.0475 per 1,000 characters
Presenter clip (Avatar 1.0)4–60 s per job, 720p, English speech in current code$0.184/s standard, $0.245/s plus, $0.55/s max (no product image)
Assembly (Timeline 1.0)1–1,800 s output, 1–200 video slots, up to 20 audio parts$0.10 per output minute

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume