AI fitness video maker: what to generate, what to film

An AI fitness video maker can make the coach, intro, B-roll and on-screen text, but not trustworthy exercise demos. What to film, what to generate.

5 min readSume
All posts

An AI fitness video maker can produce a coach talking to camera, an intro, mood B-roll and on-screen text, but it is a poor source of exercise demonstrations: a video model draws the body from a prompt, so a generated squat can show the wrong form, the wrong rep count or a joint bending the wrong way. Film the movement demos that must be correct, and use AI for the parts around them.

Below is how that split works with Sume's API; in the Agents tab you can ask for the same video in plain words, and the agent asks before it spends. Facts come from the Video generation, Models, Video frames, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code. Nothing here is health or training advice.

Can AI generate exercise videos with correct form?

Not reliably. With image-to-video, only the first frame is your picture; every frame after it is generated from the prompt, and the model does not know your program. A clip can look smooth and still teach the wrong movement, so treat every generated rep as unchecked until a coach has watched it.

The closer option is motion control: you film one real rep and Kling 3.0 Motion Control moves a still of your presenter with that clip's motion. The driving video must be at a public HTTPS URL and at most 30 seconds, and the output length follows it. The body in the output is still generated, so check it too; Motion control API covers the request.

To check form, pull stills from a Sume clip with video frames: POST /v1/video-frames takes one media.sume.com clip and 1 to 24 times in at[], returns image files, and is unbilled.

How do I make a workout video with AI?

  • Film the demos. Sume has no public upload for local files, so either cut your own footage in your own editor, or host a rep at a public HTTPS URL and use it as a motion control driving video: Sume job results come back as media.sume.com files the render can use.
  • Coach to camera: voice the cues with POST /v1/tts-1.0/generate (up to 20,000 characters per call), then send a still of the coach and that audio to POST /v1/veed/fabric-1.0. Video models do not lip-sync to a later voice-over, so every speaking shot is a still plus audio.
  • B-roll and intro: POST /v1/videos with a prompt, or a still as the first_frame in frame_images, for shots where exact movement does not matter: a gym at dawn, shoes being laced.
  • Join the shots in one POST /v1/timeline-1.0/render. In current code each clip's own sound is dropped: sound comes from the render's audio spine and an optional soundtrack.
  • Put exercise names, reps and rest times on screen as caption cues, which burn exactly the text you send at the times you set.

How much does an AI fitness video cost?

Each part is billed separately, by what it produces. A B-roll clip's price depends on the model you choose.

From Video generation, Timeline 1.0, Video captions and the Sume API reference, read 2026-09-29. Each rate is plus a 5.5% agent fee by default; see API pricing.
PartCallPrice
Movement from your filmed repPOST /v1/kling/3.0/motion-control$0.1575 per output second
Coach voicePOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
Talking coachPOST /v1/veed/fabric-1.0$0.1875 per audio second (720p)
B-roll clipPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
JoinPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
On-screen textPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds
Form-check stillsPOST /v1/video-framesUnbilled

What are the limits?

  • No video model accepts a seed, so you cannot regenerate the exact same clip; keep the takes you approve.
  • Generated clips run up to 30 seconds on the longest models; check supported_durations on GET /v1/videos/models.
  • One render runs 1 to 1,800 seconds with 1 to 200 video slots, and takes only this workspace's media.sume.com files.
  • In current code the caption job refuses a video over 60 seconds or one with no audio stream, so caption the finished, voiced cut in parts under a minute. Add captions to a long video shows the split.
  • Fabric audio must be Sume-hosted and at most 10 MB, and its declared duration_seconds at most 300. Sume's Avatar 1.0 presenters speak English only in current code; for other languages, set language on the TTS call and use Fabric. For a presenter-led course, see AI training video generator.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume