AI fitness video maker: what to generate, what to film
An AI fitness video maker can make the coach, intro, B-roll and on-screen text, but not trustworthy exercise demos. What to film, what to generate.

An AI fitness video maker can produce a coach talking to camera, an intro, mood B-roll and on-screen text, but it is a poor source of exercise demonstrations: a video model draws the body from a prompt, so a generated squat can show the wrong form, the wrong rep count or a joint bending the wrong way. Film the movement demos that must be correct, and use AI for the parts around them.
Below is how that split works with Sume's API; in the Agents tab you can ask for the same video in plain words, and the agent asks before it spends. Facts come from the Video generation, Models, Video frames, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code. Nothing here is health or training advice.
Can AI generate exercise videos with correct form?
Not reliably. With image-to-video, only the first frame is your picture; every frame after it is generated from the prompt, and the model does not know your program. A clip can look smooth and still teach the wrong movement, so treat every generated rep as unchecked until a coach has watched it.
The closer option is motion control: you film one real rep and Kling 3.0 Motion Control moves a still of your presenter with that clip's motion. The driving video must be at a public HTTPS URL and at most 30 seconds, and the output length follows it. The body in the output is still generated, so check it too; Motion control API covers the request.
To check form, pull stills from a Sume clip with video frames: POST /v1/video-frames takes one media.sume.com clip and 1 to 24 times in at[], returns image files, and is unbilled.
How do I make a workout video with AI?
- Film the demos. Sume has no public upload for local files, so either cut your own footage in your own editor, or host a rep at a public HTTPS URL and use it as a motion control driving video: Sume job results come back as
media.sume.comfiles the render can use. - Coach to camera: voice the cues with
POST /v1/tts-1.0/generate(up to 20,000 characters per call), then send a still of the coach and that audio toPOST /v1/veed/fabric-1.0. Video models do not lip-sync to a later voice-over, so every speaking shot is a still plus audio. - B-roll and intro:
POST /v1/videoswith a prompt, or a still as thefirst_frameinframe_images, for shots where exact movement does not matter: a gym at dawn, shoes being laced. - Join the shots in one
POST /v1/timeline-1.0/render. In current code each clip's own sound is dropped: sound comes from the render's audio spine and an optionalsoundtrack. - Put exercise names, reps and rest times on screen as caption
cues, which burn exactly the text you send at the times you set.
How much does an AI fitness video cost?
Each part is billed separately, by what it produces. A B-roll clip's price depends on the model you choose.
| Part | Call | Price |
|---|---|---|
| Movement from your filmed rep | POST /v1/kling/3.0/motion-control | $0.1575 per output second |
| Coach voice | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Talking coach | POST /v1/veed/fabric-1.0 | $0.1875 per audio second (720p) |
| B-roll clip | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Join | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
| On-screen text | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
| Form-check stills | POST /v1/video-frames | Unbilled |
What are the limits?
- No video model accepts a
seed, so you cannot regenerate the exact same clip; keep the takes you approve. - Generated clips run up to 30 seconds on the longest models; check
supported_durationsonGET /v1/videos/models. - One render runs 1 to 1,800 seconds with 1 to 200 video slots, and takes only this workspace's
media.sume.comfiles. - In current code the caption job refuses a video over 60 seconds or one with no audio stream, so caption the finished, voiced cut in parts under a minute. Add captions to a long video shows the split.
- Fabric audio must be Sume-hosted and at most 10 MB, and its declared
duration_secondsat most 300. Sume's Avatar 1.0 presenters speak English only in current code; for other languages, setlanguageon the TTS call and use Fabric. For a presenter-led course, see AI training video generator.
Sources
Related posts
More in Use cases
- AI game music generator: one track per level, cut to WAV
An AI game music generator makes each level, menu or boss track from a text brief. How to brief contrasting scenes, get WAV files, and what it costs.
- Can AI make a lyric video? Yes, if you time the lines
AI can make a lyric video: pictures under the song, plus each lyric line burned in at the time it is sung. You supply the line timings.
- AI meditation music generator: calm tracks, longer sessions
An AI meditation music generator makes calm instrumental tracks from a text brief. How to prompt one, join tracks into a session, and add a voice.
- AI meditation video generator: voice, calm loop, music bed
An AI guided meditation video is a slow narration over a calm visual loop and a soft music bed that dips under the voice. How to build one, and costs.
Written by Sume