Gym promo video with AI: from your own gym photos

Make a gym promo video without a shoot: animate photos of your own floor and classes, add a voiced offer, music and captions, and render a vertical cut.

5 min readSume
All posts

A gym promo video is a short, usually vertical clip that shows your real floor, classes and equipment and states one offer, made for Reels, TikTok, Shorts or a membership drive. You can make one from photos you already have: animate each photo into a few seconds of motion with AI, put a voiced offer and a music bed under the clips, burn the offer on screen as text, and export a 15 to 30 second vertical cut.

The Sume steps below come from the Video generation, Timeline 1.0, Video captions and Music Router docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code.

What are good gym promo video ideas?

Give each photo one small, plain movement:

  • The floor: a slow pan along the racks or the cardio line.
  • A class: the studio with the lights up, a slow push-in toward the instructor's spot.
  • Equipment close-ups: plates, ropes, a rower, chalk on a bar.
  • Your coaches and members, only with their consent to appear in an ad.
  • The offer, spoken and on screen, and only if it is real: a free first week, a joining fee waived, a class timetable.
  • Skip generated body transformations and before/after bodies. They show results no member achieved.

How long should a gym promo video be?

Each platform publishes its own recommended ad lengths; How long should a video ad be? collects them. At 15 to 30 seconds, a cut holds four to six shots of a few seconds each. In the render below, audio.duration_seconds sets the output length exactly, so you choose the length when you write the voice-over, not in an editor.

How do I make a gym promo video with AI?

  • Animate each photo: send it at a public HTTPS URL as the first_frame in frame_images on POST /v1/videos, with resolution and aspect_ratio: "9:16" set. Most catalog models top out at 15 seconds per clip.
  • Voice the offer with Sume text-to-speech, and make a bed with the Music Router (steer length in the prompt; duration is rejected).
  • Join them with one POST /v1/timeline-1.0/render. Its default output is 1080×1920, already vertical. Every URL must be a media.sume.com file in your workspace, such as the clips, voice and track Sume returned.
  • In current code the render keeps only the voice spine and the soundtrack; any sound a clip generated is dropped.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: gym-promo-001" \
  -d '{
    "audio": {
      "url": "https://media.sume.com/artifacts/artf_demo/offer.wav",
      "duration_seconds": 20
    },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/floor.mp4", "start": 0, "duration": 5 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/class.mp4", "start": 5, "duration": 5 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/racks.mp4", "start": 10, "duration": 5 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/entry.mp4", "start": 15, "duration": 5 }
    ],
    "soundtrack": {
      "url": "https://media.sume.com/artifacts/artf_demo/bed.mp3",
      "loop": true,
      "duck_db": 10
    }
  }'

How do I put the offer on screen?

Send the finished render's video_url to POST /v1/video-captions with cues: each cue is a text with a start and an end in seconds, and the job burns that copy at those times, with no speech-to-text. Or leave the cues out and let the job caption the voice-over. With no style, Latin text gets slam, which in current code sets the words in capitals, so proofread the offer on the render. Event promo video shows a full cue request.

In current code the caption job refuses a video longer than 60 seconds or one without an audio stream. A 15 to 30 second cut with a voice-over meets both.

How much does a gym promo video cost?

You pay per step from your workspace USD balance: four to six clips, one voice-over, one music bed, one render and one caption job.

From Video generation, Timeline 1.0, Video captions, the Sume API reference and the API pricing rate card, read 2026-09-29. Each rate is plus a 5.5% agent fee by default.
StepCallPrice
Animate each gym photoPOST /v1/videosBy model, at provider list × 1.25 per clip; see pricing_skus on GET /v1/videos/models
Voiced offerPOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
Music bedPOST /v1/music-router/generate$0.125 per audio
Join into one MP4POST /v1/timeline-1.0/render$0.10 per output minute
Offer text on screenPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume