School promotional video with AI: show the real campus
A school promotional video should show the real campus and programs. Use AI for motion, voice and text on your own photos, not a generated campus. Steps, costs.

A school promotional video should show the real school: your campus, your programs, and students and staff who agreed to appear, ending on the open-day date and the application link. AI's honest job is motion, voice and text: animate your own campus photos into short clips, add a narrator or an on-camera presenter, and cut a vertical and a wide version. A generated campus or generated students would show families a school that doesn't exist.
The Sume facts below come from the Video generation, Media inputs, Generate avatar video, Timeline 1.0 and Video captions docs and Sume's Terms of Service, read on 2026-09-29. Points marked as current behavior are read from Sume's code. The same steps fit a college, university or training academy.
What should a school promo video include?
Pick one audience per video, such as families applying next year, and answer what they ask first:
- Who the school is for: age range, focus, and what makes a day there different, in one or two sentences.
- What a day looks like: classrooms, labs, the library, sports fields, from your own photos.
- Programs: one short shot and one line per program you want to feature.
- A welcome from the head of school or an admissions officer.
- The next step: open-day date, application deadline and link, shown on screen as exact text.
- For one event only, such as an open day, event promo video covers the poster-based version.
Can I use AI-generated students or campus shots?
Not as if they were your school. Sume's terms say generated outputs should not be treated as a factual record of real people or events, and not to use Sume to mislead people about the authenticity of generated media. Families judge a school by its buildings and people, so both should be real.
Real students and staff need permission. The terms say you represent that you have permission to use any person's likeness or voice. For minors, get a parent's or guardian's consent under your school's own media policy before a photo goes into any tool.
How do I make a school promotional video from campus photos?
Only the first frame of each clip is your photo; what moves after it is generated, so check every clip for rooms, signs or people that aren't yours.
- Put each campus photo at a public HTTPS URL and send it as the
first_frameinframe_imagesonPOST /v1/videos, with a prompt for one slow camera move. Most catalog models top out at 15 seconds per clip. - For a presenter, send a script and your avatar's
avatar_handletoPOST /v1/avatar-1.0/talking-videowith a campus photo as thescene(type: "photo"). Sume accepts a script when it estimates the video at 4 to 60 seconds; 720p is the documented resolution. In current code the avatar speaks English only, so use text-to-speech narration withlanguagefor other languages. - Join the clips under the voice with a Timeline 1.0 render. It reads only your workspace's
media.sume.comfiles, and in current code each clip's own audio is dropped, so a presenter's speech has to come in as the voice track: pull it out withPOST /v1/audio-detachfirst. - Render twice for two shapes: output width and height are any even numbers from 256 to 2160, and the default is 1080×1920.
video[].fittakescover,contain,stretchorblurfor clips made in the other shape. - Burn the dates and link as authored
cues. In current code the caption job refuses videos over 60 seconds or without an audio stream.
How much does a school promotional video cost?
Each piece is billed per call from one prepaid balance; a second size adds one more render and one more caption job.
| Piece | Call | Price |
|---|---|---|
| Campus photo into a moving clip | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Presenter clip, default quality | POST /v1/avatar-1.0/talking-video | $0.245 per second |
| Narration instead of a presenter | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Join clips and voice (once per size) | POST /v1/timeline-1.0/render | $0.10 per output minute |
| Dates and application link on screen | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
Sources
Related posts
More in Use cases
- Skincare video ads with AI: three Formats, one packshot
Skincare video ads can be made with AI from a packshot: a model-and-product studio film, an application or texture demo, or a before-and-after with proof.
- Virtual twilight: turn a daytime exterior into a dusk photo
Virtual twilight turns a daytime house photo into a dusk shot with a sunset sky and lit windows. How to make one with AI, and what to check after.
- Safety training videos for employees, made with AI
Make safety training videos for employees with an AI presenter: one hazard per clip, your own site photos, captions, and versions in your crew's languages.
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
Written by Sume