Six dental education Shorts from stills, voice and captions

A patient-education video from four stills, a 750-character voice-over and burned captions, with no video model. Cost per video and for six topics.

3 min readSume
All posts

A 50-second education video made from four generated stills, a 750-character voice-over, one Timeline render and one caption job costs about $0.40 on Sume ($0.4016 before rounding); six topics cost about $2.41. It uses no video model, because Timeline holds stills as static slots.

The recipe

Stills at GPT Image 2.5 medium and 1024 by 1024 cost $0.0165 each. Timeline accepts a still as source_url and holds it; a motion field on a still is accepted but ignored and reported as motion_ignored. Fit cover fills the 9:16 frame.

{
  "audio": { "url": "https://media.sume.com/artifacts/artf_demo/floss.wav", "duration_seconds": 50 },
  "video": [
    { "source_url": "https://media.sume.com/artifacts/artf_demo/s1.png", "start": 0,  "duration": 12 },
    { "source_url": "https://media.sume.com/artifacts/artf_demo/s2.png", "start": 12, "duration": 13,
      "transition": { "type": "dissolve", "duration": 0.5 } },
    { "source_url": "https://media.sume.com/artifacts/artf_demo/s3.png", "start": 25, "duration": 13,
      "transition": { "type": "dissolve", "duration": 0.5 } },
    { "source_url": "https://media.sume.com/artifacts/artf_demo/s4.png", "start": 38, "duration": 12,
      "transition": { "type": "dissolve", "duration": 0.5 } }
  ]
}

Cost per video

Sume list price, read 2026-10-08
ItemQtyRateSubtotal
Stills, GPT Image 2.5 medium 1024x10244$0.0165$0.066
Voice-over, 750 characters1$0.0475 per 1,000 chars$0.0356
Timeline render, 50 s1$0.10 per minute$0.10
Captions on the finished video1$0.20 per job$0.20
Total per video$0.4016

Content discipline

Keep it general education, such as how long to brush or why flossing matters, and have a clinician review each script. Do not generate before-and-after smiles or results that imply a specific outcome for a patient.

Captions are worth the $0.20 because a video you publish should read without sound, and the caption job reads the speech for you.

Making stills that fit

Ask for simple, friendly illustrations of a toothbrush, floss or a clean smile rather than clinical imagery, and keep the same style words in all four prompts so the set looks like one video. GPT Image 2.5 stills at 1024 by 1024 are cropped to the vertical frame by cover fit, so keep the subject centered. If you want a calmer feel, add a dissolve of half a second between slots. A clinician should read the script and the captions before publishing, because wrong dental advice is worse than no video.

Related posts

More in Use cases

All Use cases posts

Written by Sume