AI video looks fake? Check Seedance 2.5 frames for skin and light
ByteDance says Seedance 2.5 tuned textures, skin, eyes, lighting and stray subtitles. How to pull stills from a clip on Sume and check each one.

To check whether a Seedance 2.5 clip looks artificial, pull a handful of full-size stills and inspect skin, eyes, lighting, colour saturation and any stray text, which are the areas ByteDance says it tuned. On Sume, Video frames returns durable PNG or JPEG stills at the clip's own size for times you name. ByteDance's launch post says the model optimises object textures, skin and eye features, lighting and colour saturation to avoid the overly artificial look of AI video, and reduces uncontrolled subtitles and background music. Whether it worked on your clip is something a frame check answers.
What did ByteDance say it improved?
From the Seedance 2.5 launch post, read 2026-10-03: the team "systematically optimizes" object textures, skin and eye features, lighting and colour saturation, and minimises uncontrolled occurrences in subtitles and background music, to get results that resemble live-action footage. The post also notes remaining weakness in complex physical motion.
| Claim | What to look at in a still | Frames to pull |
|---|---|---|
| Skin and eye features | Pores, eye reflections, teeth, hairline | A close-up moment, full size |
| Lighting | Shadow direction and softness across the subject | Start, middle and end |
| Colour saturation | Skin tone and highlights not oversaturated | Any frame with a neutral grey or white object |
| Textures | Fabric, wood, metal keep detail | A still where the camera is slow |
| Subtitles and music | Burned-in text that you did not ask for | Every few seconds |
How do you pull the stills?
Video frames takes one workspace media.sume.com clip and exactly one of at[] (1 to 24 seconds, each at least 0) or fps (above 0, up to 2, capped at 24 frames). format is jpeg by default or png for lossless inspection, and leaving out max_edge keeps the source frame size. The submit is always 202; poll until resource_status is ready. A time outside [0, duration) fails with frame_time_out_of_range.
Pick times that matter: a close-up, a lighting change, and the last second, where a long clip is most likely to drift.
import os, requests
r = requests.post(
"https://api.sume.com/v1/video-frames",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "fake-check-001",
},
json={
"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"at": [1, 8, 15, 22, 29],
"format": "png",
},
)
print(r.status_code, r.json())
When is Video inspect the better first step?
Video inspect is the faster skim: with no frames field you get eight mid-bin stills at a default long edge of 768, plus probe facts including whether the clip has audio. Use it to decide which moments deserve a full-size pull. Its seek: "fast" mode snaps each still to an earlier keyframe, so use the default precise seek when the exact instant matters.
What about stray text and sound?
Text shows up in stills, so look for it in every frame you pull. Sound is separate: a frames: false Video inspect returns probe facts including probe.has_audio, so you can see whether a track exists without pulling any stills. If you did not want audio, generate_audio is an optional field on the request, and it defaults to the model's own audio capability, so set it explicitly when the track matters.
A transcript is available from Video inspect with transcribe: true, billed at $0.01 per audio minute per the doc. Use it only when speech is expected; a silent clip returns inspect_source_has_no_audio.
What do you do if the clip fails the check?
Rerun, since Sume does not accept a seed and a rerun is a new take. Change one thing at a time: add a reference image of the person or product, simplify the lighting description, or choose a higher resolution (seedance-2.5 offers 480p, 720p and 1080p, 4 to 30 seconds). If stray text appears, say in the prompt that no text should appear, then check again. Stop when the stills pass your own checklist, not when the clip merely looks fine at thumbnail size.
Sources
Related posts
More in Models
- How to choose an AI video model: five questions before you pin one
Length, resolution, audio, start image and references decide the model, not the leaderboard. Five questions mapped to Sume's video catalog.
- Which AI video models can't do text-to-video? Rows that need a source
Grok Imagine Video 1.5, Genjutsu Motion Transfer and H3 Max Recast refuse a prompt-only request. What each needs, and which rows accept text alone.
- AI video models with 4:3 and 3:4 aspect ratios on Sume
Seedance 2.5 and 2.0, Wan 3.0, MiniMax H3 and H3 Max list 4:3 and 3:4 on Sume; Kling 3.0 and Gemini Omni Flash 1.1 do not. Prices for the per-second rows.
- AI video reference limits: how many images and clips per model
Reference image and reference video caps for Wan 3.0, MiniMax H3, Gemini Omni Flash, Genjutsu and H3 Max Recast on Sume, in one table with the odd limits.
Written by Sume