AI Product Video Looks the Wrong Size in Hand: Scale Reference Fix
A bottle that looks twice its real size in a generated hand is a scale error. Add a size reference photo, state the dimensions, and verify with stills.

When a generated hand holds your product and the product looks too large or too small, the model has no idea how big the real thing is. A product photo on white carries no scale. The fix is to give it scale in two places: one reference photo that shows the product beside something of known size, and the real dimensions in the prompt. Then measure the result from stills instead of trusting your eye on a phone.
This is a working method, not a benchmark: no model page read on 2026-10-03 promises accurate scale. Sources: the Sume Videos docs, Google's Gemini Omni docs, and Sume's catalog constraints in the repository.
Give the model a size to work from
Sume's Gemini Omni Flash 1.1 accepts up to 10 reference images and a prompt of up to 20,000 characters, and the reference tokens are 0-based (<IMAGE_REF_0>). Google's Omni page describes subject reference as combining separate images, such as a cat and yarn, into one scene. Use that: one image of the product alone, one of the product next to a coin or a hand.
- Reference 0: the product alone, label forward.
- Reference 1: the same product beside an object everyone knows the size of, such as a credit card.
- Prompt: state the real size, for example "a 7 cm lip balm tube held between thumb and forefinger, as in <IMAGE_REF_1>".
- Ask for a hand in frame at the start, so the scale is set before the camera moves.
Measure from stills
Extract stills at fixed times with Video frames, which returns source-size images from a clip on media.sume.com. Open the first and the last still and compare the product's height with the width of the hand or card. A product that spans most of the palm is not a 7 cm tube. If it is off, regenerate with a tighter prompt rather than cropping, since cropping only hides the error.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: scale-check-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/hand.mp4",
"at": [0.5, 3, 5.5],
"format": "png"
}'What it costs to try again
At the list rate of $0.10 per second for Omni at 720p (2026-08-28), a 5-second retry is $0.50 before Sume's margin, and the frames call is billed by its own compute. Sume has no seed on any video model, so a retry is a new generation, not a small edit. Budget two or three tries per hero shot and keep the scale reference photo the same each time. The label and fingers still need the four-frame check.
Sources
Related posts
More in Models
- AI video audio by model: toggle, always on, or none on Sume
Seedance, Wan 3.0 and Kling 3 let you switch audio; MiniMax H3 and Gemini Omni Flash always make it; Grok Imagine has none. What each row does.
- AI video looks fake? Check Seedance 2.5 frames for skin and light
ByteDance says Seedance 2.5 tuned textures, skin, eyes, lighting and stray subtitles. How to pull stills from a clip on Sume and check each one.
- How to choose an AI video model: five questions before you pin one
Length, resolution, audio, start image and references decide the model, not the leaderboard. Five questions mapped to Sume's video catalog.
- Which AI video models can't do text-to-video? Rows that need a source
Grok Imagine Video 1.5, Genjutsu Motion Transfer and H3 Max Recast refuse a prompt-only request. What each needs, and which rows accept text alone.
Written by Sume