AI 360 rotation video: spin a product from two photos
From one photo, AI invents every side it can't see. For a 360 rotation video, pin the front and back as frames, make two half turns, and join them.

An AI 360 rotation video is a generated clip in which a product, or the camera around it, turns all the way around. From a single photo, the video model has to invent every side the photo doesn't show, labels included. To pin more of the turn, give it both sides: the front photo as the first frame and the back photo as the last frame make a half turn, and a second clip from the back to the front completes it.
The result is a flat video, not an interactive 360 viewer or a VR video. Facts come from Sume's Video generation, Video frames, and Timeline 1.0 docs, read on 2026-09-28. Anything described as current behavior is read from Sume's API code.
Can AI make a 360 video from one photo?
Yes, as a guess. Send the photo as the first_frame on POST /v1/videos and ask for a full turn in the prompt, for example "the bottle rotates 360 degrees on a turntable, static camera, plain white background". Only the first frame is your photo. Every later angle is generated, so the back, the sides, and any text on them are invented. Check them before you use the clip.
How do I make a 360 rotation from front and back photos?
- Shoot the front and the back on the same background, from the same distance and height, in the same light. Crop both to the shape you will ask for, such as 1:1.
- Clip 1: the front photo as
first_frame, the back photo aslast_frame, and a half turn in one direction in the prompt. - Clip 2: the back photo as
first_frame, the front photo aslast_frame, and the same direction and speed. - Join them in a Timeline 1.0 render: clip 1 at
start: 0and clip 2 right after it, with notransition. Clip 2 ends on the photo clip 1 starts from, so the joined video can loop; see Seamless loop AI video.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: spin-front-to-back-001" \
-d '{
"model": "seedance-2",
"prompt": "The perfume bottle turns 180 degrees clockwise on a turntable at a steady speed. Static camera, plain white background, soft studio light",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/front.jpg" }, "frame_type": "first_frame" },
{ "type": "image_url", "image_url": { "url": "https://example.com/back.jpg" }, "frame_type": "last_frame" }
],
"aspect_ratio": "1:1",
"resolution": "720p",
"duration": 5
}'Can I add side photos too?
Not as references in the same request. When a request carries both frame_images and input_references, the frames take precedence and the request is treated as image-to-video; in current code the references are then not sent to the model. To use more angles, chain quarter turns instead: front to right side, right side to back, back to left side, and left side to front. That is four clips joined in order, each its own generation.
How do I check the label and join the halves?
Extract stills from each finished clip with POST /v1/video-frames and compare them with your photos, especially midway through each half turn, where neither photo pins the view. Product logo warping in image-to-video covers what to look for. Then join the clips:
| Step | Call | Input | Cost |
|---|---|---|---|
| Two half turns | POST /v1/videos | Frame images at public HTTPS URLs | By model, per pricing_skus on GET /v1/videos/models |
| Label check | POST /v1/video-frames | One workspace media.sume.com clip and the times you name | Unbilled |
| Join | POST /v1/timeline-1.0/render | Workspace media.sume.com clips, 1–200 slots; audio.mode: "silence" for no sound | Reserved at $0.10 per output minute, rounded up to whole minutes |
What are the limits?
- Only the ends of each half turn are your photos. The angles between are generated: nothing promises correct geometry or an intact label, and the two halves may not meet exactly.
- A
last_frameneeds afirst_framein current code, andgrok-imagine-video-1.5takes no last frame. - Video frames and Timeline read only your workspace's
media.sume.comfiles. Generated clips qualify, because Sume mirrors outputs to its own media URLs: once a job iscompleted, read the clip's URL fromGET /v1/jobs/{id}/result. - Timeline's default output is 1080×1920. For square clips, set
output.widthandoutput.height, such as 1080 and 1080. - A Timeline render takes sound only from its audio spine and optional soundtrack. In current code each clip's own audio is dropped.
Sources
Related posts
More in Use cases
- AI avatar for events: a host for screens and lobby loops
An AI avatar can host an event's scripted parts: the welcome, housekeeping, sponsor thanks, and session intros as short 16:9 clips, joined or looped.
- AI avatar for healthcare: patient education videos
Use an AI avatar in healthcare for general patient education: short, clinician-reviewed videos, captioned for waiting rooms, with no patient data.
- AI avatar for website videos: a talking welcome clip
An AI avatar for a website is a short presenter video. Render a captioned talking clip in your page's shape, then embed the MP4 with a poster image.
- AI B-roll generator: make cutaway clips from a prompt
An AI B-roll generator makes cutaway clips from a text prompt or a still. Match your edit's shape, pick a length the model allows, and skip the sound.
Written by Sume