AI 360 rotation video: spin a product from two photos

From one photo, AI invents every side it can't see. For a 360 rotation video, pin the front and back as frames, make two half turns, and join them.

5 min readSume
All posts

An AI 360 rotation video is a generated clip in which a product, or the camera around it, turns all the way around. From a single photo, the video model has to invent every side the photo doesn't show, labels included. To pin more of the turn, give it both sides: the front photo as the first frame and the back photo as the last frame make a half turn, and a second clip from the back to the front completes it.

The result is a flat video, not an interactive 360 viewer or a VR video. Facts come from Sume's Video generation, Video frames, and Timeline 1.0 docs, read on 2026-09-28. Anything described as current behavior is read from Sume's API code.

Can AI make a 360 video from one photo?

Yes, as a guess. Send the photo as the first_frame on POST /v1/videos and ask for a full turn in the prompt, for example "the bottle rotates 360 degrees on a turntable, static camera, plain white background". Only the first frame is your photo. Every later angle is generated, so the back, the sides, and any text on them are invented. Check them before you use the clip.

How do I make a 360 rotation from front and back photos?

  • Shoot the front and the back on the same background, from the same distance and height, in the same light. Crop both to the shape you will ask for, such as 1:1.
  • Clip 1: the front photo as first_frame, the back photo as last_frame, and a half turn in one direction in the prompt.
  • Clip 2: the back photo as first_frame, the front photo as last_frame, and the same direction and speed.
  • Join them in a Timeline 1.0 render: clip 1 at start: 0 and clip 2 right after it, with no transition. Clip 2 ends on the photo clip 1 starts from, so the joined video can loop; see Seamless loop AI video.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: spin-front-to-back-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "The perfume bottle turns 180 degrees clockwise on a turntable at a steady speed. Static camera, plain white background, soft studio light",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/front.jpg" }, "frame_type": "first_frame" },
      { "type": "image_url", "image_url": { "url": "https://example.com/back.jpg" }, "frame_type": "last_frame" }
    ],
    "aspect_ratio": "1:1",
    "resolution": "720p",
    "duration": 5
  }'

Can I add side photos too?

Not as references in the same request. When a request carries both frame_images and input_references, the frames take precedence and the request is treated as image-to-video; in current code the references are then not sent to the model. To use more angles, chain quarter turns instead: front to right side, right side to back, back to left side, and left side to front. That is four clips joined in order, each its own generation.

How do I check the label and join the halves?

Extract stills from each finished clip with POST /v1/video-frames and compare them with your photos, especially midway through each half turn, where neither photo pins the view. Product logo warping in image-to-video covers what to look for. Then join the clips:

From Video generation, Video frames, Timeline 1.0, and API pricing, read 2026-09-28.
StepCallInputCost
Two half turnsPOST /v1/videosFrame images at public HTTPS URLsBy model, per pricing_skus on GET /v1/videos/models
Label checkPOST /v1/video-framesOne workspace media.sume.com clip and the times you nameUnbilled
JoinPOST /v1/timeline-1.0/renderWorkspace media.sume.com clips, 1–200 slots; audio.mode: "silence" for no soundReserved at $0.10 per output minute, rounded up to whole minutes

What are the limits?

  • Only the ends of each half turn are your photos. The angles between are generated: nothing promises correct geometry or an intact label, and the two halves may not meet exactly.
  • A last_frame needs a first_frame in current code, and grok-imagine-video-1.5 takes no last frame.
  • Video frames and Timeline read only your workspace's media.sume.com files. Generated clips qualify, because Sume mirrors outputs to its own media URLs: once a job is completed, read the clip's URL from GET /v1/jobs/{id}/result.
  • Timeline's default output is 1080×1920. For square clips, set output.width and output.height, such as 1080 and 1080.
  • A Timeline render takes sound only from its audio spine and optional soundtrack. In current code each clip's own audio is dropped.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume