Edits keyframes vs Sume Timeline: stills are static, no zoom
Edits supports keyframes. Sume Timeline holds a still static and ignores a motion field; animate the still first with Video Router image-to-video.

Instagram's Edits page lists keyframe support among the app's features (read 2026-10-03), so a slow push-in on a photo is a few taps there. Sume Timeline 1.0 cannot do that: it treats a still as a static hold, accepts a motion field on a slot and ignores it, reporting motion_ignored as a soft warning rather than an error. If you want a still to move, animate it first into a clip, then place the clip. The docs list one route for that: image-to-video on the Video Router.
What a Timeline slot actually does with a still
video[].source_url can be a hosted still, which holds for the slot's duration with fit set to cover, contain, stretch or blur. Transitions between slots are limited to fade, wipeleft, wiperight, slideup, slidedown and dissolve, at most one second and half the shorter neighbour. Nothing in that list is a zoom or pan, so a render of stills is a slideshow with transitions.
| Need | Edits | Sume |
|---|---|---|
| Slow push-in on a photo | Keyframes | Not on Timeline; generate a clip first |
| Slide between photos | Transitions | Timeline transitions, up to 1 s |
| Photo to short clip | AI image animation | Video Router image_url, 3-10 s on gemini-omni-flash-1.1 |
| Hold a still | Yes | Yes, static, motion_ignored warning |
Animate a still, then place it
gemini-omni-flash-1.1 accepts image_url for image-to-video at 3-10 seconds, resolution 360p to 4K, 16:9 or 9:16, with native audio always on. The request below asks for a 5-second 9:16 clip at 720p. Poll the job, take the clip from result.artifacts[], and use it as a video[] slot with source_in 0.
import json, os, urllib.request
API = "https://api.sume.com/v1"
def post(path, body, key):
req = urllib.request.Request(
f"{API}{path}",
data=json.dumps(body).encode(),
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": key,
},
method="POST",
)
with urllib.request.urlopen(req) as res:
return json.load(res)
job = post(
"/video-router/generate",
{
"model": "gemini-omni-flash-1.1",
"prompt": "Slow push-in on the product, soft daylight, no text",
"image_url": os.environ["SUME_STILL_URL"],
"duration": 5,
"resolution": "720p",
"aspect_ratio": "9:16",
"mode": "async",
},
"animate-still-001",
)
print(job["request_id"])
Cost and choice
Billing for that model is provider list times 1.25 per output second by resolution, so read GET /v1/video-router/models for the current rate before you queue a batch. A generated clip is new footage, and the model may change details of the photo, so check the result. For a faithful push-in on a fixed image with no changes at all, keyframes in Edits are the more predictable tool; for many stills in a scripted run, the generate-then-place route is the one Sume documents.
Checking the plan for ignored motion
Because motion_ignored is a warning, not an error, a render can succeed with a still that never moves. The pre-flight POST /v1/timeline-1.0/plan cannot predict every warning (short-source pad or loop warnings are only known at render), so read warnings[] on the result and treat motion_ignored as a signal you asked for something Timeline does not do. If you saw it, you wanted an animated clip, not a hold.
One more fallback is to leave the push-in to an editor and use Sume for everything around it. That is a reasonable split: keyframes for a single hero shot, Timeline for the sequence, captions and soundtrack.
Three ways to get a moving shot
Pick the route by how much control you need over the movement itself.
- Keyframes in Edits: exact start and end positions on a photo you already have, no generation.
- Image-to-video through Sume: a generated clip from the still plus a prompt, with the model deciding how it moves.
- Start and end frames: for models that list
first_frameandlast_framein the catalog, pin both ends so the motion travels between two images you chose. Check the model's capabilities onGET /v1/videos/modelsbefore relying on it.
Doing the move before the render
If you want the effect of a keyframed zoom on a still, the practical route on Sume is to produce motion before Timeline sees it. Generate a short clip from the still with the Video Router image-to-video path, then place that clip in the video[] slots. The still itself stays a static hold, as the docs say, with a motion_ignored warning if you send motion fields. Whether the generated movement matches what you would have keyframed is a creative check you make by eye.
Cost stays small either way: the image-to-video clip is billed by the model you pick, and the Timeline render is $0.10 per ceiling minute, so a ten-second move inside a thirty-second Reel adds one generation and one render. Compare that with the price of reshooting, and with doing the keyframe in the Edits app, which Instagram offers for its own timeline (read 2026-10-03).
Sources
Related posts
More in Comparisons
- Eleven v4 90+ languages vs Sonic 3.6 44 languages: which fits yours
ElevenLabs says Eleven v4 covers 90+ languages. Cartesia says Sonic 3.6 covers 44. How to check your language before you pick a TTS model or API.
- Eleven v4 and Sume: a step-by-step feature map for voiceover work
Eleven v4 is not a model id on Sume. Here is what v4 does per ElevenLabs, and which Sume endpoint covers each step of a voiceover pipeline today.
- Eleven v4: 90+ languages or 99? ElevenLabs' own pages differ
ElevenLabs' launch post says more than 90 languages for Eleven v4; its docs page lists 99. How to plan around the gap, and what Sume lists instead.
- Eleven v4 drops the native accent; Sume keeps a language tag per voice
ElevenLabs' docs say v4 gives fluent target-language speech, not a preserved accent. Sume tags each voice with a language and checks for mismatch.
Written by Sume