E-bike Black Friday video ad: start and end frame reveal, three trims
An e-bike Black Friday ad from two stills: a 10-second start/end-frame clip on Gemini Omni Flash, then three trimmed cutdowns, about $1.51 in Sume jobs.

For an e-bike Black Friday video ad, make two stills (the bike parked in a garage, then the same bike at a trailhead), send them to Gemini Omni Flash 1.1 as image_url and end_image_url, and let the model write the ride between them. One 10-second 720p master costs $1.25, the two stills $0.20 on Nano Banana 2, and three video-trim cutdowns $0.06. Total $1.51 for a master and three feed lengths.
Start and end frames suit bikes because the product has to stay the same object from the first frame to the last. A text-only clip can drift the frame, the wheel size or the colour, while two pinned frames give the model less room to invent.
What the jobs cost
Sume bills generation at the provider list times 1.25: Omni 720p is $0.10 a second on the list, so $0.125 here, and nano-banana-2 is $0.08 a still on the list, so $0.10. Trim is a flat $0.02 a job. Confirm live rates in GET /v1/catalog.
| Item | Basis | Sume price |
|---|---|---|
| Start still and end still | 2 x $0.10 | $0.20 |
| Master clip | 10 s x $0.125 at 720p | $1.25 |
| Cutdowns at three lengths | 3 trim jobs x $0.02 | $0.06 |
| Total | $1.51 |
Make the two frames match
Omni's image-to-video accepts image_url and an optional end_image_url, in the same 3 to 10 second envelope as text-to-video. Describe the motion between the frames in the prompt, and keep the camera simple: a slow tracking move is easier to hold than a whip pan.
Image generation is where the two frames get their sameness. Make the second still as an edit of the first, with the first as an input_references entry, instead of prompting two fresh bikes. The Image API docs describe references as public HTTPS URLs and say a model whose reference range is 0 to 0 is text-to-image only, so read the catalog before you pin a model.
POST /v1/images blocks for up to 30 seconds and returns 200 with the image, or 202 with a job envelope when the render runs long. Branch on the status code, not on the body shape, and poll GET /v1/jobs/{id}/status for the 202 case. Check both stills by eye before the Omni job starts: a generation you did not like still bills if it completed, while a failed or cancelled one does not.
Cut it to length
video_url for trim must be this workspace's media.sume.com clip, so use the artifact URL that the Omni job returned. You send start and exactly one of end or duration. The default precision is exact, a frame-accurate re-encode. keyframe is a stream copy that can start a GOP early, and then you re-base your timestamps against actual_start_seconds in the result.
For a paid ad, keep exact. A keyframe cut starts a GOP early, so the cutdown can open on slightly earlier footage than you chose, which is easy to miss and costly in the first second of a feed ad. Trim adds no provider inference, only worker ffmpeg, which is why the price is a flat $0.02 whatever the length.
endpast the end of the source clamps and returns the warningtrim_clamped_to_source.- Sending both
endanddurationis a 400,video_trim_range_conflict. - Output must be at least 0.2 s and at most 900 s, and the source at most 1800 s.
One trim call
import os
import uuid
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": f"ebike-trim-{uuid.uuid4()}"}
body = {
"video_url": "https://media.sume.com/artifacts/artf_demo/ebike-master.mp4",
"start": 0,
"duration": 6,
}
r = requests.post("https://api.sume.com/v1/video-trim", headers=H, json=body, timeout=60)
print(r.status_code, r.json().get("next_action"))
Idempotency and where to post
Do not reuse an Idempotency-Key across the three cutdowns, because the key is bound to one operation and payload. Generate one per call and store it with the job id.
Before you post, read the cutdown's duration_seconds from GET /v1/jobs/:id/result. For vertical placements, 9:16 is the Omni aspect that matches Shorts and Reels, and YouTube's help page says Shorts can be up to 3 minutes long (YouTube Help, read 2026-10-05), so a 10-second master is far inside every limit. The longer-form route is a Timeline job that joins several masters; see the cutdown guide for sizing.
Sources
Related posts
More in Use cases
- Eleven v4 says 75% prefer it: run your own 20-listener test
ElevenLabs reports ~75% preference for Eleven v4 in blind tests. A pre-named voice winning 15 of 20 is ~2% by chance. Run your own test on Sume.
- Eleven v4 claims 90+ languages: check Korean before you ship
A 90+ language claim is not a pronunciation guarantee. How to QA Korean voiceover on Eleven v4, and what Sume's voice-language guard and receipts prove.
- 10 seconds to clone a voice: ElevenLabs IVC vs cloning in Sume
ElevenLabs Instant Voice Clones start from 10 seconds of audio. Sume voice cloning is free of charge and sets a language per voice; here is how they differ.
- Email header image 600x200: render 1920x640 at exactly 3:1
A 600x200 email header is 3:1, the widest shape GPT Image 2.5 accepts. Render 1920x640, then export 1200x400 for retina and 600x200 for 1x with Pillow.
Written by Sume