Mixed-model clips judder in Timeline: conform each to 30 fps first
Timeline resamples sources that run at different frame rates and warns output_fps_resamples_sources. Conform each clip with video-trim output, $0.02 a clip.

If an episode is built from clips that do not share a frame rate, Timeline will still render it, but it resamples the odd ones by repeating or dropping frames, and you see that as judder on motion. The fix on Sume is to conform each clip first with POST /v1/video-trim and its output object, which sets width, height and fps for $0.02 a clip (read 2026-10-03), then hand the conformed files to Timeline.
The topic matters more as Reel guidance moves toward longer cuts. Metricool's Instagram page reports that Instagram's updated Reels guide gives ideal clip lengths from 3 seconds to 3 minutes (reported, read 2026-10-03). A three-minute episode is a lot of slots, and the more sources you join, the more likely one of them runs at a different rate than the rest.
What Timeline does with mixed rates
Timeline's default output is a 1080 by 1920 MP4. When you omit output.fps, it renders at the rate the sources already run at, and the longest video sources decide. Stills have no rate, and the fallback is 30 only when nothing has one. A source at a different rate is met by repeating or dropping a frame every few frames, and the response reports output_fps_resamples_sources with the rate the sources wanted.
That warning is the signal to act on. It is a warning, not a failure, so a job that returns it still completes and still bills. The risk is that nobody reads it and the judder ships. Treat output_fps_resamples_sources as a gate in your pipeline: if it appears in a render, conform and render again.
| Field | Allowed values | Note |
|---|---|---|
| output.width | 256 to 2160 | Use the same value for every clip |
| output.height | 256 to 2160 | 1920 for a 1080 by 1920 vertical episode |
| output.fps | 24, 25, 30 or 60 | Pick the rate Timeline will render at |
| precision | exact (default) | Required for output; keyframe refuses it |
| audio | keep (default) or drop | Exact keeps audio as AAC |
Conform one clip, then all of them
A trim needs a start and exactly one of end or duration. To conform a whole clip, use start: 0 and a duration that covers it. If the duration runs past the source, the end clamps and the result carries trim_clamped_to_source, so a generous duration is safe. The longest cut is 900 seconds and the source may be up to 1,800 seconds.
Because exact mode is a frame-accurate re-encode, this is also your last chance to pick the size and rate once for the whole episode. Choose 1080 by 1920 at 30 fps, apply it to every clip, and then set output.fps to 30 in the Timeline program as well, so a stray source cannot change the rate.
import os, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
clips = os.environ["EPISODE_CLIPS"].split(",") # media.sume.com URLs
ids = []
for i, url in enumerate(clips, 1):
r = requests.post(f"{API}/v1/video-trim",
headers={**H, "Idempotency-Key": f"ep4-conform-{i}"},
json={"video_url": url, "start": 0, "duration": 900,
"output": {"width": 1080, "height": 1920, "fps": 30}})
r.raise_for_status()
ids.append(r.json()["request_id"])
print(ids)
# For each id: GET /v1/jobs/<id>/status until completed, then
# GET /v1/jobs/<id>/result and use result.video_url as the Timeline source_url.What to do and what it costs
Conform only the clips that need it. A cheap first step is to render the Timeline once and read the warnings: if output_fps_resamples_sources is absent, you do not need to conform anything. If it appears, the response names the rate the sources wanted, which tells you which rate to standardise on.
Price it before the season. A 24-clip episode conformed in full is 24 jobs at $0.02, which is $0.48, and the Timeline render of a three-minute output is $0.30 at $0.10 per output minute. Conforming costs more than the final render, so it is worth skipping when the warning is absent, and worth doing once when it is present.
Use one idempotency key per clip per version, as in the sketch. A rerun of the script returns the existing jobs instead of billing again, and a deliberate change to the target size is a new key such as ep4-conform-v2-1.
Do not use keyframe mode for this. It stream-copies, so it cannot change geometry or rate, and a request that asks for output with it is refused with video_trim_output_requires_exact.
What Sume does and does not do
Sume re-encodes a hosted clip to the size and rate you name and tells you in Timeline when sources disagree. It does not choose the target for you, it does not interpolate new frames to make a lower-rate clip look smoother, and it does not take a clip from outside media.sume.com; import it first. Whether a given model's clip runs at 24 or 30 is something to confirm from the file you get back, not an assumption.
Sources
Related posts
More in Use cases
- Mixed real and AI footage: what Gemini's SynthID check reports
A video that mixes camera footage and Gemini Omni clips shows SynthID in some segments only. What Google says it reports, and a test edit on Sume.
- Mobile game ad cutdowns: one master to 6, 15, 30 and 60 seconds
Mintegral recommends 6/15/30/60 s, Unity and Axon cap at 60 s, Liftoff at 180 s. Cut one portrait master to each length with Sume's trim API for $0.02 a job.
- Move an object in a photo with AI: erase, then place with two masks
No Sume image model has a move operation. Move an object by erasing it with one mask and placing it with a second, using GPT Image 2.5 and Python.
- Movember 2026 team update videos from phone photos, one per week
Movember asks people to talk about men's health. Make four weekly update clips for a fundraising team page from phone photos, with caption cues you wrote.
Written by Sume