video_filter_crop_out_of_bounds: FFmpeg crop pixels to fractions

Sume video-filter crop takes fractions of the frame, not pixels. Convert FFmpeg crop=w:h:x:y with a runnable Python check of the bounds.

4 min readSume
All posts

Sume's video-filter crop op takes x, y, width and height as fractions of the source frame, not pixels. video_filter_crop_out_of_bounds comes back when a value is outside its range, a side is under 0.05, or x + width or y + height exceeds 1. Divide your FFmpeg crop=w:h:x:y values by the source width and height to convert them.

What are the exact bounds?

From the Video filter page: x and y in 0 to 1, width and height in 0.05 to 1, x + width <= 1, y + height <= 1. The compiler even-rounds the result for yuv420p, so you do not need to worry about odd pixel counts. A program holds at most 8 ops, and a crop can be combined with a dim op in the same ops[].

The FFmpeg filters documentation defines crop as w/out_w, h/out_h, x, y, where x and y are the left and top edge of the output inside the input and default to centred. Those four numbers map one to one onto Sume's four fractions.

FFmpeg crop parameters and Sume crop op fields, read 2026-10-03 from the FFmpeg filters documentation and Video filter page
FFmpeg cropSume crop opAllowed range
x (left edge, pixels)x = pixels / source width0 to 1
y (top edge, pixels)y = pixels / source height0 to 1
w (width, pixels)width = pixels / source width0.05 to 1
h (height, pixels)height = pixels / source height0.05 to 1

What does the conversion look like in code?

For a 1920x1080 clip, a centred 9:16 strip is crop=608:1080:656:0 in FFmpeg terms: 608 is roughly 9/16 of 1080, and 656 centres it. As fractions that is x 0.3417 and width 0.3167. The second call below pastes a portrait pixel box onto a landscape clip, and the check flags it, which is the situation that produces the refusal.

Run it as written; it needs only the standard library.

def crop_fractions(src_w, src_h, w, h, x, y):
    op = {"op": "crop", "x": round(x / src_w, 4), "y": round(y / src_h, 4),
          "width": round(w / src_w, 4), "height": round(h / src_h, 4)}
    ok = (0 <= op["x"] <= 1 and 0 <= op["y"] <= 1
          and 0.05 <= op["width"] <= 1 and 0.05 <= op["height"] <= 1
          and op["x"] + op["width"] <= 1 + 1e-9
          and op["y"] + op["height"] <= 1 + 1e-9)
    return op, ok

# ffmpeg -vf crop=608:1080:656:0 on a 1920x1080 clip
print(crop_fractions(1920, 1080, 608, 1080, 656, 0))
# a 1080x1920 pixel box pasted onto a 1920x1080 clip is out of bounds
print(crop_fractions(1920, 1080, 1080, 1920, 0, 0))

Check first, then encode

POST /v1/video-filter/check validates the same program for free and returns valid, diagnostics and an estimate, with next_action set to submit_video_filter or fix_program_and_recheck. Only then send the encode, which is $0.02 per job according to the Video filter page and needs an Idempotency-Key. The source must be a hosted clip of at most 300 seconds.

If you wanted a dim as well, amount is in (0, 1]: 0 and anything above 1 are refused with video_filter_amount_out_of_range.

Why fractions instead of pixels?

A fraction survives a change of source size. The same 0.3167-wide strip is a 9:16 slice of a 1080p clip and of a 4K clip, and the compiler resolves it against the real frame at encode time, rounding to even sizes for yuv420p. Pixels would have to be re-derived whenever the upstream render changes resolution.

The cost is that you must know the source size when you convert. Read it from video-inspect, whose probe returns width and height, then compute the four fractions once and reuse them for every clip of that size.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume