Crop the browser bar off a screen recording with video filter
Crop browser chrome or a sidebar off a SaaS screen recording using Sume video-filter crop fractions. Free /check first, $0.02 per encode, 300-second cap.

To crop the browser bar off a screen recording, send the clip to Sume's video filter with one crop op: a rectangle given as fractions of the frame, not pixels. Dry-run it on the free /check endpoint, then encode for $0.02. You get a new MP4 and the original stays as it was.
This is the cleanup step in a SaaS onboarding video: a product tour recorded in a browser tab carries the tab strip, the address bar and a bookmarks row that viewers do not need and that can show things you did not mean to share.
How do I turn pixels into crop fractions?
A crop op takes x, y, width and height. x and y are in 0 to 1, width and height are in 0.05 to 1, and x + width and y + height must each stay at or below 1. Anything outside that is refused with video_filter_crop_out_of_bounds, so rounding up by a hair can fail a valid-looking crop. Round width and height down.
Measure the chrome in pixels on one frame of the recording, then divide. This script does it and, if SUME_API_KEY is set, posts the free check.
import json, math, os, urllib.request
def frac(v):
return math.floor(v * 10000) / 10000
def crop(w, h, top=0, bottom=0, left=0, right=0):
return {"op": "crop", "x": frac(left / w), "y": frac(top / h),
"width": frac((w - left - right) / w), "height": frac((h - top - bottom) / h)}
body = {"video_url": "https://media.sume.com/artifacts/artf_demo/tour.mp4",
"ops": [crop(1920, 1080, top=96)]}
print(json.dumps(body["ops"]))
key = os.environ.get("SUME_API_KEY")
if key:
req = urllib.request.Request("https://api.sume.com/v1/video-filter/check",
data=json.dumps(body).encode(),
headers={"Authorization": f"Bearer {key}", "Content-Type": "application/json"})
print(urllib.request.urlopen(req).read().decode())For a 1920 by 1080 recording with a 96-pixel bar on top, that prints y 0.0888 and height 0.9111. The compiler even-rounds the result for yuv420p, so the output is close to, not exactly, 1920 by 984.
What does the free check tell me?
POST /v1/video-filter/check runs the same schema, op allowlist and source preflight as the encode. It returns object: video_filter_check with valid, diagnostics[], the compiled filters, an estimate when valid, and a next_action of submit_video_filter or fix_program_and_recheck. It creates no job and reserves no credits, and it needs no idempotency key.
A passing check is not a guarantee. A program can still fail on the worker, and that comes back as a structured job error. So the check proves the request is well-formed, not that the render will succeed.
How do I encode and what are the limits?
Send the same body to POST /v1/video-filter with an Idempotency-Key. The default mode is async; mode: sync waits up to 30 seconds. Poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result, which gives a new video_url, duration_seconds and ops_applied. See Jobs and results.
| Item | Limit | Note |
|---|---|---|
| Source length | 300 seconds | Longer recordings fail with output_duration_exceeded |
| Ops per request | 8 | More fails with video_filter_too_many_ops |
| Source location | media.sume.com in your workspace | Import first with POST /v1/media-imports |
| Price | $0.02 per encode; check is free | Confirm in GET /v1/catalog |
A 10-minute walkthrough does not fit in one job. Cut it into ranges of at most 300 seconds with video trim (also $0.02), crop each piece with the same fractions, and join them on Timeline 1.0. One checklist step per clip, as many onboarding tours are already structured, avoids the problem entirely.
What can crop not do?
Crop is a rectangle. It removes everything outside it and keeps everything inside, so it cannot hide a customer email in the middle of a dashboard. If the sensitive part is inside the frame you keep, re-record with test data. Sume documents no dedicated redaction op, and a whole-frame blur is not a privacy tool.
Captions are a separate step. Burn them after the crop with video captions, $0.20 for clips up to 60 seconds. A silent screen recording needs authored cues with text, start and end, because speech-to-text fails a silent clip with caption_no_speech. Do the crop first so the caption's placement is computed against the frame viewers will actually see.
Sources
Related posts
More in Use cases
- Testimonial video from a written review: quote cards over B-roll
Turn a written customer review into a 20-second video: generate B-roll, burn the quote with caption cues, and credit the source honestly. Costs and limits.
- Deezer: 85% of AI-track streams fraudulent. Use Sume Music for beds
Deezer says up to 85% of streams on fully AI tracks were fraudulent in 2025 and demonetized. Why Sume Music fits video beds better than streaming releases.
- Deezer tags AI music: what to keep from a Sume Music job
Deezer tags fully AI-generated tracks and removes them from recommendations and editorial playlists. What a Sume Music job returns and what record to keep.
- Demand Gen dynamic product videos from Merchant Center: a clip per SKU
Google says videos uploaded to Merchant Center can run across Demand Gen based on live interest. Make one short product clip per SKU on Sume and upload each.
Written by Sume