Automate video editing in Python with an editing API
Automate video editing in Python by calling an editing API with Requests: submit a caption, cut, or crop job, poll until it ends, chain the output.

To automate video editing in Python, either drive a local library and FFmpeg on your own machine, or call a video editing API from a script. With an API, the script sends each edit (a caption, a cut, a crop) as a job, waits for it to finish, and feeds the output file into the next edit, with no FFmpeg install and no rendering on your hardware.
This post shows the API route with Sume and the Requests library. The Sume facts come from the Video captions, Video trim, Video filter, and Jobs and results docs and the Sume API reference, read on 2026-09-29. Anything described as current behavior is read from Sume's code.
Should I edit locally or call an API from Python?
Edit locally when your files live on your machine, you need full FFmpeg control, and you have the CPU to spare. Call an API when you'd rather not install or scale an encoder, when the edits are standard (cut, crop, caption, overlay, join), and when the source videos are already online. The trade-off with Sume: its cut, crop, and join tools read only files already on Sume's media host, so a pipeline starts from a public URL that the caption job can take, or from an earlier Sume output.
What does an editing script look like?
One helper submits a job and polls it; each edit is one call. This one captions a public clip, then cuts the captioned file to its first 15 seconds; the caption job's output is a media.sume.com file, which in current code the trim tool admits. Keep the key in a server-side environment variable, never in browser or mobile code.
import os, time, requests
API = "https://api.sume.com/v1"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def call(method, url, key=None, **kw):
headers = {**AUTH, **({"Idempotency-Key": key} if key else {})}
r = requests.request(method, url, headers=headers, timeout=30, **kw)
r.raise_for_status()
return r.json()["data"]
def run(path, body, key):
job = call("POST", f"{API}/{path}", key, json=body)
status = call("GET", job["status_url"])
while not status["terminal"]:
time.sleep(status["next_poll_after_seconds"] or 5)
status = call("GET", job["status_url"])
if status["sume_status"] != "completed":
raise RuntimeError(f"{path}: {status['sume_status']}")
return job
cap = run("video-captions", {"video_url": "https://example.com/clip.mp4"}, "clip-cap-1")
captioned = call("GET", f"{API}/video-captions/{cap['video_caption_id']}")["video_url"]
cut = run("video-trim", {"video_url": captioned, "start": 0, "duration": 15}, "clip-cut-1")
print(call("GET", cut["result_url"])["result"]["video_url"])Why poll the status before reading the result?
Because only a completed job has a result. terminal is true once the job is completed, failed, or canceled, and sume_status says which. In current code GET /v1/jobs/{id}/result answers job_not_completed, marked retryable, even for a failed or canceled job, so a script that retries the result call alone can loop forever. Store each job id as soon as the submit returns, so a restarted script resumes polling instead of paying for the edit again.
Send an Idempotency-Key on every submit and reuse it, with the same body, only when retrying after a timeout or network failure. To skip polling on long pipelines, pass a webhook_url and receive a signed callback when each job ends.
Which edits can the script call, and what do they cost?
Each call is a separate job, billed on its own, plus a 5.5% agent fee by default. The media tools reserve their listed rate and capture their own compute, never above the reservation.
- In current code a caption job refuses a source over 60 seconds or one with no audio stream.
- Trim and filter return a new MP4 and leave the source untouched; filter programs can be checked free at
POST /v1/video-filter/checkfirst. - There is no public route to upload a file from your computer, so local footage has to be reachable at a public HTTPS URL for the caption step.
| Edit | Endpoint | Input it takes | Price |
|---|---|---|---|
| Burn captions | POST /v1/video-captions | Public HTTPS URL, ≤ 60 s, with sound | $0.20 per job |
| Cut a range, resize | POST /v1/video-trim | Sume-hosted file | up to $0.02 per job |
| Crop, dim, other allowlisted filters | POST /v1/video-filter | Sume-hosted file, ≤ 300 s | up to $0.02 per job |
| Logo over video | POST /v1/timeline-1.0/compose | Sume-hosted files | up to $0.02 per job |
| Join clips over audio | POST /v1/timeline-1.0/render | Sume-hosted files | up to $0.10 per output minute |
Sources
Related posts
More in Developers
- Bash for loop with curl: one API request per line
Loop over a file with while IFS= read -r, build each JSON body with jq --arg, send it with curl --fail-with-body, and pace it under the API's rate limit.
- C# HttpClient default timeout: 100 seconds, and how to set it
HttpClient.Timeout defaults to 100 seconds per request and throws TaskCanceledException. How to set it, and why slow API jobs need polling instead.
- How to delete my data from an AI tool, and what stays
Delete your data from an AI tool in two steps: delete the account, then send a deletion request for stored files. How it works on Sume, and its limits.
- Do AI companies sell your data? What to read in the policy
Some may; the privacy policy is where to check. How to read its sale and sharing sections, and what Sume's policy says it collects and shares.
Written by Sume