Mux Robots find-scenes vs Sume: what scene detection covers
Mux promoted find-scenes out of experimental on Sept 17, 2026. Sume has no scene-detection job on dest; here is the probe, stills and trim route instead.

Mux's find-scenes workflow splits a video into ordered, timestamped scenes with structured metadata, and Mux promoted it out of experimental status on September 17, 2026. Sume does not ship an equivalent scene-detection job on dest: video inspect returns probe facts, sampled stills and an optional transcript, and you decide where the scenes break.
This page compares the two honestly, using the Mux changelog, the Mux Robots guide and the Mux pricing page on the Mux side, and the Sume docs on the other.
What did Mux change for find-scenes on September 17, 2026?
The Mux changelog lists five Robots workflows that left experimental status that day: translate-audio, find-scenes, find-best-thumbnails, generate-premium-captions and generate-engagement-insights. Mux's July 2026 changelog entry says they run through the Robots Jobs API, Directives or the dashboard.
Mux's guide puts find-scenes under a Structure category, next to timestamped chapter generation. A job is a POST to https://api.mux.com/robots/v0/jobs/{workflow} with an asset id, so the video has to be a Mux Video asset first. The pricing page says the find-scenes workflow relies on the Shots primitive, billed at $0.001 per input minute on top of Robots units.
Does Sume have a scene-detection endpoint?
No, not on dest. The older video analyses resource did return typed scenes[], but its docs say dest answers 410 video_analysis_retired and that production keeps it only until a planned removal. The docs tell new work not to start there. Dest-only video_analyze and video_segment tools exist only when they appear in tools_list, so do not build a product on them.
What you can rely on is video inspect: POST /v1/video-inspect takes one media.sume.com clip up to 1800 seconds and returns a probe, stills and, if you ask, a transcript. Probe and stills are unbilled. It does not return scene boundaries, shot types or summaries.
How do you approximate scenes with Sume today?
The practical route is to sample the clip, look at the stills yourself (or with your own vision model), and cut where the content changes. Frames can be sampled by interval or at explicit timestamps, and seek: "fast" snaps each still to the keyframe at or before its instant, which is quicker for skimming but can be earlier by up to one GOP.
Once you have boundaries, video trim cuts a [start, end) range into a new MP4 for $0.02 per job, and Timeline 1.0 reassembles clips against an audio spine. That is a manual pipeline, not a hosted workflow.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: scene-skim-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": { "fps": 0.5, "seek": "fast", "max_edge": 512 }
}'Side by side
The table lists only what each vendor's own documentation states.
| Question | Mux Robots find-scenes | Sume video inspect |
|---|---|---|
| Output | Ordered, timestamped scenes with structured metadata | Probe, stills, optional transcript |
| Input | A Mux Video asset id | One workspace media.sume.com clip, up to 1800 s |
| Billing | Robots units plus $0.001 per input minute for Shots | Probe and stills unbilled; transcript $0.01 per audio minute |
| Job states | pending, processing, completed, errored, cancelled | queued, processing, completed, failed, canceled |
| Result retention | Jobs deleted after 30 days | Not stated in the Sume docs |
Which one should you pick?
Pick Mux Robots if your videos already live in Mux and you want hosted scene and chapter structure. Pick Sume if the job is to prepare, trim and assemble clips, and you are comfortable owning the cut decisions. If you only need chapter titles from speech, a transcript from transcribe: true plus your own model is the closer fit; see YouTube chapters from a transcript.
Sume does not promise scene accuracy, because it does not detect scenes. Mux's own docs do not publish an accuracy figure for find-scenes either, so test both on your footage before committing.
What should you verify before choosing?
Run both on three clips from your real library: one talking head, one fast-cut ad and one long screen recording. Mux's playground and Robots jobs let you see its structured scenes. For Sume, count how many stills you need per minute to see every cut; the frames program allows up to 24 stills per call and 0.5 fps sampling is a reasonable starting point, but Sume does not recommend a rate.
Also check the data path. Mux jobs are stored for 30 days per its guide, so copy results out. On Sume, probe and stills come back as durable media.sume.com artifacts, but the Sume docs I read do not state a retention period, so confirm with your plan before depending on old URLs.
- Mux job statuses: pending, processing, completed, errored, cancelled.
- Sume job statuses: queued, processing, completed, failed, canceled.
- Sume stills are capped at 24 per call; a longer film needs several calls.
Sources
Related posts
More in Comparisons
- Mux Robots playground vs Sume: test a video job before paying
Mux's Robots playground (Oct 1, 2026) bills real units. Sume's free dry runs are timeline plan, video inspect probe and stills, and a filter check.
- Mux thumbnail time and fit_mode vs Sume video frames at and max_edge
Mux gets a thumbnail from a playback ID URL with time, width and fit_mode. Sume video-frames takes at[] times or fps and returns durable image files.
- Nano Banana 2: five characters, 14 objects vs Sume reference slots
Google says Nano Banana 2 keeps up to five characters and 14 objects consistent. Sume caps reference images per model; see what you can send and where it stops.
- Novita AI API alternative for video and image jobs: Sume
Novita offers model APIs, agent sandboxes and GPU deployment. Sume offers managed media jobs only. Where they overlap and where Novita does more.
Written by Sume