Mux Robots find-scenes vs Sume: what scene detection covers

Mux promoted find-scenes out of experimental on Sept 17, 2026. Sume has no scene-detection job on dest; here is the probe, stills and trim route instead.

5 min readSume
All posts

Mux's find-scenes workflow splits a video into ordered, timestamped scenes with structured metadata, and Mux promoted it out of experimental status on September 17, 2026. Sume does not ship an equivalent scene-detection job on dest: video inspect returns probe facts, sampled stills and an optional transcript, and you decide where the scenes break.

This page compares the two honestly, using the Mux changelog, the Mux Robots guide and the Mux pricing page on the Mux side, and the Sume docs on the other.

What did Mux change for find-scenes on September 17, 2026?

The Mux changelog lists five Robots workflows that left experimental status that day: translate-audio, find-scenes, find-best-thumbnails, generate-premium-captions and generate-engagement-insights. Mux's July 2026 changelog entry says they run through the Robots Jobs API, Directives or the dashboard.

Mux's guide puts find-scenes under a Structure category, next to timestamped chapter generation. A job is a POST to https://api.mux.com/robots/v0/jobs/{workflow} with an asset id, so the video has to be a Mux Video asset first. The pricing page says the find-scenes workflow relies on the Shots primitive, billed at $0.001 per input minute on top of Robots units.

Does Sume have a scene-detection endpoint?

No, not on dest. The older video analyses resource did return typed scenes[], but its docs say dest answers 410 video_analysis_retired and that production keeps it only until a planned removal. The docs tell new work not to start there. Dest-only video_analyze and video_segment tools exist only when they appear in tools_list, so do not build a product on them.

What you can rely on is video inspect: POST /v1/video-inspect takes one media.sume.com clip up to 1800 seconds and returns a probe, stills and, if you ask, a transcript. Probe and stills are unbilled. It does not return scene boundaries, shot types or summaries.

How do you approximate scenes with Sume today?

The practical route is to sample the clip, look at the stills yourself (or with your own vision model), and cut where the content changes. Frames can be sampled by interval or at explicit timestamps, and seek: "fast" snaps each still to the keyframe at or before its instant, which is quicker for skimming but can be earlier by up to one GOP.

Once you have boundaries, video trim cuts a [start, end) range into a new MP4 for $0.02 per job, and Timeline 1.0 reassembles clips against an audio spine. That is a manual pipeline, not a hosted workflow.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: scene-skim-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": { "fps": 0.5, "seek": "fast", "max_edge": 512 }
  }'

Side by side

The table lists only what each vendor's own documentation states.

Mux find-scenes vs Sume clip inspection, read 2026-10-02
QuestionMux Robots find-scenesSume video inspect
OutputOrdered, timestamped scenes with structured metadataProbe, stills, optional transcript
InputA Mux Video asset idOne workspace media.sume.com clip, up to 1800 s
BillingRobots units plus $0.001 per input minute for ShotsProbe and stills unbilled; transcript $0.01 per audio minute
Job statespending, processing, completed, errored, cancelledqueued, processing, completed, failed, canceled
Result retentionJobs deleted after 30 daysNot stated in the Sume docs

Which one should you pick?

Pick Mux Robots if your videos already live in Mux and you want hosted scene and chapter structure. Pick Sume if the job is to prepare, trim and assemble clips, and you are comfortable owning the cut decisions. If you only need chapter titles from speech, a transcript from transcribe: true plus your own model is the closer fit; see YouTube chapters from a transcript.

Sume does not promise scene accuracy, because it does not detect scenes. Mux's own docs do not publish an accuracy figure for find-scenes either, so test both on your footage before committing.

What should you verify before choosing?

Run both on three clips from your real library: one talking head, one fast-cut ad and one long screen recording. Mux's playground and Robots jobs let you see its structured scenes. For Sume, count how many stills you need per minute to see every cut; the frames program allows up to 24 stills per call and 0.5 fps sampling is a reasonable starting point, but Sume does not recommend a rate.

Also check the data path. Mux jobs are stored for 30 days per its guide, so copy results out. On Sume, probe and stills come back as durable media.sume.com artifacts, but the Sume docs I read do not state a retention period, so confirm with your plan before depending on old URLs.

  • Mux job statuses: pending, processing, completed, errored, cancelled.
  • Sume job statuses: queued, processing, completed, failed, canceled.
  • Sume stills are capped at 24 per call; a longer film needs several calls.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume