Help-center article from a product video: one still per step
Combine video inspect sentence segments with video frames at[] to get a screenshot at each step start. 24 stills max per call, 300-second clip cap.

To turn a product walkthrough into a help-center article, run Video inspect with segmentation.mode: "sentence" for sentence-level timing, choose the sentences that start a step, then pass those start times to Video frames at[] for one still each. Frames are unbilled, and a call takes up to 24 times from a clip of at most 300 seconds.
Sume extracts the pictures and the timed text. It does not write the article, and it does not know which sentence is a step. You pick those, or write the rule.
Why two calls?
Inspect returns stills too, but at times it chooses: 8 mid-bin samples by default, or frames.at[] you supply, resized to a 768 long edge unless you set max_edge. Video frames is the exact-time route, returns source-size images by default and accepts max_edge from 16 to 2160. For documentation screenshots you want legible UI text, so use the frames route and omit the clamp.
The 300-second cap on frames is lower than the 1800 seconds inspect allows. A 12-minute tutorial can be transcribed whole but needs cutting into pieces of at most 300 seconds for frame extraction, which Video trim does.
What does the frames call look like?
Say the transcript shows steps starting at 4.2, 19.8, 41.0 and 63.5 seconds. Request PNG for crisp text. Idempotency-Key is sent on REST so a retry does not queue a second extract, and submit is always 202; this route does not honour mode: sync.
import json, os, time, urllib.request
key = os.environ.get("SUME_API_KEY")
if not key:
raise SystemExit("set SUME_API_KEY")
H = {"Authorization": f"Bearer {key}", "Content-Type": "application/json"}
def call(url, body=None, extra=None):
req = urllib.request.Request(url, headers={**H, **(extra or {})},
data=json.dumps(body).encode() if body else None)
return json.load(urllib.request.urlopen(req))
start = call("https://api.sume.com/v1/video-frames", {
"video_url": "https://media.sume.com/artifacts/artf_demo/setup.mp4",
"at": [4.2, 19.8, 41.0, 63.5], "format": "png"},
{"Idempotency-Key": "help-setup-frames-1"})
rid = start["request_id"]
while True:
r = call(f"https://api.sume.com/v1/video-frames/{rid}")
if r["video_frames"]["resource_status"] == "ready":
break
time.sleep(2)
for f in r["video_frames"]["frames"]:
print(f["t"], f["url"])The loop has no timeout or failed-status branch, so add both before production use.
What can go wrong?
| Symptom | Cause | Fix |
|---|---|---|
| frame_time_out_of_range | A time at or beyond the probed duration | Keep every t below source_duration_seconds |
| A frame has url null | That one instant failed to extract | Job still succeeds; retry that time alone |
| Rejected at admit | More than 24 times, or both at and fps | Batch into several calls |
| Source refused | Clip over 300 seconds or not on media.sume.com | Trim, or import first |
Is the result ready to publish?
Not without a review. Screenshots show whatever was on screen at that instant, including a half-loaded page or a notification banner. Look at each still, and nudge the time by a second if the UI has not settled. Frames are durable media.sume.com artifacts, so you can reference or download them. Read Jobs and results for the polling model.
Also keep the original recording out of the article if it contains real customer data. Extracting a still does not redact anything.
Sources
Related posts
More in Use cases
- Higgsfield's Seedance outage on Sept 30: slow vs failed jobs on Sume
Higgsfield said Seedance 2.0 failed more and 2.5 ran slow on Sept 30, now fixed. On Sume, a slow job is queued or processing; rerun only after failed.
- Check 200 holiday clips before upload with a free probe-only inspect
A video-inspect call with frames false is a probe with no stills and no charge. Use it as a pre-upload audio gate, then pull stills only for clips you doubt.
- 300 holiday UGC clips on one plan: pace around queue_full 429s
A Pro workspace holds 4 processing and 20 queued jobs. Submit 300 avatar clips, treat queued as normal, and retry queue_full with the same idempotency key.
- Hour-long podcast video to clips: the 1800-second source cap
Sume's trim, detach and inspect tools read sources up to 1800 seconds, so a 60-minute episode must be split before import. Limits and a safe split plan.
Written by Sume