Reference ingest uncertain[]: five kinds and what to do
The reference-ingest manifest lists five uncertain kinds, each with a suggested next step. Read the list, look again once per entry, and skip the rest.

The uncertain[] array in a Sume reference-ingest manifest is the only list of reasons to look at a frame again. Each entry has a kind, an optional ref and at time, and a suggested string. The contract defines five kinds: text_low_confidence, gradual_transition_candidate, ocr_language_unsupported, ocr_unavailable and audio_ambiguous.
Reference ingest is dest first (docs say development is auto-on and production is opt-in), so check tools_list for reference_ingest before you build on it. It reads one clip of up to 300 seconds that your workspace already holds on media.sume.com. That fits series episodes, which YouTube says are rolling out across web, mobile and TV (YouTube Blog, read 2026-10-04).
The five kinds
The reference ingest docs define the list as the only reason to re-inspect and do not describe each kind at length. The right column below is an inference from the kind names plus the documented rule; the manifest's own suggested string is the authority for each entry.
| kind | Reads as | First move |
|---|---|---|
| text_low_confidence | An OCR line under the confidence threshold | Open the attached native crop and read it yourself |
| gradual_transition_candidate | A cut that may be a fade or dissolve | Pull a frame at the entry's at time |
| ocr_language_unsupported | Text in a script OCR does not read | Treat on-screen text as unread |
| ocr_unavailable | The OCR pass did not run | Treat on-screen text as unread; re-run later |
| audio_ambiguous | The audio facts are not clear-cut | Listen to that stretch before planning music |
Look once per entry
The documented flow is: read the attached crop, then call video_frames_create at a manifest time, at most once per entry. video_frames takes 1 to 24 times per call, so one request can cover every entry that carries an at. Lines under the confidence threshold (default 0.85, set by ocr.min_confidence_attach_crop) come back as needs_verification with a crop, which is how most text_low_confidence entries are settled without a second call.
Collect the times
The sample manifest below is shortened to the fields this step reads. Replace it with the manifest your read returned.
uncertain = [
{"kind": "text_low_confidence", "ref": "t3", "at": 2.4,
"suggested": "read the crop"},
{"kind": "gradual_transition_candidate", "at": 6.1,
"suggested": "look at a frame near 6.1 s"},
{"kind": "audio_ambiguous", "suggested": "listen to the track"},
]
times = sorted({
round(item["at"], 3)
for item in uncertain
if "at" in item and item["kind"] != "audio_ambiguous"
})
others = [i["kind"] for i in uncertain if "at" not in i]
print("video_frames at:", times[:24])
print("no frame to pull:", others)
Frames come from video frames, which is billed by its Modal compute and reserves its ceiling at submit. If the clip is over 300 seconds, ingest answers source_too_long_for_reference_ingest; use video_inspect for that clip.
Sources
Related posts
More in Media tools
- Remove a video background: Firefly vs Sume options
Adobe Firefly can now remove a video background across a whole clip. The Sume docs list a cutout tool for stills, not video, so plan a clean backdrop instead.
- Roku ad frame rates: only three allowed, probe first
Roku ads reportedly accept 23.98, 25 or 29.97 fps. Probe your clip before upload with Sume's video inspect, and cut it to length if it is outside 6 to 92 s.
- Seedance 2.5 secondary edit vs trim-and-regenerate on Sume
BytePlus describes timestamp-level edits to Seedance 2.5 clips. Sume has no such edit field, so here is the trim-and-regenerate route and its limits.
- SeedVR2 10x video upscale vs Sume's upscale tool
SeedVR2 is listed with 1x to 10x video upscaling. Sume exposes a video_upscale_create tool, but its docs name no scale factors, so check the live schema first.
Written by Sume