video_inspect 400: frames program and transcribe flag errors

Fix video_inspect_frames_program_conflict, _required and video_inspect_transcribe_required: send at or fps, and transcribe true with its language options.

5 min readSume
All posts

A video_inspect call returns a 400 for four mistakes about its options: video_inspect_frames_program_conflict (both at and fps), video_inspect_frames_program_required (a frames object with neither), video_inspect_transcribe_required (transcript options without transcribe: true) and ffmpeg_fields_rejected (any ffmpeg-style key). Two more codes arrive after the probe: frame_time_out_of_range and inspect_source_has_no_audio.

Everything below comes from the Video inspect reference, read 2026-10-10. The route is POST /v1/video-inspect, it takes a media.sume.com clip of your workspace (import it first), and a write needs an Idempotency-Key.

The refusal table

These are the stable codes the docs list for this route, with the one-line fix for each.

Video inspect refusals and fixes, from the Video inspect docs (read 2026-10-10)
CodeWhat triggers itFix
video_inspect_frames_program_conflictframes.at[] and frames.fps togetherSend only one of the two
video_inspect_frames_program_requiredframes is an object with no at[] and no fpsAdd one, or send frames: false or omit frames
video_inspect_transcribe_requiredlanguage_code, segmentation or duration_seconds without transcribe: trueAdd transcribe: true, or drop those fields
inspect_source_has_no_audiotranscribe: true on a clip with no audio trackCheck probe.has_audio first
frame_time_out_of_rangeAn at value outside [0, duration)Use times below the clip duration, which the error reports
ffmpeg_fields_rejectedvf, filter, ffmpeg, cmd, codec, crf or similarRemove them; the server compiles ffmpeg
source_not_foundThe media.sume.com URL is dead or from another workspaceRe-import the clip

Choosing a frames program

The frames field has four shapes, and the confusion usually comes from treating them as one. Leaving it out gives 8 mid-bin stills (1 fps if the clip is under 8 seconds). frames: false gives the probe only. { at: [...] } gives explicit timestamps, 1 to 24 values, each 0 or more. { fps: n } samples at a rate above 0 and up to 2, mid-bin, capped at 24 stills.

The cap matters for long clips. Twenty-four stills at 2 fps cover only 12 seconds, so a two-minute clip needs an at list or a lower fps. At 0.2 fps, 24 stills reach 120 seconds. The object also takes format (jpeg by default, or png) and max_edge from 64 to 2160, with a default of 768.

  • Add seek: "fast" for a quick look: each still moves to the keyframe at or before its instant, up to about one GOP early (roughly 0 to 5 seconds on typical sources), never later.
  • Keep seek on its default precise when the timestamp must be accurate, such as explicit at instants.
  • The source can be up to 1800 seconds, and each call returns at most 24 stills.
  • For an accurate frame at one time at source size, use Video frames instead; inspect stills are clamped to 768 pixels by default.

The transcript flag has three dependents

transcribe: true runs Sume STT 1.0 on the clip audio. Three other fields only make sense with it: language_code (a hint such as en, otherwise auto-detect), segmentation (mode: "sentence" returns gapless sentence segments[], with silence_split_seconds from 0.2 to 3) and duration_seconds (the reservation hint, up to 600). Send any of them alone and you get video_inspect_transcribe_required.

The public speech rate is $0.01 per audio minute. Without duration_seconds, Sume reserves 1 minute, so an 8-minute clip with no hint is reserved for 1 minute; send the length so the hold matches the work. Probe and stills are billed separately by Modal compute, and the live numbers are in GET /v1/catalog.

A request that passes

The sample asks for three explicit stills, a 1080-pixel long edge, and an English transcript. It sets at without fps, and sets transcribe alongside language_code. The default mode is sync: the handler waits up to 30 seconds and returns 200 with the result, or 202 with a queued job that you then poll with jobs_wait and jobs_result over MCP, or GET /v1/video-inspect/:id over REST.

import asyncio, os
import httpx

async def main():
    key = os.environ["SUME_API_KEY"]
    body = {
        "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
        "frames": {"at": [0, 2.5, 5], "max_edge": 1080},
        "transcribe": True,
        "language_code": "en",
    }
    async with httpx.AsyncClient(base_url="https://api.sume.com", timeout=60) as c:
        r = await c.post("/v1/video-inspect", json=body,
            headers={"Authorization": f"Bearer {key}", "Idempotency-Key": "inspect-001"})
        print(r.status_code, r.json())

asyncio.run(main())

Avoid the silent-clip error

If you are not sure a clip has sound, run a probe-only inspect first with frames: false, read probe.has_audio, and only then ask for a transcript. For a clip that is silent, captions can still be burned from authored cues; see Video captions.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume