Which agent LLM can read video frames? Image input by model

GPT-6.1 Sol, Claude 5.5, Grok 4.7 and DeepSeek Flash accept images; one DeepSeek row is text only. How a Sume agent checks a render with stills.

5 min readSume
All posts

Yes, for stills. GPT-6.1 Sol, every current Claude model, Grok 4.7 and DeepSeek's deepseek-flash all take image input on the vendors' pages, so an agent running any of them can look at frames pulled from a render. DeepSeek's deepseek-v4-pro does not accept images, and Sume keeps one text-only DeepSeek row, V4 Flash 0731, for stored picks.

None of these models is handed an MP4 by Sume. The agent calls a frame tool, gets image artifacts back, and reads those. That distinction decides how you build a quality check.

What does each vendor say about image input?

Input is only half of it. Image tokens are billed as input, so a check that sends 24 full-size frames on every turn is not free. Each vendor prices images in its own way, and I left those rules out because they were not on the pages above.

Image input per vendor page, read 2026-10-02
ModelImage inputWhat the page says
GPT-6.1 SolYesImage input is in its supported-features list
Claude Opus 5.5 and Sonnet 5.5YesAll current models take text and image input
Grok 4.7YesModalities listed as text, image to text
DeepSeek deepseek-flashYesVision is marked supported
DeepSeek deepseek-v4-proNoVision is marked not supported

How does a Sume agent get frames out of a video?

Two Sume tools return stills. Video inspect probes one hosted clip and returns sampled stills, with 8 mid-bin stills by default and up to 24 per call, plus an optional transcript from Sume STT 1.0. Probe and stills are unbilled. Video frames pulls exact stills at the times you name in at[] or at an fps, as durable media.sume.com images at source size unless you cap them with max_edge. It is also unbilled and always answers 202.

Both take a media.sume.com clip. They do not fetch an arbitrary web URL, so import first with POST /v1/media-imports. The frames come back as artifacts the agent can open, which is why image input on the orchestrator is the capability that matters.

curl -X POST https://api.sume.com/v1/video-frames \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-frames-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "at": [0, 2.5],
    "max_edge": 1024
  }'

When do you use inspect and when frames?

Use inspect when you do not know where to look. With no frames field it returns 8 mid-bin stills at a default max_edge of 768, a size that keeps the image input small, and a probe with the clip's facts. Pass frames: false for probe only, which is enough to check probe.has_audio before you ask for a transcript. The call is synchronous by default: it waits up to 30 seconds, then answers 200 with the result or 202 with a job to poll.

Use frames when you know the instant, such as the moment a price card should land, and want the source-size image. Inspect also takes seek: "fast" to skim a clip by snapping each still to the keyframe at or before the instant, up to one group of pictures early. Keep the default precise when a timestamp has to match.

A request for 24 stills at source size on every turn is the expensive way to do this. Ask for the few frames that decide the question.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-inspect-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": { "at": [0, 2.5, 5], "max_edge": 768 }
  }'

Which Sume rows cannot take images?

In the Sume catalog every row you would pick for a video agent is marked image-capable except one: DeepSeek V4 Flash 0731, which the repo labels text only, matching OpenRouter's text-only listing when it was added. It is off the picker and survives only so older stored picks keep resolving. Auto starts on DeepSeek V4.1 Flash, which the repo marks image-capable.

Image-capable on the row is not the same as the agent reading every frame well. Sume makes no claim about how accurately any model judges motion, lip sync, or text on screen, and a still cannot show timing at all.

How should you check a render with stills?

Pick frames that answer a question. For a product shot, ask for the first frame, the frame where the price card should appear, and the last frame, then ask the model one yes-or-no question per frame. Use fps only when you really do need coverage, because every extra frame is paid input on the next turn.

If a check hinges on audio, send for the transcript instead, because a model cannot hear a still. Video inspect's transcribe: true runs Sume STT on the clip's audio, and that is billed as speech-to-text. Then the LLM compares the words with the script you gave it.

Sources

Related posts

More in Models

All Models posts

Written by Sume