Which agent LLM can read video frames? Image input by model
GPT-6.1 Sol, Claude 5.5, Grok 4.7 and DeepSeek Flash accept images; one DeepSeek row is text only. How a Sume agent checks a render with stills.

Yes, for stills. GPT-6.1 Sol, every current Claude model, Grok 4.7 and DeepSeek's deepseek-flash all take image input on the vendors' pages, so an agent running any of them can look at frames pulled from a render. DeepSeek's deepseek-v4-pro does not accept images, and Sume keeps one text-only DeepSeek row, V4 Flash 0731, for stored picks.
None of these models is handed an MP4 by Sume. The agent calls a frame tool, gets image artifacts back, and reads those. That distinction decides how you build a quality check.
What does each vendor say about image input?
Input is only half of it. Image tokens are billed as input, so a check that sends 24 full-size frames on every turn is not free. Each vendor prices images in its own way, and I left those rules out because they were not on the pages above.
| Model | Image input | What the page says |
|---|---|---|
| GPT-6.1 Sol | Yes | Image input is in its supported-features list |
| Claude Opus 5.5 and Sonnet 5.5 | Yes | All current models take text and image input |
| Grok 4.7 | Yes | Modalities listed as text, image to text |
| DeepSeek deepseek-flash | Yes | Vision is marked supported |
| DeepSeek deepseek-v4-pro | No | Vision is marked not supported |
How does a Sume agent get frames out of a video?
Two Sume tools return stills. Video inspect probes one hosted clip and returns sampled stills, with 8 mid-bin stills by default and up to 24 per call, plus an optional transcript from Sume STT 1.0. Probe and stills are unbilled. Video frames pulls exact stills at the times you name in at[] or at an fps, as durable media.sume.com images at source size unless you cap them with max_edge. It is also unbilled and always answers 202.
Both take a media.sume.com clip. They do not fetch an arbitrary web URL, so import first with POST /v1/media-imports. The frames come back as artifacts the agent can open, which is why image input on the orchestrator is the capability that matters.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-frames-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"at": [0, 2.5],
"max_edge": 1024
}'When do you use inspect and when frames?
Use inspect when you do not know where to look. With no frames field it returns 8 mid-bin stills at a default max_edge of 768, a size that keeps the image input small, and a probe with the clip's facts. Pass frames: false for probe only, which is enough to check probe.has_audio before you ask for a transcript. The call is synchronous by default: it waits up to 30 seconds, then answers 200 with the result or 202 with a job to poll.
Use frames when you know the instant, such as the moment a price card should land, and want the source-size image. Inspect also takes seek: "fast" to skim a clip by snapping each still to the keyframe at or before the instant, up to one group of pictures early. Keep the default precise when a timestamp has to match.
A request for 24 stills at source size on every turn is the expensive way to do this. Ask for the few frames that decide the question.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-inspect-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": { "at": [0, 2.5, 5], "max_edge": 768 }
}'Which Sume rows cannot take images?
In the Sume catalog every row you would pick for a video agent is marked image-capable except one: DeepSeek V4 Flash 0731, which the repo labels text only, matching OpenRouter's text-only listing when it was added. It is off the picker and survives only so older stored picks keep resolving. Auto starts on DeepSeek V4.1 Flash, which the repo marks image-capable.
Image-capable on the row is not the same as the agent reading every frame well. Sume makes no claim about how accurately any model judges motion, lip sync, or text on screen, and a still cannot show timing at all.
How should you check a render with stills?
Pick frames that answer a question. For a product shot, ask for the first frame, the frame where the price card should appear, and the last frame, then ask the model one yes-or-no question per frame. Use fps only when you really do need coverage, because every extra frame is paid input on the next turn.
If a check hinges on audio, send for the transcript instead, because a model cannot hear a still. Video inspect's transcribe: true runs Sume STT on the clip's audio, and that is billed as speech-to-text. Then the LLM compares the words with the script you gave it.
Sources
Related posts
More in Models
- Which Flow model can extend a video? Veo 3.1 Lite today, Omni soon
Flow's model table lets only Veo 3.1 Lite extend a clip, with an 8-second limit; Omni lists Extend as coming soon. Here is what to do on Sume meanwhile.
- Which AI image model edits a photo best? Reference limits compared
Reference-image limits for edits on GPT Image 2.5, Nano Banana, Grok Imagine Image 2.0 and FLUX, from each vendor's page, and how to call them on Sume.
- Which Sume image models accept aspect_ratio auto on an edit?
Only GPT Image, Nano Banana and Seedream 4.0 list auto on Sume. Grok, Seedream 4.5 and 5.0 Lite, and FLUX.2 return 400. Table plus a nearest-ratio helper.
- Zalando designer minimum 1800x2600: request 1808x2608 on Sume
Zalando designer brands need at least 1800x2600 px in 1:1.44 JPEG. GPT custom sizes need multiples of 16, so ask Sume for 1808x2608, which clears both edges.
Written by Sume