DeepSeek V4.1 Flash sees images: hand one to a Sume Agent Completion
V4.1 Flash is described as natively multimodal. To act on an image with Sume, pass it as an input_image attachment on an Agent Completion, up to 30 per run.

If your planner is DeepSeek V4.1 Flash and it has looked at a reference image, send the image on to Sume as an input_image part of an Agent Completion. DeepSeek's September 10 notice describes V4.1 Flash as having native multimodal and visual understanding (read 2026-10-05). The Sume side takes up to 30 images per call and runs them in the Sume Agent with its own tools and sandbox.
What the DeepSeek notice says
The page names the model DeepSeek-V4.1-Flash with API model id deepseek-flash and describes a 552B-parameter mixture-of-experts design with native visual understanding (read 2026-10-05). It does not state tool-calling details or a context window in the text I could read, so this post makes no claim about either.
| Item | Value | Source |
|---|---|---|
| API model id | deepseek-flash | DeepSeek notice |
| Visual input | Native multimodal understanding | DeepSeek notice |
| Image attachments per Sume completion | Maximum 30 | Sume docs |
| Size limits | 30 MB each, 500 MB in total | Sume docs |
| Part type | input_image with an HTTPS image_url or an asset_id | Sume docs |
Passing the image to Sume
The Agent Completion request uses an OpenAI-style messages[] shape. A turn can carry an input_text part and one or more input_image parts. Sume merges top-level attachments and inline image parts into one list. An image-only turn is allowed; Sume then tells the agent to use the attached files.
Each image needs a public HTTPS URL or an asset_id, not both. If Sume cannot fetch the URL, the call fails with 502 attachment_fetch_failed, which the docs attribute to an unreachable host, hotlink protection or a non-2xx response.
import json, os, urllib.request
body = {
"messages": [{"role": "user", "content": [
{"type": "input_text",
"text": "Describe this product shot in one line."},
{"type": "input_image",
"image_url": "https://cdn.example.com/shot.jpg"}]}],
"generation_spend_cap_usd": 1,
}
req = urllib.request.Request(
"https://api.sume.com/v1/agent/completions",
data=json.dumps(body).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"},
method="POST")
with urllib.request.urlopen(req) as r:
print(json.load(r)["data"]["status_url"])Keep the planner and the worker apart
Your DeepSeek agent and the Sume Agent are different models with different inputs. Describing the image in your own prompt does not give Sume the pixels. Attach them. Equally, model on the Sume call accepts only sume-agent, so the DeepSeek id never goes in that field.
Cap the run. generation_spend_cap_usd is required, and a vision task that triggers an edit or a render can spend.
Structured results from an image task
If the DeepSeek planner needs a machine-readable answer, bind the run to your own schema with output_schema. Attachments and output_schema work together: the images go to the Sume Agent, and after the run completes Sume parses the output against your schema. The default output carries the agent's last text in output.text and any generated media in output.images, output.videos, output.audio and output.files.
Because the planner and the worker are separate models, return only the fields the planner needs. A short caption string is easier for the planner to consume than a full receipt, and it keeps the planner's context small.
Checklist
Before you wire the two together, confirm the points below.
- Image URLs are public HTTPS and not hotlink-protected.
- Total images are at most 30, each at most 30 MB.
- The cap fits the worst case you accept for this one task.
- You poll the receipt's
status_urlinstead of re-submitting.
Sources
Related posts
More in Agents
- Does the Sume run spend cap include the LLM turn? No, only generation
On a Sume run receipt, billable_amount_usd_micros is generation spend only; the agent's LLM turn bills a separate Agent wallet. How to budget both.
- Fire a Sume schedule from a decision, not a clock: API trigger
Set api_trigger_enabled on a Sume schedule, let Clef or Decider decide when it is worth running, then POST /v1/actions/{action_id}/runs with an idempotency key.
- Five generate_image calls in one turn: Sume MCP create budget
Hosted Sume MCP gives paid creates their own budget (20 per principal, 64 per process) and queues a call up to 20 seconds before wait_busy.
- usage.cap on a Format run: an agent reads remaining_usd_micros
A Format run receipt splits its spend cap into limit, counted and remaining USD micros. A supervising agent can read headroom before it asks for more work.
Written by Sume