DeepSeek V4.1 Flash: the model Sume Auto starts on, and what it takes
Sume's Auto agent model reserves and starts on DeepSeek V4.1 Flash. DeepSeek's page lists a 1M window, 384K output and vision. How the served model settles.

Sume's Auto choice in the Agents model picker reserves and starts on DeepSeek V4.1 Flash, according to the registry's comment for the agents/auto row. Auto then settles on whichever model actually served the turn, so "Auto" and "DeepSeek" are not the same promise. Pin an id when you need to know.
What is the Auto model in the Sume Agents picker?
Auto is a router row, not a model. Its tooltip says Sume routes automatically, and the registry marks it as the default for composing and for requests that omit a model on the Agents catalog. The start model, the one a reservation is made against, is DeepSeek V4.1 Flash. Usage then records the served model after the turn.
Two cautions follow. This is the agent LLM that plans and calls tools, not the model that renders video: video, image and audio models are chosen separately, and sume/auto for video has its own rules. And a Format run's omitted model is a different default from Auto's, covered in the GPT-6.1 Sol default post.
What does DeepSeek say V4.1 Flash is?
DeepSeek's pricing page lists two API model names, deepseek-flash for DeepSeek-V4.1-Flash and deepseek-v4-pro for DeepSeek-V4-Pro-0813:
- DeepSeek's ranges reflect peak and off-peak pricing, with off-peak at half the peak rate; peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday.
- DeepSeek says legacy model names are still accepted but are served by DeepSeek-V4.1-Flash and billed at the Flash price.
| Item | deepseek-flash | deepseek-v4-pro |
|---|---|---|
| Model | DeepSeek-V4.1-Flash | DeepSeek-V4-Pro-0813 |
| Context length | 1M tokens | 1M tokens |
| Max output | 384K | 384K |
| Vision | Supported | Not supported |
| Output price per 1M tokens | $0.6 to $1.2 | $1.98 to $3.96 |
How does Sume name the DeepSeek models?
Sume carries several DeepSeek rows, and only some are live. The live pick is DeepSeek V4.1 Flash, with a tooltip of 1M context and 384k max output. A beta spelling of V4.1 Flash is kept for already-reserved turns and prints as V4.1 Flash. A dated V4 Flash 0731 row is text-only and no longer on the picker, and an experimental V4 Flash Vision row is off the picker as well; both remain resolvable for stored picks.
The practical consequence is about images. The registry marks V4.1 Flash as accepting image input and the 0731 snapshot as text-only. If you attach images to an agent run, make sure you are not on a stored 0731 pick that quietly drops them. The Sume docs allow up to 30 images on a Format run through attachments, per Create a run.
Is DeepSeek's list price what I pay through Sume?
No. DeepSeek's page is DeepSeek's own rate card, with peak and off-peak ranges, and it applies to calls made directly against DeepSeek's API. A Sume run is billed under Sume's rates for the model that served it, and the amount is on the receipt's usage object. Use the DeepSeek numbers to see why Flash is a cheap default, such as the 1M context and low output price, and use the receipt for what a run actually cost.
The same goes for caching and the off-peak discount: they are facts about DeepSeek's API, not promises about a Sume invoice.
When should I not rely on Auto?
When you must reproduce a result, compare two orchestrators, or explain a bill to a client. Auto settles on the served model, so two Auto runs can be on different models. Pin gpt-6.1-sol or another catalog id, run both with the same input, and compare each receipt's model and usage.
Auto is a good default for exploring, because it needs no decision. Keep the pinned run for anything you will repeat. The related Sume Auto video post makes the same point for the media models.
Sources
Related posts
More in Models
- Does H3 Max Recast keep the original audio? What to check on Sume
fal says H3 Max Recast preserves the source audio. Sume rejects generate_audio and audio references on it. Confirm a result has sound with video inspect.
- Does sume/auto pick Seedance? Sume does not say which model ran
Sume says sume/auto echoes sume/auto in the response and never discloses the family that served the request. To get Seedance, pin the id.
- FastH3 on vLLM-Omni: a 10-second H3 video in 8.7 seconds
vLLM's team reports 10.1 seconds of H3 video and audio in about 8.7 s on 8 B300 GPUs. What the number covers, and what Sume's hosted minimax-h3 does.
- Flow Agent picks the image model for you: Sume's sume/auto does too
Flow's Agent routes image requests to the best model automatically. Sume's sume/auto does too but never says which family ran; pin an id for a matched series.
Written by Sume