DeepSeek V4.1 Flash: the model Sume Auto starts on, and what it takes

Sume's Auto agent model reserves and starts on DeepSeek V4.1 Flash. DeepSeek's page lists a 1M window, 384K output and vision. How the served model settles.

5 min readSume
All posts

Sume's Auto choice in the Agents model picker reserves and starts on DeepSeek V4.1 Flash, according to the registry's comment for the agents/auto row. Auto then settles on whichever model actually served the turn, so "Auto" and "DeepSeek" are not the same promise. Pin an id when you need to know.

What is the Auto model in the Sume Agents picker?

Auto is a router row, not a model. Its tooltip says Sume routes automatically, and the registry marks it as the default for composing and for requests that omit a model on the Agents catalog. The start model, the one a reservation is made against, is DeepSeek V4.1 Flash. Usage then records the served model after the turn.

Two cautions follow. This is the agent LLM that plans and calls tools, not the model that renders video: video, image and audio models are chosen separately, and sume/auto for video has its own rules. And a Format run's omitted model is a different default from Auto's, covered in the GPT-6.1 Sol default post.

What does DeepSeek say V4.1 Flash is?

DeepSeek's pricing page lists two API model names, deepseek-flash for DeepSeek-V4.1-Flash and deepseek-v4-pro for DeepSeek-V4-Pro-0813:

  • DeepSeek's ranges reflect peak and off-peak pricing, with off-peak at half the peak rate; peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday.
  • DeepSeek says legacy model names are still accepted but are served by DeepSeek-V4.1-Flash and billed at the Flash price.
DeepSeek API models and pricing page (read 2026-10-02)
Itemdeepseek-flashdeepseek-v4-pro
ModelDeepSeek-V4.1-FlashDeepSeek-V4-Pro-0813
Context length1M tokens1M tokens
Max output384K384K
VisionSupportedNot supported
Output price per 1M tokens$0.6 to $1.2$1.98 to $3.96

How does Sume name the DeepSeek models?

Sume carries several DeepSeek rows, and only some are live. The live pick is DeepSeek V4.1 Flash, with a tooltip of 1M context and 384k max output. A beta spelling of V4.1 Flash is kept for already-reserved turns and prints as V4.1 Flash. A dated V4 Flash 0731 row is text-only and no longer on the picker, and an experimental V4 Flash Vision row is off the picker as well; both remain resolvable for stored picks.

The practical consequence is about images. The registry marks V4.1 Flash as accepting image input and the 0731 snapshot as text-only. If you attach images to an agent run, make sure you are not on a stored 0731 pick that quietly drops them. The Sume docs allow up to 30 images on a Format run through attachments, per Create a run.

Is DeepSeek's list price what I pay through Sume?

No. DeepSeek's page is DeepSeek's own rate card, with peak and off-peak ranges, and it applies to calls made directly against DeepSeek's API. A Sume run is billed under Sume's rates for the model that served it, and the amount is on the receipt's usage object. Use the DeepSeek numbers to see why Flash is a cheap default, such as the 1M context and low output price, and use the receipt for what a run actually cost.

The same goes for caching and the off-peak discount: they are facts about DeepSeek's API, not promises about a Sume invoice.

When should I not rely on Auto?

When you must reproduce a result, compare two orchestrators, or explain a bill to a client. Auto settles on the served model, so two Auto runs can be on different models. Pin gpt-6.1-sol or another catalog id, run both with the same input, and compare each receipt's model and usage.

Auto is a good default for exploring, because it needs no decision. Keep the pinned run for anything you will repeat. The related Sume Auto video post makes the same point for the media models.

Sources

Related posts

More in Models

All Models posts

Written by Sume