DeepSeek legacy names now run V4.1 Flash: what Sume offers
DeepSeek still accepts deepseek-v4-flash and the vision-exp name but serves V4.1 Flash. Which DeepSeek rows Sume lists, and why there is no v4-pro row.

DeepSeek's pricing page says the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but their models are retired and requests are served by DeepSeek-V4.1-Flash at the Flash price. The current names are deepseek-flash and deepseek-v4-pro.
This is from DeepSeek's Models & Pricing page, read on 2026-10-02. The page does not say when the legacy names stop being accepted.
What does DeepSeek list today?
Both are on one page with the same limits. Pro costs about 4.4 times Flash on input and 3.3 times on output at these rates. The page quotes off-peak rates, so check it for the peak rate before you budget.
| Model | Context | Max output | Input per M | Output per M |
|---|---|---|---|---|
| deepseek-flash | 1M | 384K | $0.15 | $0.60 |
| deepseek-v4-pro | 1M | 384K | $0.66 | $1.98 |
Which DeepSeek rows does Sume have?
The repo's registry has these DeepSeek rows, and all of them route through OpenRouter and are behind the OpenRouter catalog gate:
DeepSeek V4.1 Flash is the live row, listed with the note 1M context and 384k max output, and it is the row Auto starts on. DeepSeek V4 Flash 0731 is a dated snapshot, described as text only; it is off the picker and kept for stored picks and alias resolution. DeepSeek V4 Flash Vision Exp is the experimental image-understanding row. A beta spelling of V4.1 Flash exists only for turns already reserved, and every spelling of the beta runs as V4.1 Flash on a new request.
There is no deepseek-v4-pro row in the registry, so the Formats model field cannot name it. An id outside the catalog is 400 invalid_request per the Formats API.
How do the names line up?
Two naming schemes meet here, DeepSeek's own and Sume's catalog ids, and they are not the same strings. A Sume id carries the openrouter/ prefix and a hyphenated slug, for example openrouter/deepseek-deepseek-v4.1-flash, while DeepSeek's own API names the same family deepseek-flash. Use the catalog id on Sume and the vendor name only on the vendor's API.
- On Sume, send the catalog id in
model, and read the id back from the receipt. - On DeepSeek's API, use
deepseek-flashordeepseek-v4-pro, not the legacy names. - Do not infer a Sume route from a vendor name. The repo's row list is the source.
Should a video agent use DeepSeek Flash?
It is the cheapest orchestrator in this comparison, and the same Flash row accepts images in the repo's catalog, which matters when the agent has to look at frames. Whether its plans are good enough for your Format is a test, not a claim: run the same input on it and on a pricier model and compare the receipts, as in the A/B recipe.
If you need the 0731 text-only row, remember it cannot take images, so a Format that sends attachments will not suit it.
What does a cheap orchestrator cost you?
The price gap is easy to read and the quality gap is not. At DeepSeek's off-peak rates a turn of 20,000 input tokens and 4,000 output tokens costs 20,000 times $0.15 per million, $0.003, plus 4,000 times $0.60 per million, $0.0024, so about half a cent. The numbers are DeepSeek's list prices from its page and ignore caching and peak rates, which are double.
For a video agent the orchestrator turn is usually a small share of a run that also pays for video generation, so saving a few cents per turn matters most when an agent loops through many turns. If a cheaper model needs more turns or more re-renders, the saving disappears. That is why the comparison is on usage.debited_usd_micros for a finished run, not on token price alone.
Sume does not claim DeepSeek is the best or the worst model for your Format. It lists the row, starts Auto there, and leaves the result to your own test.
Sources
Related posts
More in Models
- Does H3 Max Recast keep the original audio? What to check on Sume
fal says H3 Max Recast preserves the source audio. Sume rejects generate_audio and audio references on it. Confirm a result has sound with video inspect.
- Does sume/auto pick Seedance? Sume does not say which model ran
Sume says sume/auto echoes sume/auto in the response and never discloses the family that served the request. To get Seedance, pin the id.
- FastH3 on vLLM-Omni: a 10-second H3 video in 8.7 seconds
vLLM's team reports 10.1 seconds of H3 video and audio in about 8.7 s on 8 B300 GPUs. What the number covers, and what Sume's hosted minimax-h3 does.
- Flow Agent picks the image model for you: Sume's sume/auto does too
Flow's Agent routes image requests to the best model automatically. Sume's sume/auto does too but never says which family ran; pin an id for a matched series.
Written by Sume