DeepSeek V4 Pro API after Sept 14: is there a Sume row?
DeepSeek says V4 Pro API service continues past Sept 14 at unchanged billing. Sume's agent catalog lists V4.1 Flash and older Flash rows, but no V4 Pro.

DeepSeek confirmed that it will keep serving DeepSeek V4 Pro through its API beyond September 14, 2026, with billing unchanged. Sume's agent catalog, however, has no V4 Pro row; the DeepSeek entries are V4.1 Flash, V4 Flash 0731 and a vision experiment, so you cannot choose V4 Pro as the orchestrator for a Sume agent or Format.
The statement comes from the DeepSeek API updates page (entry of September 10, 2026), read 2026-10-03, and the catalog claim from the Sume model catalog source read the same day.
What else did that DeepSeek entry say?
The same entry announces DeepSeek-V4.1-Flash, described as the smallest model of a new architecture family with native multimodal visual understanding. It tells users to set the model name deepseek-flash for the latest V4.1 Flash, says the earlier names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash, and says API prices were reduced with the release. The page quotes benchmark scores, which this post does not repeat.
Which DeepSeek rows does Sume list?
From the catalog source, the DeepSeek rows are below. They sit behind the OpenRouter gate, so even V4.1 Flash may be absent from a given workspace's picker; check the picker before you plan around it.
| Row label | Notes in the catalog |
|---|---|
| DeepSeek V4.1 Flash | 1M context, 384k max output; the one new picks use |
| DeepSeek V4 Flash 0731 | Text only; off the picker, kept for stored picks and aliases |
| DeepSeek V4 Flash Vision Exp | Experimental, off the picker; aliases stay resolvable |
| DeepSeek V4 Pro | No row |
What happens if I send a V4 Pro id to a Format?
The model field on Calling a Format takes an Agents catalog id. An id outside the catalog is 400 invalid_request. A V4 Pro request therefore fails immediately instead of being routed to a Flash row, which is better than a silent swap but means you need a different model name in the body.
curl -X POST https://api.sume.com/v1/formats/sume/sume-video-hook/runs \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: deepseek-pin-001" \
-d '{"instruction":"Three hook ideas for a vitamin C serum",
"model":"openrouter/deepseek-deepseek-v4.1-flash",
"generation_spend_cap_usd":5}'When does V4 Pro matter for video work at all?
An agent that orchestrates video jobs spends its LLM tokens on planning, tool arguments and review notes, not on pixels. A larger text model helps when the plan is complicated (many shots, strict brand rules); it does not change what the video models render. If V4 Pro is the model you trust for planning, run it against DeepSeek's API directly and hand the finished shot list to Sume through the public API; if you want the whole loop inside Sume, V4.1 Flash is the DeepSeek option today.
What should I do if I depend on V4 Pro?
Keep it where it is served. The updates page says DeepSeek's API continues to serve V4 Pro with unchanged billing, so there is no deadline forcing a move this month. If part of your pipeline is a V4 Pro planning step, leave that step on DeepSeek's API and treat Sume as the place the finished plan goes next: one request to a Format or a model endpoint per shot, with the plan carried in the request body.
If you later want the planning inside Sume, run the same fixed set of prompts on V4.1 Flash and compare the plans yourself. DeepSeek's page quotes benchmark numbers for the Flash release, but those say little about whether your shot lists come out the same, and I have no basis for claiming they do.
Sources
Related posts
More in Models
- Did OpenAI change GPT Image 2.5 since launch? Changelog check, Oct 3
OpenAI's changelog shows GPT Image 2.5 shipped Sept 8 and no October image entry. The Sept 25 image fix covers GPT-6 inputs, not generation.
- DMAD 4-step MiniMax H3: 50-step baseline vs Turbo LoRA
DMAD distills MiniMax H3 from 50 steps to 4. Why that 4 is not the Turbo LoRA's 4, what the card says, what it omits, and where hosted jobs fit.
- End-frame AI video: which Sume models take a last frame
Seedance, Wan 3.0, Kling 3, MiniMax and Gemini Omni Flash accept a first and last frame; Grok Imagine and the swap rows do not. How to send both frames.
- Fastest Seedance or Kling model on Sume: latency labels
Sume labels seedance-2-fast and seedance-2-mini fast, seedance-2.5 and seedance-2 medium, and kling-3 medium-slow. What the labels mean and how to test.
Written by Sume