Grok 4.7 doubles price at 200k tokens: long video agent threads
xAI bills a Grok 4.7 request at the higher rate for every token once the prompt reaches 200k. What that does to a long video agent thread, with the arithmetic.

Grok 4.7 costs $2 per million input tokens and $6 per million output tokens while the prompt stays under 200k tokens, and $4 and $12 once it reaches 200k. xAI's page says the higher rate applies to all tokens in that request, not only the part above the line, so one request crossing the threshold can cost more than double.
These numbers come from xAI's Grok 4.7 page, read on 2026-10-02. The same page lists a 500,000-token context window, which means the tier boundary sits well inside it.
What does the jump look like in dollars?
Ten percent more prompt costs 2.2 times as much. Cached input follows the same rule: $0.50 per million below the line and $1 above it. If your thread is mostly cache hits, the jump is smaller in dollars but just as sharp in ratio.
| Prompt tokens | Rate tier | Input cost | Output cost (5,000 tokens) | Total |
|---|---|---|---|---|
| 190,000 | Under 200k: $2 in, $6 out | $0.380 | $0.030 | $0.410 |
| 210,000 | 200k and over: $4 in, $12 out | $0.840 | $0.060 | $0.900 |
Do other agent models have the same cliff?
Not all of them. Anthropic's pricing page says Claude 4.6 and later models include the full 1M-token window at standard pricing, so a 900k-token request costs the same per token as a 9k one. DeepSeek's pricing page lists a 1M window with no length tiers. OpenAI's GPT-6.1 Sol page lists 1,050,000 tokens of context and has its own cliff: prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.
So the cliff is not unique to Grok, only its position and size differ. The Anthropic and DeepSeek statements are about those pages today, not a guarantee. Re-read the page before you build a budget on it, and note that tokenizers differ: Anthropic says its 4.7-and-later tokenizer produces about 30% more tokens for the same text than earlier ones.
How does a video agent end up past 200k?
Rarely from one prompt. It happens when a thread keeps growing: every tool result, every shot list revision, every image the agent looked at stays in the next turn's input. A project with many scenes and a few image checks per scene can climb for a long time before anyone notices.
On Sume, continuing a run with previous_run_id replays what the agent produced so it can redo one scene and leave the rest alone. That is the right tool for a one-scene retry, and it is also how a thread grows. Sume's docs do not describe automatic trimming of a continued thread, so do not assume one.
Can you even run Grok 4.7 on Sume?
Check the picker for your workspace before you plan around it. In the repo, the grok-4.7 row sits behind a catalog gate that needs a Grok Build harness and the Grok catalog flag, and a request for an id outside the admitted catalog is a 400 invalid_request, never a silent switch to another model. The Formats API documents model as an Agents catalog id and says the receipt echoes the id that ran.
The repo also records that Grok 4.7 takes low, medium, high and xhigh effort, defaults to normal service tier only, and that Sume's picker offers Low, Medium and High for it.
How do you know where a thread sits against 200k?
Count, do not guess. Anthropic's pricing FAQ gives a rough rule of one token to about four English characters, and xAI's tokenizer is a different one, so 200k tokens is only loosely 800,000 characters. Images, tool definitions and tool results all count toward the prompt as well, and a video thread is heavy on the last two.
The reliable number is the prompt-token count xAI returns on each response. Sume's public receipt shows spend, not prompt size: it carries usage with the cap, the counted amount and the wallet amounts, but the receipt fields do not list a prompt-token count. So the practical signal on Sume is cost per turn climbing faster than the work done, which you can read from debited_usd_micros run over run.
What should you do about it?
The cliff is not a reason to avoid Grok 4.7. It is a reason to know your prompt size per turn, which the receipt lets you do.
- Split long projects into separate threads, one per scene group, instead of one endless conversation.
- Keep image checks small: a few stills with
max_edgecapped, not whole sequences. - Read
usage.debited_usd_microson the receipt after each run, because it includes the orchestrator's own turn. - If a single request must carry a very large prefix, compare it against a model without a length tier, such as one of the flat-priced 1M-window models above, on the same input before you commit.
Sources
Related posts
More in Models
- Grok Imagine on Sume: 4:5 returns 400, 1080x1350 does not
Grok Imagine has no 4:5 on Sume. aspect_ratio 4:5 returns 400, but 1080x1350 snaps to the nearest native ratio, 2:3. What to send for Instagram portrait.
- Grok Imagine video in Sume: why it needs a start-frame image
Sume's Grok Imagine entry is image-to-video only: it blocks a submit without a start frame, tops out at 10 seconds, and sends no audio or aspect ratio.
- H3 Max 3D to Video: previs to photoreal on fal, not on Sume
fal's H3 Max 3D-to-Video turns a blockout render into photoreal video for $0.50 a request plus per second. Sume lists no such endpoint; what it offers instead.
- H3 Max Insert-Video: add a scene mid-clip on fal, and on Sume
fal's H3 Max Insert-Video adds 5 to 13 s to a source up to 60 s, billed on the new seconds only. Sume lists no such row; here is a trim, generate, join route.
Written by Sume