Grok 4.7 doubles price at 200k tokens: long video agent threads

xAI bills a Grok 4.7 request at the higher rate for every token once the prompt reaches 200k. What that does to a long video agent thread, with the arithmetic.

5 min readSume
All posts

Grok 4.7 costs $2 per million input tokens and $6 per million output tokens while the prompt stays under 200k tokens, and $4 and $12 once it reaches 200k. xAI's page says the higher rate applies to all tokens in that request, not only the part above the line, so one request crossing the threshold can cost more than double.

These numbers come from xAI's Grok 4.7 page, read on 2026-10-02. The same page lists a 500,000-token context window, which means the tier boundary sits well inside it.

What does the jump look like in dollars?

Ten percent more prompt costs 2.2 times as much. Cached input follows the same rule: $0.50 per million below the line and $1 above it. If your thread is mostly cache hits, the jump is smaller in dollars but just as sharp in ratio.

Grok 4.7 list price for one request, xAI page, read 2026-10-02
Prompt tokensRate tierInput costOutput cost (5,000 tokens)Total
190,000Under 200k: $2 in, $6 out$0.380$0.030$0.410
210,000200k and over: $4 in, $12 out$0.840$0.060$0.900

Do other agent models have the same cliff?

Not all of them. Anthropic's pricing page says Claude 4.6 and later models include the full 1M-token window at standard pricing, so a 900k-token request costs the same per token as a 9k one. DeepSeek's pricing page lists a 1M window with no length tiers. OpenAI's GPT-6.1 Sol page lists 1,050,000 tokens of context and has its own cliff: prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.

So the cliff is not unique to Grok, only its position and size differ. The Anthropic and DeepSeek statements are about those pages today, not a guarantee. Re-read the page before you build a budget on it, and note that tokenizers differ: Anthropic says its 4.7-and-later tokenizer produces about 30% more tokens for the same text than earlier ones.

How does a video agent end up past 200k?

Rarely from one prompt. It happens when a thread keeps growing: every tool result, every shot list revision, every image the agent looked at stays in the next turn's input. A project with many scenes and a few image checks per scene can climb for a long time before anyone notices.

On Sume, continuing a run with previous_run_id replays what the agent produced so it can redo one scene and leave the rest alone. That is the right tool for a one-scene retry, and it is also how a thread grows. Sume's docs do not describe automatic trimming of a continued thread, so do not assume one.

Can you even run Grok 4.7 on Sume?

Check the picker for your workspace before you plan around it. In the repo, the grok-4.7 row sits behind a catalog gate that needs a Grok Build harness and the Grok catalog flag, and a request for an id outside the admitted catalog is a 400 invalid_request, never a silent switch to another model. The Formats API documents model as an Agents catalog id and says the receipt echoes the id that ran.

The repo also records that Grok 4.7 takes low, medium, high and xhigh effort, defaults to normal service tier only, and that Sume's picker offers Low, Medium and High for it.

How do you know where a thread sits against 200k?

Count, do not guess. Anthropic's pricing FAQ gives a rough rule of one token to about four English characters, and xAI's tokenizer is a different one, so 200k tokens is only loosely 800,000 characters. Images, tool definitions and tool results all count toward the prompt as well, and a video thread is heavy on the last two.

The reliable number is the prompt-token count xAI returns on each response. Sume's public receipt shows spend, not prompt size: it carries usage with the cap, the counted amount and the wallet amounts, but the receipt fields do not list a prompt-token count. So the practical signal on Sume is cost per turn climbing faster than the work done, which you can read from debited_usd_micros run over run.

What should you do about it?

The cliff is not a reason to avoid Grok 4.7. It is a reason to know your prompt size per turn, which the receipt lets you do.

  • Split long projects into separate threads, one per scene group, instead of one endless conversation.
  • Keep image checks small: a few stills with max_edge capped, not whole sequences.
  • Read usage.debited_usd_micros on the receipt after each run, because it includes the orchestrator's own turn.
  • If a single request must carry a very large prefix, compare it against a model without a length tier, such as one of the flat-priced 1M-window models above, on the same input before you commit.

Sources

Related posts

More in Models

All Models posts

Written by Sume