Context window and max output for agent LLMs, compared
Context and output limits for GPT-6.1 Sol, Opus 5.5, Sonnet 5.5, Grok 4.7 and DeepSeek Flash from vendor pages, and what they mean for a long video thread.

Four of the five models here advertise a context window of about a million tokens: GPT-6.1 Sol at 1,050,000, Claude Opus 5.5 and Sonnet 5.5 at 1M, and DeepSeek's deepseek-flash at 1M. Grok 4.7 lists 500,000. Maximum output is 128K on GPT-6.1 Sol and both Claude models, and 384K on DeepSeek; xAI's page does not state one for Grok 4.7.
The figures are from the vendors' own pages, read on 2026-10-02: GPT-6.1 Sol, GPT-6 Sol, Anthropic's models overview, Grok 4.7 and DeepSeek's Models & Pricing.
What are the limits side by side?
OpenAI's pages split the window: of 1,050,000 tokens, 922,000 can be input and 128,000 output, which adds up to the whole window. For the others, the output ceiling counts inside the window.
| Model | Context window | Max input | Max output |
|---|---|---|---|
| GPT-6.1 Sol | 1,050,000 | 922,000 | 128,000 |
| GPT-6 Sol | 1,050,000 | 922,000 | 128,000 |
| Claude Opus 5.5 | 1M | not stated | 128K |
| Claude Sonnet 5.5 | 1M | not stated | 128K |
| Grok 4.7 | 500,000 | not stated | not stated |
| DeepSeek deepseek-flash | 1M | not stated | 384K |
Why is a token count not comparable across vendors?
Each vendor counts with its own tokenizer. Anthropic says its current tokenizer, introduced with Claude Opus 4.7, makes 1M tokens roughly 555k words, where earlier models fit about 750k words in the same count, and its pricing page says the newer tokenizer produces about 30% more tokens for the same text. So a 1M window on Claude holds less text than 1M on a model with a more compact tokenizer.
Tokens are also not the only thing in the window. Tool definitions, tool results, reasoning and images all fill it. For a video agent the heavy parts are usually tool results, such as job status payloads and probe data, and any stills the agent looked at.
Does max output matter for a video agent?
Rarely for the final answer. A shot list, a script and a JSON result fit in a few thousand tokens. Where it matters is thinking: on reasoning models, thinking counts toward the output limit. Anthropic's effort page notes thinking counts toward max_tokens even when the thinking content is not returned, and for agentic coding on Sonnet 5.5 it suggests setting max_tokens to the 128,000 maximum.
So the output ceiling is a guard against a high-effort turn being cut off mid-thought, not a measure of how long a script you can ask for.
What does Sume promise about thread size?
Nothing beyond what its docs say, and the docs do not give a usable-window figure per model. A Format run is one agent turn, and a new run with previous_run_id continues the conversation by replaying what the agent produced, per Runs and results. The Formats API limits the request: an input object up to 2 MiB, at most 64 top-level keys, an instruction up to 8,000 characters, and up to 30 image attachments.
The repo's picker copy for DeepSeek V4.1 Flash reads 1M context, 384k max output, the same figures as DeepSeek's page. For the other rows it carries the vendors' one-line descriptions and no window figures, so use the vendor numbers as upper bounds, not as promises about a Sume thread.
How should you use the window?
Pick on price per turn and quality first. The window only decides whether a thread can grow, and a well-run agent rarely needs most of it.
If you want a number for your own work, run one representative project on two models and read the receipt for each. The token counts differ by tokenizer, so compare the dollar figure, not the raw count, and keep the one that finishes the job for less.
- Treat the window as a ceiling and budget for a fraction of it, because cost grows with every turn that resends the thread.
- Keep big payloads out of the conversation. Pass a URL or an asset id, not a pasted file.
- Start a fresh thread for an unrelated project rather than continuing a long one.
- For Grok 4.7, remember the price tier changes at 200k, well inside its 500,000 window.
Sources
- OpenAI API: GPT-6.1 Sol model page (read 2026-10-02)
- OpenAI API: GPT-6 Sol model page (read 2026-10-02)
- Anthropic: Models overview (read 2026-10-02)
- Anthropic: Pricing (read 2026-10-02)
- xAI Docs: Grok 4.7 (read 2026-10-02)
- DeepSeek API Docs: Models & Pricing (read 2026-10-02)
- Create a run (Formats API)
- Runs and results (Formats API)
Related posts
More in Models
- Dialogue in Korean, Japanese or Spanish: Kling 3.0 vs Seedance 2.5
Which spoken languages Kling 3.0 Omni and Seedance 2.5 list, what Sume's generate_audio flag does, and how to test a Korean line before a full render.
- Which AI video model for a 15-second single take on Sume?
Kling 3.0, Wan 3.0, MiniMax H3 and Seedance 2 reach 15 seconds on Sume; Auto and Grok stop at 10. Seedance 2.5 and Wan 3.0 go to 30. Table of limits.
- POST /v1/avatar-1.0/fabric is test-only: use veed/fabric-1.0 instead
The experimental avatar-1.0/fabric route is a temporary comparison endpoint that may be removed. How it differs from image-to-video and what to build on.
- ByteDance MPA copyright pact: what Seedance users should check
ByteDance signed a copyright agreement with the MPA on Aug 17, 2026 covering Seedance. What was disclosed, what was not, and how to handle a refused request.
Written by Sume