Agent model token prices Oct 4: Argon, Opus, Sonnet, Sol, Grok

Per-million-token prices for the five agent models launched in September, and why the receipt's usage.billable figure leaves the LLM turn out.

5 min readSume
All posts

As of October 4, 2026, input prices for the five new agent models range from $2 to $4 per million tokens and output prices from $6 to $20. Gemini 4 Argon and Claude Opus 5.5 sit at the top, while Claude Sonnet 5.5, GPT-6.1 Sol and Grok 4.7 start at $2 input. Every figure below comes from the vendor's own page, read today.

For a video agent the table matters less than it looks: the orchestrating model's tokens are a small line next to generation spend, and Sume reports the two differently.

What does each model cost per million tokens?

Standard rates only; batch, flex and regional options are left out.

Standard API prices per million tokens, from each vendor's page, read 2026-10-04.
ModelInputOutputNotes
Gemini 4 Argon$2 introductory, $4 standard$10 introductory, $20 standardNot yet publicly available; cached input has a 95% discount
Claude Opus 5.5$4$20Cache reads $0.20; fast mode $8 input and $40 output
Claude Sonnet 5.5$2$10Cache reads $0.20; cache writes $2.50
GPT-6.1 Sol$2$10Cached input $0.10
Grok 4.7$2$6$4 and $12 at 200k-token prompts and above

Does a Sume run's usage figure include the model's tokens?

Not in usage.billable_amount_usd_micros. Sume's Runs and results page says that field is the generation spend attributed to the run, the total its spend cap is enforced against, and that it excludes the agent's own LLM turn, so it is not the run's total cost.

The cost is usage.debited_usd_micros: what the wallet actually deducted for the run and its thread, across every captured ledger row type, the turn's own LLM row included. held_usd_micros and refunded_usd_micros are holds, not spend, and final turns true once no hold is open.

Which number should I use to budget?

Use each for a different job.

  • Set generation_spend_cap_usd from the media you expect the run to make; that is what the cap counts.
  • Read usage.debited_usd_micros after the run when you want the real wallet cost.
  • Wait for usage.final to be true before you reconcile.
  • Compare models on debited from receipts of the same input, not on the list prices above.

What would change this table?

Argon's introductory price is time-limited by Google's own wording, so its row will move. The other four are standard rates and can also change; re-read the vendor page before you commit a budget to them.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume