GPT-6 Sol vs Luna vs Astra: context, prices, cutoffs

GPT-6 Astra, Sol and Luna share a 1,050,000-token window but differ 100x in price. A table from OpenAI's own pages, plus which one Sume runs Formats on.

5 min readSume
All posts

GPT-6 Astra, Sol and Luna all have a 1,050,000-token context window and 128,000 maximum output tokens, but their API prices span a factor of 100 on input: Luna is $0.1 per million input tokens, Sol $2 and Astra $10. On Sume, a Format run orchestrates on gpt-6-sol unless you pass a different Agents catalog id.

What do OpenAI's pages list side by side?

Figures below are copied from each model page read on 2026-10-02. Prices are per one million tokens. The Sol page also lists a 922,000-token maximum input; the other two pages I read did not state one, so that cell is left open.

GPT-6 model pages, OpenAI (read 2026-10-02)
ModelInputCached inputOutputContextKnowledge cutoff
gpt-6-astra$10$1$501,050,000April 30, 2026
gpt-6-sol$2$0.2$101,050,000Apr 20, 2026
gpt-6-luna$0.1$0.01$0.51,050,000May 18, 2026

What else is shared?

All three accept text and image input and return text only, and each lists Chat Completions, Responses and Batch among its endpoints. Astra's page lists streaming, structured outputs, function calling, file search, image input, web search and prompt caching. Luna's page lists cache writes at $0.125 per million; Astra's lists $12.50.

Notice the cutoffs: Luna's is the newest at May 18, 2026, even though it is the cheapest. Price does not track recency, so check the date if your prompt depends on recent facts.

Which one runs a Sume Format?

The Create a run page says the model field is the Agents catalog id for the orchestrating LLM, and omitting it gives the gpt-6-sol default. A request for the retired gpt-5.6-sol runs on gpt-6-sol. It selects the orchestrator only; the Format's tools choose the image, video and audio models.

So the OpenAI price you pay for orchestration is separate from the price of each render. A run's spend is bounded by generation_spend_cap_usd, which defaults to the Format's cap and is limited to $500 platform-wide, per the same page.

How do I choose between them?

Use the cheap tier for high-volume, narrow tasks such as classifying or routing, the middle tier as the default agent, and the top tier when a run fails on reasoning rather than on a media step. Test with your own prompts; this post does not publish benchmarks, because OpenAI's pages I read did not either.

Sume's docs do not publish a per-token markup for these orchestrators, so I do not state one here. Read the receipt for what a run actually cost.

For tool-calling quirks on Chat Completions, see Sol function calling needs reasoning none.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume