GPT-6 Astra ultrafast tier vs Sume's Fast checkbox
OpenAI added service_tier ultrafast for GPT-6 Astra on Sep 29. Sume's picker has Fast but no ultrafast option.

OpenAI's Sep 29 changelog entry launches an ultrafast mode for GPT-6 Astra in the Responses API: set service_tier: "ultrafast" to reduce latency between output tokens, available to API customers subject to rate limits. Sume's agent picker for GPT-6 Astra has a Fast checkbox, but no ultrafast setting, so on Sume you cannot select this tier.
OpenAI's wording comes from the OpenAI API changelog and Sume's from the model catalog source, both read on 2026-10-03. The changelog entry does not give an ultrafast price, so none is stated here.
What does ultrafast change for an agent?
The page describes it as lower latency between output tokens, which is the streaming speed you notice when an agent writes a long shot list or a tool-call argument. It does not claim a lower time to first token, so a thread that stalls waiting on a render is not helped. Video agents spend most of their wall-clock waiting for media jobs, which no token-speed tier shortens.
What does Sume expose for Astra?
In the model catalog the GPT-6 Astra row carries three parameters: Effort, Fast and Thinking. A code comment on the row says OpenAI documents Fast mode at 2x Standard, which is why the Fast checkbox is a real service-tier twin rather than an invented knob. The same comment notes the picker offers Low, Medium and High effort even though OpenAI documents more levels. Astra is never the default and never recommended; it is opt-in.
The catalog source contains no ultrafast tier, so the honest summary is: Standard and Fast are selectable, ultrafast is not.
| Tier | OpenAI changelog | Sume agent picker |
|---|---|---|
| Standard | Default | Selectable |
| Fast | Documented in the catalog code as 2x Standard | Fast checkbox |
| Ultrafast | service_tier ultrafast, Responses API, rate limited, Sep 29 | Not offered |
When would I want it anyway?
If a product streams the agent's words to a user in real time and token speed is the bottleneck, calling the Responses API directly with the new tier is the only route today. If the agent is a producer that runs a Format and returns a finished video, the speed of its own words barely moves total time; the render dominates.
For that producer case, put the effort on cost control instead: the Formats API takes a generation_spend_cap_usd per run and lets you name the orchestrating model with model, as described on Calling a Format. An id outside the catalog is a 400, so you will not get a silent substitute.
Is anything else new for Astra this week?
The same changelog page lists, on Sep 3, that Astra does not support none reasoning, custom temperature or top_p, or logprobs, and that tool calling requires the Responses API. If you were porting a Chat Completions agent, those are the reasons a call can fail before ultrafast even matters.
How would I compare tiers on my own thread?
Measure before you pay. Pick one representative agent turn, run it on Standard, then on whichever faster tier you can reach, and record time from request to final token and the cost on the receipt or usage line. Do it at least five times each, because latency varies with load, and report the median, not the best run.
For Sume runs, the comparison that is available today is Standard against the Fast checkbox on a GPT-6 row. The Format receipt documents the model that ran and the debited spend, so the cost half of the comparison is on the receipt; time you will need to take from your own clock between submit and the terminal status.
Sources
Related posts
More in Models
- Grok Imagine Image 2.0 policy review: cost of a refused image
xAI says generated media gets content policy review. Sume's docs say failed image generations are not billed; here is what is and is not documented.
- Higgsfield Soul on Sume: text to image, 1 or 4 images, 720p or 1080p
Soul is a text-to-image row in Sume's image catalog: seven aspect ratios, no references, 1 or 4 images per call, 720p or 1080p. Limits and price.
- Is Ideogram 4.0 open source? Apache code, non-commercial weights
Ideogram 4.0's inference code is Apache 2.0, but the weights on Hugging Face ship under a non-commercial agreement. Selling the images needs a paid licence.
- Irodori-TTS-v4-Large: Japanese cloning, 120 s reference, terms
Irodori-TTS-v4-Large is a 3.29B Japanese TTS model with emoji style control and Gemma terms. What the card says and how Sume's audio tools fit.
Written by Sume