GPT-6 Astra ultrafast tier vs Sume's Fast checkbox

OpenAI added service_tier ultrafast for GPT-6 Astra on Sep 29. Sume's picker has Fast but no ultrafast option.

4 min readSume
All posts

OpenAI's Sep 29 changelog entry launches an ultrafast mode for GPT-6 Astra in the Responses API: set service_tier: "ultrafast" to reduce latency between output tokens, available to API customers subject to rate limits. Sume's agent picker for GPT-6 Astra has a Fast checkbox, but no ultrafast setting, so on Sume you cannot select this tier.

OpenAI's wording comes from the OpenAI API changelog and Sume's from the model catalog source, both read on 2026-10-03. The changelog entry does not give an ultrafast price, so none is stated here.

What does ultrafast change for an agent?

The page describes it as lower latency between output tokens, which is the streaming speed you notice when an agent writes a long shot list or a tool-call argument. It does not claim a lower time to first token, so a thread that stalls waiting on a render is not helped. Video agents spend most of their wall-clock waiting for media jobs, which no token-speed tier shortens.

What does Sume expose for Astra?

In the model catalog the GPT-6 Astra row carries three parameters: Effort, Fast and Thinking. A code comment on the row says OpenAI documents Fast mode at 2x Standard, which is why the Fast checkbox is a real service-tier twin rather than an invented knob. The same comment notes the picker offers Low, Medium and High effort even though OpenAI documents more levels. Astra is never the default and never recommended; it is opt-in.

The catalog source contains no ultrafast tier, so the honest summary is: Standard and Fast are selectable, ultrafast is not.

GPT-6 Astra tiers: OpenAI changelog vs Sume picker (read 2026-10-03)
TierOpenAI changelogSume agent picker
StandardDefaultSelectable
FastDocumented in the catalog code as 2x StandardFast checkbox
Ultrafastservice_tier ultrafast, Responses API, rate limited, Sep 29Not offered

When would I want it anyway?

If a product streams the agent's words to a user in real time and token speed is the bottleneck, calling the Responses API directly with the new tier is the only route today. If the agent is a producer that runs a Format and returns a finished video, the speed of its own words barely moves total time; the render dominates.

For that producer case, put the effort on cost control instead: the Formats API takes a generation_spend_cap_usd per run and lets you name the orchestrating model with model, as described on Calling a Format. An id outside the catalog is a 400, so you will not get a silent substitute.

Is anything else new for Astra this week?

The same changelog page lists, on Sep 3, that Astra does not support none reasoning, custom temperature or top_p, or logprobs, and that tool calling requires the Responses API. If you were porting a Chat Completions agent, those are the reasons a call can fail before ultrafast even matters.

How would I compare tiers on my own thread?

Measure before you pay. Pick one representative agent turn, run it on Standard, then on whichever faster tier you can reach, and record time from request to final token and the cost on the receipt or usage line. Do it at least five times each, because latency varies with load, and report the median, not the best run.

For Sume runs, the comparison that is available today is Standard against the Fast checkbox on a GPT-6 row. The Format receipt documents the model that ran and the debited spend, so the cost half of the comparison is on the receipt; time you will need to take from your own clock between submit and the terminal status.

Sources

Related posts

More in Models

All Models posts

Written by Sume