Can GPT-6 Sol, Luna or Astra generate images or video?
No. OpenAI lists GPT-6 Astra, Sol and Luna as text-output models with image input. Where image and video generation live instead, and how Sume splits it.

No. OpenAI's model pages for GPT-6 Astra, GPT-6 Sol and GPT-6 Luna, all read on 2026-10-02, list the input as text and image and the output as text only. None of the three generates images, video or audio. Images come from the GPT Image 2.5 models, and OpenAI says there is no one-to-one replacement for the shut-down Sora video API.
What do the three model pages say?
Each page lists modalities, endpoints and limits. The Luna page states image generation is not a supported endpoint, and the Astra page says most audio and image endpoints are unsupported.
| Model | Input | Output | Endpoints listed |
|---|---|---|---|
| gpt-6-astra | Text, image | Text | Chat Completions, Responses, Batch |
| gpt-6-sol | Text, image | Text | Chat Completions, Responses, Batch |
| gpt-6-luna | Text, image | Text | Chat Completions, Responses, Batch |
Where do OpenAI's images come from?
From the GPT Image family. The image generation guide lists gpt-image-2.5-sunburst and gpt-image-2.5-flare, with generation, masked editing and PNG, JPEG or WebP output. The Flare page lists the v1/images/generations and v1/images/edits endpoints plus Batch.
On the video side, OpenAI's video generation guide states that the Sora 2 models and the Videos API were shut down on September 24, 2026, with no one-to-one replacement API available. So a GPT-6 model cannot make a clip and neither can the retired Sora endpoints.
How does Sume split the work?
Sume separates the model that thinks from the models that render. In the Formats run request, model is the Agents catalog id of the LLM that orchestrates the run, and the default is gpt-6-sol. The docs add that it selects the orchestrator only, and that image, video and audio models are chosen by the Format's tools.
That matches OpenAI's own split. A GPT-6 model plans, writes and calls tools. Rendering happens in separate media models you can see in the Image API and video catalog, each billed as its own job.
What does this mean when I pick a model?
Do not choose between GPT-6 Sol and Luna to change image quality; neither renders pixels. Choose them for planning, tool use and cost of the orchestration, and choose the media model for how the picture or clip looks. An id outside the Agents catalog is a 400 invalid_request, per the same docs.
Sume does not claim GPT-6 can produce media, and you should be wary of any wrapper that implies it. If a tool called by the orchestrator generates an image, the image is made by an image model, and the receipt echoes the orchestrator id that ran.
For the longer story on how long video work interacts with GPT-6 tool calls, read GPT-6 Astra async tool calls and video jobs.
Sources
Related posts
More in Models
- Does H3 Max Recast keep the original audio? What to check on Sume
fal says H3 Max Recast preserves the source audio. Sume rejects generate_audio and audio references on it. Confirm a result has sound with video inspect.
- Does sume/auto pick Seedance? Sume does not say which model ran
Sume says sume/auto echoes sume/auto in the response and never discloses the family that served the request. To get Seedance, pin the id.
- FLUX.2 takes 8 references by API, 10 in Playground: Sume's range
Black Forest Labs says FLUX.2 pro, max and flex take up to 8 references by API (10 in its Playground), klein 4. Sume's input_references range is read per model.
- flux-2-pro-preview vs flux-2-pro: which id does Sume send?
BFL has flux-2-pro-preview (latest) and flux-2-pro (fixed snapshot). Sume lists black-forest-labs/flux.2-pro and does not let you pick the BFL endpoint.
Written by Sume