Can GPT-6 Sol, Luna or Astra generate images or video?

No. OpenAI lists GPT-6 Astra, Sol and Luna as text-output models with image input. Where image and video generation live instead, and how Sume splits it.

5 min readSume
All posts

No. OpenAI's model pages for GPT-6 Astra, GPT-6 Sol and GPT-6 Luna, all read on 2026-10-02, list the input as text and image and the output as text only. None of the three generates images, video or audio. Images come from the GPT Image 2.5 models, and OpenAI says there is no one-to-one replacement for the shut-down Sora video API.

What do the three model pages say?

Each page lists modalities, endpoints and limits. The Luna page states image generation is not a supported endpoint, and the Astra page says most audio and image endpoints are unsupported.

GPT-6 modalities and endpoints from OpenAI model pages (read 2026-10-02)
ModelInputOutputEndpoints listed
gpt-6-astraText, imageTextChat Completions, Responses, Batch
gpt-6-solText, imageTextChat Completions, Responses, Batch
gpt-6-lunaText, imageTextChat Completions, Responses, Batch

Where do OpenAI's images come from?

From the GPT Image family. The image generation guide lists gpt-image-2.5-sunburst and gpt-image-2.5-flare, with generation, masked editing and PNG, JPEG or WebP output. The Flare page lists the v1/images/generations and v1/images/edits endpoints plus Batch.

On the video side, OpenAI's video generation guide states that the Sora 2 models and the Videos API were shut down on September 24, 2026, with no one-to-one replacement API available. So a GPT-6 model cannot make a clip and neither can the retired Sora endpoints.

How does Sume split the work?

Sume separates the model that thinks from the models that render. In the Formats run request, model is the Agents catalog id of the LLM that orchestrates the run, and the default is gpt-6-sol. The docs add that it selects the orchestrator only, and that image, video and audio models are chosen by the Format's tools.

That matches OpenAI's own split. A GPT-6 model plans, writes and calls tools. Rendering happens in separate media models you can see in the Image API and video catalog, each billed as its own job.

What does this mean when I pick a model?

Do not choose between GPT-6 Sol and Luna to change image quality; neither renders pixels. Choose them for planning, tool use and cost of the orchestration, and choose the media model for how the picture or clip looks. An id outside the Agents catalog is a 400 invalid_request, per the same docs.

Sume does not claim GPT-6 can produce media, and you should be wary of any wrapper that implies it. If a tool called by the orchestrator generates an image, the image is made by an image model, and the receipt echoes the orchestrator id that ran.

For the longer story on how long video work interacts with GPT-6 tool calls, read GPT-6 Astra async tool calls and video jobs.

Sources

Related posts

More in Models

All Models posts

Written by Sume