Which Sume image models make 2K or 4K output, by model

FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.

5 min readSume
All posts

On Sume, two kinds of image model give you 2K or 4K. Nano Banana 2 and Nano Banana Pro publish a resolution tier of 512, 1K, 2K, or 4K, and Imagen 4 Ultra publishes 1K or 2K. ChatGPT Image 2.5 has no tier but takes custom pixels through image_size, up to a 3840-pixel edge. Most other rows publish neither, so asking them for "4K" returns 400 unsupported_parameter and the output size is whatever the model picks.

FLUX 3 Image, launched Oct 1, lists 4K as a top resolution tier, but it is not in Sume's catalog today. Here is what the models that are in it can do.

Which models publish a resolution tier?

Sume's docs say a model only accepts the values its catalog descriptors list, so the resolution enum on each row is the authority. In the code that builds the catalog, only three kinds of row have one: the two Nano Banana models, Imagen 4 Ultra, and Higgsfield Soul (720p or 1080p). Every other row has no resolution field at all.

The public value for the smallest tier is 512, and the catalog code treats the legacy 0.5K spelling as the same tier. The default is 1K for Nano Banana and 2K for Imagen 4 Ultra.

Which models take custom pixels instead?

The ChatGPT Image 2.5 rows document an image_size rule: both edges are multiples of 16, the maximum edge is 3840, the aspect ratio is at most 3:1, and the total is 655,360 to 8,294,400 pixels. 8,294,400 is 3840 by 2160, so a 4K frame is the ceiling. The Image 1.0 page says custom pixels are accepted on GPT, Seedream, Flux, Qwen, and Recraft; it publishes a pixel rule only for GPT, so I cannot give you a FLUX.2 ceiling, and I will not guess one.

Large-output options in Sume's image catalog, read 2026-10-03
Model idHow to ask for sizePublished range
google/nano-banana-2resolution512, 1K, 2K, 4K
google/nano-banana-proresolution512, 1K, 2K, 4K
google/imagen-4-ultraresolution1K, 2K (text to image only)
openai/gpt-image-2.5image_size in pixelsEdges multiple of 16, max 3840, 655,360 to 8,294,400 px
black-forest-labs/flux.2-proaspect_ratio; custom pixels per Image 1.0 pageNo ceiling published
higgsfield/soulresolution720p, 1080p

What happens when you ask for 4K?

Sume's docs say POST /v1/images waits up to 30 seconds and then returns a 202 job, and that slow configurations such as 4K, high quality, and a large n are the most likely to do so. Your code should treat 200 and 202 as two normal outcomes. A failed generation is not billed. For a 4K Nano Banana call the docs add that admission scales above the default list price, so check pricing on the endpoint record before you loop.

Press coverage of FLUX 3 Image (Tech Times, reported) puts its native output near 5,456 by 3,072 pixels, about 16.8 megapixels, against a 4-megapixel ceiling for FLUX.2. Sume's catalog ceilings are lower, and for the 4K tier I could not confirm exact pixel dimensions, so measure the files you get back.

Does writing 4K in the prompt do anything?

No. Words like "4K, ultra detailed" in a prompt change style cues, not the pixel count; size comes from the resolution tier, from image_size, or from the model's own default. The same holds for FLUX 3 Image: the Replicate page exposes size as a resolution input of 768sq to 4k, not as prompt text. The AI video 4K prompt note makes the same point for video.

Aspect ratio and tier combine. A 4K tier at 16:9 and a 4K tier at 1:1 are different pixel counts, and models pick the long edge differently, so pin both fields and read the returned width and height. On edit calls Sume's docs recommend aspect_ratio: "auto" so the result follows the reference, and note that leaving the field out is not the same as auto.

Which of these also take reference images?

Size and editing are separate questions. Nano Banana 2 and Pro accept up to 10 input_references alongside a tier, so a 4K edit from a reference is one call. Imagen 4 Ultra is text-to-image only and rejects references. ChatGPT Image 2.5 takes up to 16 references plus image_size, so it can edit toward a large frame, within the pixel ceiling above. The live-catalog script below prints the reference maximum next to each tier list so you do not have to remember which is which.

List the tiers from the live catalog

Catalog rows change, so read them instead of copying a table. This script prints every model that advertises a resolution tier, plus whether it accepts reference images.

import os
import requests

r = requests.get(
    "https://api.sume.com/v1/images/models",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
    p = m["supported_parameters"]
    if "resolution" in p:
        refs = p.get("input_references", {}).get("max", 0)
        print(m["id"], p["resolution"]["values"], "refs:", refs)

When should you upscale instead?

If the model you like has no tier, generate at its native size and upscale after. Sume's upscaler is POST /v1/image-upscale-1.0/upscale, and the edit first, upscale after note explains why that order keeps detail. The 4K generate or upscale post compares cost, and draft at 1K, finish at 2K shows how to read the descriptor for a two-step workflow.

Sources

Related posts

More in Models

All Models posts

Written by Sume