Text to image API: send a prompt, get image URLs back

A text-to-image API turns a prompt sent over HTTPS into generated images. How to call one: the request, the response, slow jobs, and the cost.

5 min readSume
All posts

A text-to-image API is a web endpoint that takes a text prompt in an HTTPS request and returns generated images. With Sume, you send POST https://api.sume.com/v1/images with a model and a prompt, and your API key as a Bearer token. The call waits up to 30 seconds and answers 200 with links to the finished images, or 202 with a job to poll when they take longer.

The details below come from Sume's Image API, Jobs and results and Authentication docs, read on 2026-09-28. The same call in Python, with the files saved to disk, is in Python image generation API.

How do I call a text-to-image API?

Create a key on the API keys page, keep it in a server-side environment variable, and send one request. The request below is the docs' own cURL example.

  • model picks the model: an id from GET /v1/images/models, or sume/auto to let Sume choose. Sume never discloses which family ran an Auto request.
  • prompt describes the image. model and prompt are the only required fields, and any other generation parameter must be one the model lists, or the call fails with 400 unsupported_parameter. To start from a photo instead, add input_references: that is image to image AI.
  • Send the key as Authorization: Bearer or as x-api-key, never both: a request that carries both is refused with 401 unauthorized.
  • The docs say not to place API keys in frontend JavaScript or mobile apps, so call the API from your backend.
curl -X POST "https://api.sume.com/v1/images" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "bytedance-seed/seedream-4.5",
    "prompt": "a red panda astronaut floating in space, studio lighting"
  }'

What does the API send back?

A finished call returns 200 with Sume-hosted image URLs rather than inline base64, in the shape below: the docs' example, with the billed amount left out.

  • data[].url links are signed, so download the files you keep. media_type gives each file's type.
  • usage.cost is the USD amount billed to your wallet. The token counts are always 0, because image models are metered per image.
  • model echoes the id you sent, so sume/auto stays sume/auto.
{
  "created": 1748372400,
  "model": "bytedance-seed/seedream-4.5",
  "data": [
    { "url": "https://media.sume.com/img/01J.../0.png", "media_type": "image/png" }
  ],
  "usage": { "prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0, "cost": … }
}

What if the image takes longer than 30 seconds?

Then the same call returns 202 with a job envelope instead of images, so check the status code, not the body shape. Poll GET /v1/jobs/{id}/status until the job is completed, failed or canceled, and read GET /v1/jobs/{id}/result only for a completed job; Python image generation API shows that loop. 4K, high quality and a large n are the settings most likely to end in a 202.

To skip the wait, send mode: "async". To be told instead of polling, send mode: "webhook" with a public HTTPS webhook_url: Sume posts the terminal event, job.completed, job.failed or job.canceled, with no progress events, and signs it so you can verify it. The docs say to keep polling as a backup. The whole flow:

From Image API, Jobs and results and Authentication, read 2026-09-28.
StepCallWhat to know
Get a keyAPI keys page in the dashboardWorkspace-scoped; send it as a Bearer token.
Choose a modelGET /v1/images/modelsLists each model's supported_parameters. Or send sume/auto.
GeneratePOST /v1/imagesmodel and prompt required; the call blocks for up to 30 seconds.
Read a finished call200 responseSigned URLs in data[].url; the billed amount in usage.cost.
Read a slow call202 job envelopePoll GET /v1/jobs/{id}/status, then fetch GET /v1/jobs/{id}/result.
Check the priceGET /v1/images/models/{model_id}/endpointsThe pricing line gives cost_usd per output image; usage.cost shows what a call billed.

How much does a text-to-image API cost?

On Sume, image calls are paid per image. A model's endpoints record carries a pricing line with cost_usd per output image, which already includes Sume's margin, plus a 5.5% agent fee by default. On ChatGPT Image 2.5, that per-image price is estimated from tokens, so it moves with each request's size and quality. usage.cost in the response is the USD amount billed, and a failed or cancelled generation is not billed. Model-by-model prices are in AI image generator API cost.

What can't Sume's text-to-image API do?

Two gaps a text-to-image caller meets first; image generation with reference images covers the other fields the Image API doesn't serve yet.

  • Streaming partial images: stream: true returns 400 streaming_not_supported. A seed returns 400 unsupported_parameter.
  • SVG: no model lists svg among its output_format values today, so ask for a format the model lists, such as png, jpeg or webp.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume