Text to image API: send a prompt, get image URLs back
A text-to-image API turns a prompt sent over HTTPS into generated images. How to call one: the request, the response, slow jobs, and the cost.

A text-to-image API is a web endpoint that takes a text prompt in an HTTPS request and returns generated images. With Sume, you send POST https://api.sume.com/v1/images with a model and a prompt, and your API key as a Bearer token. The call waits up to 30 seconds and answers 200 with links to the finished images, or 202 with a job to poll when they take longer.
The details below come from Sume's Image API, Jobs and results and Authentication docs, read on 2026-09-28. The same call in Python, with the files saved to disk, is in Python image generation API.
How do I call a text-to-image API?
Create a key on the API keys page, keep it in a server-side environment variable, and send one request. The request below is the docs' own cURL example.
modelpicks the model: an id fromGET /v1/images/models, orsume/autoto let Sume choose. Sume never discloses which family ran an Auto request.promptdescribes the image.modelandpromptare the only required fields, and any other generation parameter must be one the model lists, or the call fails with400 unsupported_parameter. To start from a photo instead, addinput_references: that is image to image AI.- Send the key as
Authorization: Beareror asx-api-key, never both: a request that carries both is refused with401 unauthorized. - The docs say not to place API keys in frontend JavaScript or mobile apps, so call the API from your backend.
curl -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bytedance-seed/seedream-4.5",
"prompt": "a red panda astronaut floating in space, studio lighting"
}'What does the API send back?
A finished call returns 200 with Sume-hosted image URLs rather than inline base64, in the shape below: the docs' example, with the billed amount left out.
data[].urllinks are signed, so download the files you keep.media_typegives each file's type.usage.costis the USD amount billed to your wallet. The token counts are always0, because image models are metered per image.modelechoes the id you sent, sosume/autostayssume/auto.
{
"created": 1748372400,
"model": "bytedance-seed/seedream-4.5",
"data": [
{ "url": "https://media.sume.com/img/01J.../0.png", "media_type": "image/png" }
],
"usage": { "prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0, "cost": … }
}What if the image takes longer than 30 seconds?
Then the same call returns 202 with a job envelope instead of images, so check the status code, not the body shape. Poll GET /v1/jobs/{id}/status until the job is completed, failed or canceled, and read GET /v1/jobs/{id}/result only for a completed job; Python image generation API shows that loop. 4K, high quality and a large n are the settings most likely to end in a 202.
To skip the wait, send mode: "async". To be told instead of polling, send mode: "webhook" with a public HTTPS webhook_url: Sume posts the terminal event, job.completed, job.failed or job.canceled, with no progress events, and signs it so you can verify it. The docs say to keep polling as a backup. The whole flow:
| Step | Call | What to know |
|---|---|---|
| Get a key | API keys page in the dashboard | Workspace-scoped; send it as a Bearer token. |
| Choose a model | GET /v1/images/models | Lists each model's supported_parameters. Or send sume/auto. |
| Generate | POST /v1/images | model and prompt required; the call blocks for up to 30 seconds. |
| Read a finished call | 200 response | Signed URLs in data[].url; the billed amount in usage.cost. |
| Read a slow call | 202 job envelope | Poll GET /v1/jobs/{id}/status, then fetch GET /v1/jobs/{id}/result. |
| Check the price | GET /v1/images/models/{model_id}/endpoints | The pricing line gives cost_usd per output image; usage.cost shows what a call billed. |
How much does a text-to-image API cost?
On Sume, image calls are paid per image. A model's endpoints record carries a pricing line with cost_usd per output image, which already includes Sume's margin, plus a 5.5% agent fee by default. On ChatGPT Image 2.5, that per-image price is estimated from tokens, so it moves with each request's size and quality. usage.cost in the response is the USD amount billed, and a failed or cancelled generation is not billed. Model-by-model prices are in AI image generator API cost.
What can't Sume's text-to-image API do?
Two gaps a text-to-image caller meets first; image generation with reference images covers the other fields the Image API doesn't serve yet.
- Streaming partial images:
stream: truereturns400 streaming_not_supported. Aseedreturns400 unsupported_parameter. - SVG: no model lists
svgamong itsoutput_formatvalues today, so ask for a format the model lists, such aspng,jpegorwebp.
Sources
Related posts
More in Developers
- Text to speech API in Java: pick a voice, save the MP3
Call a text to speech API from Java: choose a voice selector and an audio format, POST the text, wait for the job, then write the MP3 to disk.
- Text to video and image to video: how they differ
Text-to-video invents every frame from words; image-to-video starts on your picture and animates it. How they differ, and how one request does both.
- Text to video API: how it works, what it returns, cost
A text to video API takes a prompt and returns a job id, not a video: poll it or take a webhook, then download the file. Fields, flow and prices.
- Webhook security best practices: a receiver checklist
Accept HTTPS only, verify an HMAC over the raw body in constant time, reject stale timestamps, dedupe on the event id, and answer 2xx fast.
Written by Sume