Luma Photon API: 4 image refs and a character ref vs Sume references
Luma's photon-1 image API takes up to 4 image references, a style reference and a 4-image character reference, each with a weight. Here is the Sume equivalent.

Luma's earlier image API, photon-1 and photon-flash-1, takes up to four image references with a weight, a style reference with a weight, and a character reference built from up to four images of the same person. Sume's Image API has one reference field, input_references, and no separate style or character slot, so the roles are expressed in the prompt and by which references you send.
If you built on Photon and are mapping it to Sume, the table below lines the fields up.
What does the Photon API accept?
Two models are listed: photon-1, the default, and photon-flash-1. Aspect ratios are 1:1, 3:4, 4:3, 9:16, 16:9 (default), 9:21 and 21:9. The guide describes four ways to guide a generation.
- Image reference: up to 4 images, with a weight that sets how much each influences the result.
- Style reference: applies a specific style, with its own weight.
- Character reference: builds a consistent character from up to 4 images of the same person.
- Modify image: changes an existing image by prompt; the page says colour changes work best at low weights (0.0 to 0.1) and object or shape changes at higher weights.
What do the weights and uploads look like?
Weight runs from 0.0 to 1.0 and controls influence. A callback_url receives status updates and results by POST. The guide says to upload your own CDN image URLs, and that this is currently the only way to pass an image to the API. Luma's reframe guide also lists a 10 MB maximum image for photon-1 and photon-flash-1.
The practical point is that Photon separates what an image is for (content, style, identity) into three slots, and you tune each with a number. That is convenient when you want identity strongly held and style loosely held in the same call.
How do those map to Sume?
Sume's image request takes input_references, an array of public HTTPS image URLs, and each model's catalog entry lists how many it accepts. A model whose descriptor is zero to zero is text-to-image only and rejects references. ChatGPT Image 2.5 accepts up to 16 references and an optional mask_url. A single call can ask for up to 10 images with n, though per-model ceilings are lower.
| Photon feature | Photon limit | On Sume |
|---|---|---|
| Image reference | 4 images, weight 0.0-1.0 | input_references; count set per model in the catalog; no weight field |
| Style reference | 1 style with weight | No style slot; describe the style or send a style image as a reference and say so |
| Character reference | 4 images of one person | No character slot; send the images as references and name the subject in the prompt |
| Callback | callback_url POST | webhook_url on the job; job statuses are polled or pushed |
| Aspect ratios | 7 ratios from 9:21 to 21:9 | Normalized list incl. 1:1, 16:9, 9:16, 4:5, 9:21, 21:9, or auto |
How do you request identity and style on Sume?
State the role of each reference in the prompt, because there is no weight to tune. For example: Image 1 is the person, keep the face and hair identical. Image 2 is the lighting and color mood only. On edit calls use aspect_ratio auto to match the first reference, because leaving the field out is not the same as auto.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-image-2.5","prompt":"Image 1 is the person: keep face and hair identical. Image 2 is the color mood only. Portrait at a cafe window.","input_references":[{"type":"image_url","image_url":{"url":"https://example.com/person.jpg"}},{"type":"image_url","image_url":{"url":"https://example.com/mood.jpg"}}],"aspect_ratio":"auto"}'What do you lose and gain?
You lose numeric weights and the dedicated character slot, so you cannot dial identity up while holding style down by a number. Check output across several runs when likeness matters, and keep the same reference set and prompt wording for each shot in a series.
You gain a wider reference ceiling on some models, a mask field on others, and the same job envelope as Sume's video and audio routes, so one client handles every media type. Read supported_parameters on the catalog entry for the model you pin; the Image API page explains the descriptors.
Sources
Related posts
More in Models
- Lyria 3.5 is single-turn and varies per call: keep the artifact
Google says Lyria generation is single-turn, not iteratively editable and varies between calls. Why to save every Sume Music artifact you like.
- Lyria 3.5 in another language: prompt in it, then check the song
Google says Lyria 3.5 makes music in other languages when you prompt in that language. On Sume, write the brief and lyrics in it, then check the result.
- MAI-Image-2.6 output cap is 2,359,296 pixels; Sume sizes differ
MAI-Image-2.6 sets a 2,359,296-pixel ceiling and a 768-pixel minimum edge. Sume sets size per model with tiers, ratios and, for GPT models, custom pixels.
- MAI-Image-2.6 edits take 5 references; Sume takes 10 or 16
MAI-Image-2.6 in Foundry accepts up to five JPEG or PNG reference images per edit. On Sume, input_references tops out at 10, or 16 on GPT Image 2.5.
Written by Sume