Consistent character in 3 steps: anchor, sheet, scenes on Sume
Keep one character across many AI images without a seed: make an anchor, build a reference sheet, then reuse both as input_references. Limits and cost.

On Sume, the way to keep one character across images is to feed earlier results back in as input_references, because there is no seed to pin: seed is in the request schema, but no image model advertises it, so sending it returns 400 unsupported_parameter. Three steps work: make an anchor portrait, make a reference sheet from it, then make every scene with both attached.
A 10-image set (one anchor, one sheet, eight scenes) costs 0.50 USD on Seedream 4.5 and 1.00 USD on Nano Banana 2.1 at catalog base prices.
Step by step
Each step is an ordinary POST /v1/images. The result of one step is a Sume-hosted URL in data[].url, and a public HTTPS URL is exactly what the next step needs as a reference.
- Anchor: text-only prompt for a front-facing portrait with plain light. Generate a few with
nand keep one. - Sheet: send the anchor as
input_references[0]and ask for the same person in a neutral pose from front, side and back, on a plain background. - Scenes: send the anchor and the sheet as two references, then describe only what is new: place, action, light.
- Pin the same model id for all three steps so the reference handling stays the same.
Which models have the room
The number of references decides how many angles you can attach. The ceilings below are the catalog values read on 2026-10-10. Ideogram 4.5 differs: with references, it edits the first image and uses up to 4 more, five in total, so there the anchor must be first.
| Model | Max input_references | Base price per image (USD) |
|---|---|---|
| openai/gpt-image-2.5 | 16 | 0.065875 |
| google/nano-banana-2.1 | 10 | 0.10 |
| bytedance-seed/seedream-4.5 | 10 | 0.05 |
| black-forest-labs/flux.2-pro | 10 | 0.0375 |
| ideogram/ideogram-v4.5 | 5 | 0.075 (default medium quality) |
Where it can go wrong
References guide the model; they do not lock it. Faces and clothing can still drift between scenes, and the Sume docs make no promise of identity preservation. Review every scene against the anchor and regenerate the ones that drift, which costs one image each.
Keep one prompt sentence that states the stable traits, such as hair, outfit and age range, and repeat it unchanged. Changing those words from scene to scene is the most common reason a set looks like different people.
Fan-out has a cost too: sending eight scene requests at once can hit 429 queue_full when the workspace queue is full, so stagger them and back off as the Errors and rate limits page describes.
Choosing the anchor
The anchor decides everything after it, so spend a little there. Generate four candidates in one call with n: 4 (every reference-capable model in the table allows up to 4), look at them at full size, and keep the one whose face and clothes you would be happy to see fifty times. Plain light and a neutral background make it easier for later steps to take the identity and ignore the setting.
Avoid an anchor with strong props, heavy makeup effects or an unusual crop. Whatever is most distinctive in the anchor is what the model tends to copy into every scene, which can be a problem if you wanted the character to change clothes.
The arithmetic
Ten images at the base price: 10 x 0.05 = 0.50 USD on Seedream 4.5, 10 x 0.10 = 1.00 USD on Nano Banana 2.1, and 10 x 0.0375 = 0.375 USD on FLUX.2 Pro. Failed or cancelled generations are not billed, so a retried scene costs once.
These are floors. If you raise resolution or quality where a model lists those fields, read the pricing lines from the endpoint route for the exact figure.
Sources
Related posts
More in Use cases
- Coworking day-pass promo: a price card stacked over a room video
Stack a still price card over a 6-second room video with Timeline compose: Omni Flash 720p, one Nano Banana 2.1 card, 87 cents in total (read 2026-10-10).
- Craft fair booth reel: three 9:16 clips and a music bed for $2.38
A maker can promote a holiday craft fair booth with three vertical Wan 3.0 clips and one original music track for $2.38 at 720p. Costs and a shot plan.
- Crochet Pattern Designer: Turn One Photo Into a Reel for 8 Cents
Animate a photo of your finished piece. Grok Imagine 1.5 on Sume needs exactly one image and costs 8 cents for 6 seconds at 720p; Wan 3.0 takes more control.
- Descript-style delete-a-word editing with Sume STT and video trim
Descript edits video by editing its transcript. Sume has no such editor, but STT word timings, video trim at $0.02 a cut and Timeline can rebuild the result.
Written by Sume