Consistent character in 3 steps: anchor, sheet, scenes on Sume

Keep one character across many AI images without a seed: make an anchor, build a reference sheet, then reuse both as input_references. Limits and cost.

5 min readSume
All posts

On Sume, the way to keep one character across images is to feed earlier results back in as input_references, because there is no seed to pin: seed is in the request schema, but no image model advertises it, so sending it returns 400 unsupported_parameter. Three steps work: make an anchor portrait, make a reference sheet from it, then make every scene with both attached.

A 10-image set (one anchor, one sheet, eight scenes) costs 0.50 USD on Seedream 4.5 and 1.00 USD on Nano Banana 2.1 at catalog base prices.

Step by step

Each step is an ordinary POST /v1/images. The result of one step is a Sume-hosted URL in data[].url, and a public HTTPS URL is exactly what the next step needs as a reference.

  • Anchor: text-only prompt for a front-facing portrait with plain light. Generate a few with n and keep one.
  • Sheet: send the anchor as input_references[0] and ask for the same person in a neutral pose from front, side and back, on a plain background.
  • Scenes: send the anchor and the sheet as two references, then describe only what is new: place, action, light.
  • Pin the same model id for all three steps so the reference handling stays the same.

Which models have the room

The number of references decides how many angles you can attach. The ceilings below are the catalog values read on 2026-10-10. Ideogram 4.5 differs: with references, it edits the first image and uses up to 4 more, five in total, so there the anchor must be first.

Reference ceilings for character work in the Sume image catalog (read 2026-10-10)
ModelMax input_referencesBase price per image (USD)
openai/gpt-image-2.5160.065875
google/nano-banana-2.1100.10
bytedance-seed/seedream-4.5100.05
black-forest-labs/flux.2-pro100.0375
ideogram/ideogram-v4.550.075 (default medium quality)

Where it can go wrong

References guide the model; they do not lock it. Faces and clothing can still drift between scenes, and the Sume docs make no promise of identity preservation. Review every scene against the anchor and regenerate the ones that drift, which costs one image each.

Keep one prompt sentence that states the stable traits, such as hair, outfit and age range, and repeat it unchanged. Changing those words from scene to scene is the most common reason a set looks like different people.

Fan-out has a cost too: sending eight scene requests at once can hit 429 queue_full when the workspace queue is full, so stagger them and back off as the Errors and rate limits page describes.

Choosing the anchor

The anchor decides everything after it, so spend a little there. Generate four candidates in one call with n: 4 (every reference-capable model in the table allows up to 4), look at them at full size, and keep the one whose face and clothes you would be happy to see fifty times. Plain light and a neutral background make it easier for later steps to take the identity and ignore the setting.

Avoid an anchor with strong props, heavy makeup effects or an unusual crop. Whatever is most distinctive in the anchor is what the model tends to copy into every scene, which can be a problem if you wanted the character to change clothes.

The arithmetic

Ten images at the base price: 10 x 0.05 = 0.50 USD on Seedream 4.5, 10 x 0.10 = 1.00 USD on Nano Banana 2.1, and 10 x 0.0375 = 0.375 USD on FLUX.2 Pro. Failed or cancelled generations are not billed, so a retried scene costs once.

These are floors. If you raise resolution or quality where a model lists those fields, read the pricing lines from the endpoint route for the exact figure.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume