AI chibi figurine from a photo: one reference edit, four takes

Turn a portrait into a glossy chibi figurine render with one reference edit on gpt-image-2.5, four takes per call, and a plain background for the cutout.

5 min readSume
All posts

A chibi figurine from a photo is one reference edit: openai/gpt-image-2.5, the portrait in input_references, a prompt that names the proportions and the finish, and n: 4 so you can choose the best of four. Ask for an opaque, plain backdrop for the first pass and make the transparent cutout in a second step.

Four takes per call is a request, not a promise: Sume's Image API lists an n range per model, up to 10 per call at most, and a request above the model's ceiling is rejected. Read the range from GET /v1/images/models before you set it.

Prompt shape

Name what must carry over from the photo (hair colour, glasses, clothing colours) and what changes (head-to-body ratio, material). Keep the pose simple; hands and held objects are where figurine renders usually need a retake.

  • Carry over: hair, glasses, outfit colours from the reference.
  • Change: oversized head, small body, glossy vinyl finish.
  • Scene: a plain light-grey seamless backdrop, soft shadow under the feet.
  • Leave out: any text on the base, which you can add in code.

Cutout in a second pass

For a sticker or a shop listing you want transparency. gpt-image-2.5 accepts background: auto|transparent|opaque, so once you have picked a take you can re-run the same prompt with background: "transparent" and the picked image as the reference. Keep it as a second call: the first call is for choosing, the second is for the asset.

{
  "model": "openai/gpt-image-2.5",
  "prompt": "Same chibi vinyl figurine as the reference, same pose and colours, nothing behind it",
  "input_references": [
    {"type": "image_url", "image_url": {"url": "https://example.com/picked-take.png"}}
  ],
  "background": "transparent",
  "aspect_ratio": "auto",
  "output_format": "png"
}

What the two calls use

Two-pass figurine workflow on the Image API (read 2026-10-04)
PassGoalKey fields
1Choose a takeinput_references, aspect_ratio auto, n up to the model range
2Asset with no backdropbackground transparent, output_format png, picked take as reference

Hosting the picked take

Results come back as Sume-hosted URLs, and the Media inputs doc says signed or private URLs are rejected as inputs. Copy the picked take to a public HTTPS location of your own before you use it as the second reference. Check the status code on each call: 200 is the image response and 202 is the job envelope you poll through Jobs and results.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume