AI comic generator: consistent characters, panel by panel

Make an AI comic one panel at a time: reuse one reference image per character, give each panel its shape, and letter the speech bubbles yourself.

5 min readSume
All posts

To make a comic with an AI image generator, work one panel at a time. Design each character once, send that picture as a reference image with every panel's prompt, repeat the character's description word for word, and give each panel its own shape. Then arrange the panels on the page and letter the speech bubbles yourself in a layout tool.

On Sume, each panel is one POST /v1/images request to a model that edits from reference images. The facts below come from Sume's Image API docs, the request schema in the Sume API reference, and the model catalog that GET /v1/images/models serves, read on 2026-09-28.

How do I make a comic with AI, step by step?

  • Write a panel script: one line per panel with the shot (wide, medium, close-up), who is in it, and what they do. Keep the dialogue in its own column; it goes on the page at the end, not into the prompt.
  • Design each character: generate a few candidates with n and keep one clear, full-body picture per character. Save it, because result URLs are signed, and host it at a public HTTPS URL of your own to use as the reference.
  • Generate each panel with the character picture in input_references, the same style line and character description pasted unchanged, and then the panel's action. End the prompt with "no text, no speech bubbles" and say where the empty space for the balloon goes.
  • Compare every panel with the character picture, and regenerate the ones that drift.
  • Arrange the panels and add balloons, captions, and sound effects in a layout or design tool.
curl -X POST "https://api.sume.com/v1/images" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/nano-banana-pro",
    "prompt": "Comic panel, bold black ink lines, flat colors. Rex: a small green robot with round blue eyes and a red scarf. Medium shot: Rex peeks around a stack of books in a library at night. Empty space top left for a speech balloon. No text, no speech bubbles.",
    "input_references": [
      { "type": "image_url", "image_url": { "url": "https://example.com/rex.png" } }
    ],
    "aspect_ratio": "4:3",
    "n": 2
  }'

How do I keep the characters consistent across panels?

Reuse one approved picture and one fixed description per character in every panel. Sume's image API serves no seed, so the reference picture is what carries a character from panel to panel; consistent character AI image generator explains the technique.

A comic adds more characters and recurring places. Send each character's picture, and a picture of a recurring setting, in the same request, up to the model's reference limit in the table below. A reference carries only its type and URL, with no name or role, so the prompt has to say who is who by appearance ("the robot with the red scarf") and what each one does. Faces, costumes, and props can still drift, so check every panel.

Which panel shapes can I generate?

Set each panel's shape with aspect_ratio, using a value from the model's list; a model accepts only the values its catalog lists. Don't send "auto" for panels: on reference calls it matches the reference, so every panel would copy the character picture's shape. For a one-row strip, Nano Banana 2 lists 4:1, and ChatGPT Image 2.5 takes a custom image_size up to 3:1, such as 3840 × 1280 pixels.

From Image API and the GET /v1/images/models catalog in current code, read 2026-09-28.
Model idPanel shapes (`aspect_ratio`)Reference images per request
openai/gpt-image-2.51:1, 16:9, 9:16, 4:3, 3:4, 5:4, 9:8, 4:5; or a custom image_size up to 3:116
google/nano-banana-221:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16, 4:1, 1:4, 8:1, 1:810
google/nano-banana-pro21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:1610

Can I turn a photo into a comic?

Yes. Send the photo as a reference and name the comic style in the prompt, for example "redraw this photo as a comic panel with bold ink outlines and flat colors; keep the pose and the setting". Add aspect_ratio: "auto", which all three models above list, to keep the photo's shape. The model redraws the people rather than tracing them, so the likeness can change. Use photos of yourself or of people who agreed.

Should the AI write the speech bubbles?

No, letter them yourself. Words drawn into a picture have to be proofread letter by letter and can't be edited afterwards, so ask the model for empty space and set balloons, captions, and sound effects in your layout tool. If a word must be part of the art, such as a shop sign, quote it exactly in the prompt and check it; AI image with text covers that case.

What are the limits?

  • No setting guarantees an identical character from panel to panel.
  • One call returns up to 4 images on these models today. Each completed image is billed, and a failed or cancelled generation is not.
  • Each model's price is in GET /v1/images/models/{model_id}/endpoints. ChatGPT Image 2.5 estimates each request's price from the size, the quality, and the images and text you send, so read usage.cost in the response for the USD amount billed.
  • A request still running after 30 seconds returns 202 with a job to poll instead of the images.
  • Result URLs are Sume-hosted and signed, so download the panels you keep.
  • To make the finished panels move, animate them one clip at a time.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume