Demand Gen carousels: 2 to 10 matching cards from one reference

Demand Gen carousels take 2 to 10 cards. Image assets run 4:5 or 9:16 at 5 MB. How to batch matching cards from one reference image with the Sume image API.

5 min readSume
All posts

A Demand Gen carousel has 2 to 10 cards, and card images follow the image limits below, with 4:5 and 9:16 as the two portrait ratios. To build one from a single product reference, generate each card as its own image request that passes the same reference image, and ask for the ratio you want on every call.

The numbers here are from Google's image ad spec page. The Sume side comes from the Image API docs.

The image limits

Images are capped at 5 MB, a logo at 150 KB, an ad takes up to 20 images, and carousels hold 2 to 10 cards. The page lists no separate carousel ratio, so pick one ratio from the table and use it on every card.

Demand Gen image asset specs (read 2026-10-03)
RatioMinimumRecommended
1:1300x3001200x1200
1.91:1600x3141200x628
4:5480x600960x1200
9:16600x10671080x1920

One reference, many cards

The Image API accepts input_references, public HTTPS image URLs, on models whose catalog entry allows them. Reference URLs that are localhost, private-network or non-HTTPS are rejected. The aspect_ratio field takes 4:5 and 9:16, and for 4:5 the docs say exact 1080x1350 is a documented post-step rather than the native output. Hold the reference, the style line and the palette constant, and change only the card's role in the prompt: hero, close-up, in use, detail, offer.

The n parameter returns up to 10 images per call, but each model sets a lower ceiling in its catalog entry. Read the n range before relying on it. Sending one request per card instead keeps each prompt specific.

curl -X POST https://api.sume.com/v1/images \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "sume/auto",
    "prompt": "Carousel card 2 of 6: the same bottle in use at a sink, clean daylight",
    "aspect_ratio": "4:5",
    "output_format": "jpeg",
    "input_references": [
      {"type": "image_url", "image_url": {"url": "https://example.com/bottle.png"}}
    ]
  }'

File size and format

Ask for output_format of jpeg or webp when a card risks passing 5 MB. The docs list png, jpeg, webp and svg, and say output_compression is not served in v1, so format is the lever you have. Check each saved file's size before upload.

If a request runs past the 30-second blocking budget it returns 202 with a job envelope, so check the status code and read the result from the job endpoints rather than assuming an image body.

A card plan

Keep every card in one ratio. Mixed ratios inside one carousel are a layout risk that the spec page does not clear.

  • Card 1: the product alone, clean background.
  • Cards 2 to 4: use, detail and scale.
  • Card 5: the offer or proof point, with the text added after generation rather than drawn by the model.
  • Cards 6 and up only if the story needs them; Google's minimum is two.

Keeping cards consistent

Consistency across cards comes from what you hold fixed. Use the same reference image on every call, the same style sentence at the start of every prompt, and the same colour words. Change only the scene and the card's job. If a card drifts, regenerate that card alone rather than the whole set.

Review the cards side by side at phone size, in order, as a viewer would swipe them. A card that looks strong alone can still break the sequence if its lighting or crop differs from its neighbours. Fix those before checking file size.

Store each card's request, response and job ID with the card number, so a replacement can be generated later from the same inputs. The jobs docs describe the job endpoints you can read back.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume