Per-slide carousel captions: pair each image job with its caption

Instagram added per-slide carousel captions. Keep each slide, its caption and its Sume image job in one manifest so a retry never mismatches. Script included.

4 min readSume
All posts

Per-slide carousel captions turn a post from one block of text into a sequence where each image has its own line, which makes the slide-to-text pairing the thing that can go wrong. The fix is boring and effective: keep one manifest row per slide that holds the image prompt, the caption, an idempotency key and, once submitted, the Sume job id. I could not retrieve an Instagram help or newsroom page describing the feature, so this page makes no claim about per-slide limits or settings; check Instagram's own current help before you plan around a number.

What Sume documents is the image side: Image 1.0 generates 4:5 portrait slides (1080x1350) from a prompt, with num_images 1-4 per call.

One manifest, one key per slide

A retry after a client timeout must not queue a second job, so reuse the same Idempotency-Key only for the same operation and payload, as the docs say. Deriving the key from the slide number and a hash of the prompt gives you that for free, and if you edit a prompt the key changes, so you get a new image rather than a stale replay.

import json, os, urllib.request

API = "https://api.sume.com/v1"

def post(path, body, key):
    req = urllib.request.Request(
        f"{API}{path}",
        data=json.dumps(body).encode(),
        headers={
            "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
            "Content-Type": "application/json",
            "Idempotency-Key": key,
        },
        method="POST",
    )
    with urllib.request.urlopen(req) as res:
        return json.load(res)

import hashlib

slides = [
    {"prompt": "Flat vector steps diagram: cut the clip at 3 seconds", "caption": "Step 1: cut the hook"},
    {"prompt": "Flat vector steps diagram: burn captions on the clip", "caption": "Step 2: burn captions"},
]

manifest = []
for number, slide in enumerate(slides, start=1):
    digest = hashlib.sha256(slide["prompt"].encode()).hexdigest()[:8]
    job = post(
        "/image-1.0/generate",
        {"prompt": slide["prompt"], "aspect_ratio": "4:5", "quality": "medium", "num_images": 1},
        f"slide-{number}-{digest}",
    )
    manifest.append({"slide": number, "caption": slide["caption"], "job": job["request_id"]})

print(json.dumps(manifest, indent=2))

Poll, then check order

Each submit returns a job; read GET /v1/jobs/:id/result once it is terminal and take the image from result.artifacts[] where type is image. The URL is a Sume-hosted media.sume.com artifact. Sort by the manifest's slide number rather than by completion time, because jobs finish in any order.

Manifest fields to store per slide, read 2026-10-03
FieldWhy
slideFixes the order regardless of job completion
captionThe text you will paste for that slide
promptSo you can regenerate the same slide
idempotency_keyMakes a retry safe
jobThe request_id to poll
artifact_urlThe finished 4:5 image

What stays outside Sume

Sume does not post to Instagram, and the docs I read do not describe writing a per-slide caption anywhere; you paste or publish captions yourself. Music is a separate constraint: Instagram's help page says music cannot be added to carousels with videos, so keep slides as images if you want a soundtrack (read 2026-10-03). With the manifest in hand, review the images beside their captions in one pass before you publish.

Writing captions that stand alone

When each slide carries its own caption, a viewer may read one slide's text without the others, so write each caption to make sense by itself. Name the subject, give the one claim that slide makes, and avoid pronouns that point back to the previous slide. If a slide is a step in a sequence, put the step number in the caption as you did in the manifest.

Keep the image prompt and the caption describing the same thing. A mismatch is the common error with batch-generated slides: the picture shows one step and the caption names another because the lists were edited separately. Storing them in one row removes that failure.

Cost sketch for ten slides

Prices vary by model, so read the catalog before a run. As one data point, the Image API doc lists Ideogram 4.5 at fal list prices of $0.03, $0.06 or $0.22 per image by quality (Sume bills on top of list), so ten slides at medium list at $0.60 before Sume's margin. Use GET /v1/images/models for the live number and set quality low for drafts, then regenerate only the slides you keep at a higher setting.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume