One caption per carousel slide: get a typed list from a Sume Format
Write a separate caption for each carousel slide by binding an output_schema to a Sume Format run, so your app gets one typed caption per image.

To write one caption per carousel slide, bind a JSON Schema to a Sume Format run with output_schema, so the run returns a typed list with one caption per image instead of a single paragraph. Your app can then write each caption into its own field. We did not read an Instagram page on per-slide captions, so this post makes no claim about that feature; it covers how to get the data in the right shape.
The behavior below is from the Sume docs on structured output and the Format run call.
Why a schema
By default a completed Format run gives you media and a paragraph of text, which is good for a person and hard to put into a database. If you bind a schema, you get a typed object in the shape you asked for.
| Field | Direction | If the shape is wrong |
|---|---|---|
| input | Caller data, you to the run; any JSON object | The run starts and unknown keys are only more data |
| output_schema | A contract for the receipt, run to you; JSON Schema in the supported subset | 400 output_schema_invalid; nothing runs and nothing is charged |
Design the schema
Keep it flat and small. A useful shape for a carousel:
slides: an array with one object per image, each withindex,alt_textandcaption.opening_line: the line for the post text that introduces the set.disclosure: a short line stating that AI was used, when it was.- Keep only fields you will really use; each one is a place for a run to go wrong.
Pass the slides in as input
Send the slide list as input, for example the image URLs in order plus the product facts, and ask in instruction for a caption of a given length per slide. The projection never sees your input, but the run does, because input is part of the run's prompt as a data block. The receipt says whether the object was filled by the run (filled_by: agent) or by the fallback projection, so log that value and review the projection cases by hand.
Many carousels at once
If you have a catalog, the bulk runs guide covers submitting one input per row. Use the same output_schema for every row so that the results land in one table. Count the number of captions returned against the number of slides you sent, and flag any row that does not match before it goes near a post.
What stays with you
Sume returns the text. Posting it, choosing the platform's character limit and checking the platform's rules are your steps, and we did not verify any Instagram limit. If the captions are for video, the related path is burned captions: see creator caption looks from one transcription. For disclosure wording, see the disclosure checklist by platform.
Sources
Related posts
More in Use cases
- Change one of three identical bottles: position words or a mask
Edit one of three identical products in a photo. Flux 3 Image gives each element an id; on Sume, use position words (Ideogram 4.5) or a GPT Image 2.5 mask.
- Change prices on a chalkboard menu photo with Ideogram 4.5
Update dish prices on a chalkboard menu photo with ideogram/ideogram-v4.5: list old and new text, one change set per call, check each digit. $0.075 at medium.
- Change the background behind a person in a video with AI (Omni edit)
Swap the backdrop behind a person in an existing clip with Gemini Omni Flash 1.1 edit on Sume: a keep-the-subject prompt, limits, and a 10-second price table.
- Change the name and number on a jersey photo, keep the lettering
Edit the name and number on a jersey or sign photo with Ideogram 4.5 on Sume: the prompt wording, the reference rules and what a low-tier test costs.
Written by Sume