A/B test two hooks in one bulk queue: instruction beats the recipe
Run two opening hooks on one product by sending two items to the same Sume Format. Your instruction is placed after the recipe, so it wins where they disagree.

Send two items to one bulk queue on the same Format: identical input, and an instruction that differs only in the hook line. Sume places your instruction after the Format body, so where the two disagree the run does what you asked. That makes the hook the one variable.
How the run is composed
The agent receives, in order: a pointer at the Format recipe, the attached package, your instruction (or the Format's default), an unattended-run note, a pointer at your input, and any attachments. Your instruction is limited to 8000 characters, of which roughly the first 4000 are carried as prompt text, so keep it short and put data in input.
{
"concurrency": 2,
"items": [
{ "instruction": "Open on the price drop.",
"input": { "sku": "8823", "variant": "A" } },
{ "instruction": "Open on a customer question.",
"input": { "sku": "8823", "variant": "B" } }
]
}Keeping the test fair
Two runs of a generative recipe will differ in many ways beyond the hook: images, timing, voice take. One pair can show a difference that is only noise, so repeat each variant several times before you conclude anything.
| Hold constant | How |
|---|---|
| Format version | Check format.version on both receipts |
| Orchestrator model | Set the same model field, or leave both on the default |
| Input data | Same input object except the variant label |
| Spend cap | Same generation_spend_cap_usd on both items |
Reading the results
Each child has its own receipt with output, artifacts[] and usage. Pull the two primary_output_url values and review them side by side, and compare usage.billable_amount_usd_micros so a variant that costs more is visible.
For the real answer, publish both and let the ad platform's own reporting decide, with enough impressions to mean something. The queue only produces the two candidates; it does not measure them.
If the recipe itself fights your hook, the instruction still wins where they disagree, but a long list of overrides is a sign that the recipe should change. Edit the Format once instead of repeating the same correction in every item.
Keep the comparison fair by changing only the hook between the two variants, and label each item in your own records so you can tell which result came from which hook once the batch finishes.
Tradeoff
A bulk queue is a good way to start both variants together, but the queue does not compare them for you. You read each child receipt and judge the outputs yourself, and results from the ad platform decide the winner. Never claim a hook won from one render each.
Sources
Related posts
More in Formats
- AI video generator for advertising: check a Format's sample first
GET /v1/formats returns an io profile and a showcase for each Format. How to read both before you run a paid ad, and what a null means.
- AI video generator for YouTube: get 16:9 from a Sume Format
Sume's first-party example prompts are vertical. To get a 16:9 YouTube video from a Format, say so in instruction and check the width and height.
- Black Friday ads for cosmetics and apparel: which Sume Formats
Adobe expects $9.2B in cosmetics and $51.3B in apparel online this season. Which Sume beauty and fashion Formats fit, and which are still, video or try-on.
- Black Friday ads for electronics and furniture: Sume Formats
Adobe forecasts $63.3B in electronics and $33.4B in furniture online this season. Which Sume Formats show a gadget in use or a room before and after.
Written by Sume