A/B test two hooks in one bulk queue: instruction beats the recipe

Run two opening hooks on one product by sending two items to the same Sume Format. Your instruction is placed after the recipe, so it wins where they disagree.

4 min readSume
All posts

Send two items to one bulk queue on the same Format: identical input, and an instruction that differs only in the hook line. Sume places your instruction after the Format body, so where the two disagree the run does what you asked. That makes the hook the one variable.

How the run is composed

The agent receives, in order: a pointer at the Format recipe, the attached package, your instruction (or the Format's default), an unattended-run note, a pointer at your input, and any attachments. Your instruction is limited to 8000 characters, of which roughly the first 4000 are carried as prompt text, so keep it short and put data in input.

{
  "concurrency": 2,
  "items": [
    { "instruction": "Open on the price drop.",
      "input": { "sku": "8823", "variant": "A" } },
    { "instruction": "Open on a customer question.",
      "input": { "sku": "8823", "variant": "B" } }
  ]
}

Keeping the test fair

Two runs of a generative recipe will differ in many ways beyond the hook: images, timing, voice take. One pair can show a difference that is only noise, so repeat each variant several times before you conclude anything.

What to hold constant in a hook test (read 2026-10-06)
Hold constantHow
Format versionCheck format.version on both receipts
Orchestrator modelSet the same model field, or leave both on the default
Input dataSame input object except the variant label
Spend capSame generation_spend_cap_usd on both items

Reading the results

Each child has its own receipt with output, artifacts[] and usage. Pull the two primary_output_url values and review them side by side, and compare usage.billable_amount_usd_micros so a variant that costs more is visible.

For the real answer, publish both and let the ad platform's own reporting decide, with enough impressions to mean something. The queue only produces the two candidates; it does not measure them.

If the recipe itself fights your hook, the instruction still wins where they disagree, but a long list of overrides is a sign that the recipe should change. Edit the Format once instead of repeating the same correction in every item.

Keep the comparison fair by changing only the hook between the two variants, and label each item in your own records so you can tell which result came from which hook once the batch finishes.

Tradeoff

A bulk queue is a good way to start both variants together, but the queue does not compare them for you. You read each child receipt and judge the outputs yourself, and results from the ad platform decide the winner. Never claim a hook won from one render each.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume