One Format run with five variants or a five-item bulk queue?

One run gives you five variants under one cap and one receipt. A bulk queue gives five receipts, a concurrency window up to 16 and per-item retry.

4 min readSume
All posts

Use one Format run with a variants[] schema when the five variants share a brief and you want them as a set. Use a bulk queue of five items when each variant is its own job, needs its own cap or webhook, or must retry alone. A bulk request is a server-side queue of ordinary runs, so the difference is how many agent turns and receipts you get.

One run, five variants

One run is one sandbox and one agent turn. The agent sees the whole set at once, which helps with contrast between hooks, and one spend cap bounds the lot. The cost is a shared failure domain. If the run stops, every unfinished variant stops with it, and the only record is the partial output.

A queue of five items

A queue takes concurrency from 1 to 16 and 1 to 100 items, each one a full run body. The server keeps concurrency children in flight and starts the next as a slot frees. Each child binds its own output_schema, spend cap and webhook. The queue itself has no webhook, so you poll GET /v1/format-run-queues/{id} or take each child's terminal delivery.

One run vs bulk queue (Sume docs, read 2026-10-05)
QuestionOne run, variants[]Bulk queue, 5 items
Agent turns15
Receipts15 child receipts plus the queue
Spend capOne per runOne per item
WebhookOne terminal deliveryOne per child; none on the queue
Retry one variantContinue with previous_run_idCancel or re-send that item
ParallelismInside the agent turnWindow of 1 to 16
FailureWhole run fails togetherItem fails, queue goes on

Reading a finished queue

Queue status completed means every item is terminal, not that every item succeeded. Branch on counts.failed and counts.canceled, and read a failed child's receipt for the reason. Use a fresh Idempotency-Key for each batch, because replaying a spent key returns the old queue with 202.

The queue request

The queue body for five hook variants is small. Each item carries its own hook in input, and the hook is data, so keep it out of the instruction. The script prints the queue id and the count of running children from the 202 receipt.

import json, os, urllib.request

HOOKS = ["Stop scrolling", "Three things I wish I knew", "Before and after",
         "Unboxing in ten seconds", "The one setting that matters"]

def main():
    items = [{"instruction": "Make one vertical ad using the hook in input.",
              "input": {"hook": h}, "generation_spend_cap_usd": 5} for h in HOOKS]
    body = {"concurrency": 3, "items": items}
    req = urllib.request.Request(
      "https://api.sume.com/v1/formats/sume/sume-video-hook/bulk-runs",
      data=json.dumps(body).encode(), method="POST",
      headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
               "Content-Type": "application/json",
               "Idempotency-Key": "hooks-batch-2026-10-05"})
    with urllib.request.urlopen(req) as r:
        q = json.load(r)["data"]
        print(q["id"], q.get("counts"))

main()

A rule of thumb

A simple rule works. If the variants are one creative decision, run one. If they are five independent deliverables going to five places, queue five. When unsure, start with the queue, because per-item caps and retries are easier to reason about than a shared turn.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume