Claude batch scripts into a Sume bulk run: holiday video pipeline
Write 100 holiday product scripts with Anthropic Message Batches, then render each as a video with one Sume bulk run. Includes the JSONL-to-items conversion.

Can you use a batch LLM API to write holiday product scripts and then render them as videos in bulk? Yes, as two stages that meet at one JSON file. Anthropic's Message Batches (read 2026-10-04) takes up to 100,000 requests or 256 MB per batch, most batches finish within an hour, and each result is one of succeeded, errored, canceled or expired. Sume's bulk-run endpoint takes the succeeded scripts, up to 100 per queue, and runs one Format per script.
The reason to split the work is cost and shape. Writing 100 short scripts is a text job that tolerates waiting, which is what a batch API is for. Rendering a video is a different job with its own receipt, so it belongs on a Format run.
Stage one: collect the scripts
Submit one request per SKU and set custom_id to the SKU. Anthropic's page says custom_id is 1 to 64 characters from letters, digits, underscore and hyphen, so pick a SKU-safe id. Results are available once all requests have completed (or after 24 hours), and errored, canceled and expired results are not billed.
When the batch ends, you download a JSONL file with one result per line. Keep only the lines whose result.type is succeeded; the rest go on a retry list, because a bulk item built from a missing script would only fail later and cost a slot.
Stage two: turn results into queue items
A Sume bulk item is the same body as a single run: it needs at least one of instruction, input, previous_run_id or attachments. The script becomes instruction, and the SKU goes in input so your own logs can find it. The envelope has exactly two keys, concurrency (1 to 16) and items (1 to 100), so chunk the list at 100.
This script does the conversion on a small inline sample and prints the request body you would POST to /v1/formats/{handle}/{slug}/bulk-runs.
import json
lines = [
'{"custom_id":"mug-01","result":{"type":"succeeded","message":{"content":[{"type":"text","text":"Warm mug, snow outside."}]}}}',
'{"custom_id":"mug-02","result":{"type":"errored"}}',
'{"custom_id":"mug-03","result":{"type":"succeeded","message":{"content":[{"type":"text","text":"Gift-ready box."}]}}}',
]
items, retry = [], []
for line in lines:
row = json.loads(line)
res = row["result"]
if res["type"] != "succeeded":
retry.append(row["custom_id"])
continue
text = res["message"]["content"][0]["text"]
items.append({"instruction": text, "input": {"sku": row["custom_id"]}})
assert 1 <= len(items) <= 100
body = {"concurrency": 4, "items": items}
print(json.dumps(body))
print("retry the script step for", retry)Stage three: post once, key it, poll
Send Idempotency-Key on the create call and mint a fresh one per batch. A spent key replays 202 with the old queue, and the same key with a different payload is 409 idempotency_conflict. Poll GET /v1/format-run-queues/{id} with backoff until the queue status is completed, then branch on counts.failed and counts.canceled rather than treating completed as success.
There is no queue-level webhook. If you want a push, set communication.webhook_url on each item and count terminal deliveries yourself.
Why two stages beat one prompt per video
You could ask the Format to write its own script inside each run, and for a handful of videos that is simpler. Splitting the work pays off at scale for three reasons. First, the script stage is cheap to rerun on its own: if a campaign brief changes you resubmit only the batch. Second, you can read and approve scripts before spending render budget, which matters because each Format run has its own spend cap, generation_spend_cap_usd, that you can set per item. Third, the failure modes stay separate, so an errored script never turns into a failed video.
Keep one habit from the Sume side: put the approved script text, not a reference to it, in instruction. The docs say 8,000 characters are accepted but about 4,000 are carried into the run, so keep scripts short and put long data in input, which allows 2 MiB. A 30-second holiday spot script fits comfortably.
Finally, store the queue id the moment the 202 arrives. There is no list-queues endpoint, so the id is the only handle you will get.
Where each stage stops
| Stage | Unit | Limit that matters |
|---|---|---|
| Scripts | Message Batches request | 100,000 requests or 256 MB per batch |
| Join key | custom_id | 1 to 64 characters, letters digits underscore hyphen |
| Render | Bulk-run item | 1 to 100 items per queue, concurrency 1 to 16 |
| Result | Item status | queued, running, completed, failed or canceled |
Sources
Related posts
More in Developers
- claude -p says Sume needs authentication after one refused write call
Claude Code 2.1.286 stopped headless runs flagging a server as unauthenticated after one refused call. Sume's scope refusal is a tool result, not a 401.
- claude -p killed by timeout or systemd: the Sume job keeps running
A supervisor that kills claude -p does not cancel the Sume job it started. List jobs, then wait or cancel before retrying, or you pay twice.
- claude plugin validate: a clean .mcp.json entry for Sume
Claude Code v2.1.281 makes claude plugin validate check .mcp.json entries. A Sume entry needs only the URL https://mcp.sume.com/mcp and no secret.
- Cloudflare AI Search bills from Nov 1: split retrieval from renders
Cloudflare's October 1 changelog makes AI Search GA with usage billing from November 1, 2026. How to keep retrieval costs separate from Sume render costs.
Written by Sume