Turn a new video model launch into a repeatable Sume Format

A new video model dropped and one test clip looked great. Save the recipe as a Format, call it with a key, a cap and a webhook, and keep the result repeatable.

5 min readSume
All posts

Do not rebuild the prompt every launch week. Save the recipe that worked as a Format, call it at {handle}/{slug}, and send each product as input. A Format run takes an idempotency key, a spend cap and a webhook, so the same call works on day one and day ninety, whichever video model the recipe uses.

Split the recipe from the data

A one-off test mixes two things: the steps that made the clip good, and the product it was made for. A Format keeps the steps. The caller keeps the data. The run body has an instruction (up to 8000 characters, of which about 4000 reach the run as prompt text) and an input object (up to 64 top-level keys and 2 MiB). Sume writes input to a file in the run workspace and tells the agent that the file is caller data, not instructions.

  • Put the shot order, pacing, framing and tool choices in the Format body.
  • Put the product name, URLs, price and language in input.
  • Put per-run overrides (aspect ratio, no BGM) in instruction; it is placed after the Format body and wins where they disagree.

Call it the same way every time

The address is POST /v1/formats/{handle}/{slug}/runs. A Format that you authored answers at your handle. A first-party Format answers at sume/{slug}. The call below is the shape that production callers use, taken from the create-a-run page.

curl -sS -X POST "https://api.sume.com/v1/formats/acme/live-commerce/runs" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sku-8823-launch-week-v1" \
  -d '{
    "instruction": "Vertical 9:16, Korean host, no BGM.",
    "input": { "product_name": "Aurora Headphones", "vo_language": "ko" },
    "generation_spend_cap_usd": 120,
    "communication": { "webhook_url": "https://acme.example.com/hooks/sume" }
  }'

Key, cap and webhook are the repeatable part

Derive the Idempotency-Key from the item and a version that you bump on purpose, here the SKU and v1. A retry with the same key and body returns the original receipt with idempotency_hit: true and does not start or bill a second run. Bump the version when you want a new take after a model change.

generation_spend_cap_usd is the ceiling for this run. It accepts a number up to 500, and 0 is rejected. The webhook delivers one signed format.run.terminal POST when the run completes or fails, so you do not need a poll timer for every item.

What a launch changes, and what it does not

The model field is easy to misread. It names the Agents catalog LLM that orchestrates the run, and the default is gpt-6-sol. The Format's tools choose the image, video and audio models. Read this post before you try to swap a video engine with it.

Read from docs.sume.com/formats/call on 2026-10-05
You want toChange thisNot this
Try the new video modelThe Format body, in a fork of the FormatThe model field on the run
Change the host, language or priceinputThe Format
Re-run one SKU on purposeThe version in the idempotency keyThe Format slug
Limit the blast radius of a bad recipegeneration_spend_cap_usdNothing; the cap is per run

A launch-week routine

Fork the Format in the Format library, change the recipe, and run it on three real products with a low cap. Compare the receipts: each one shows format.version, usage.billable_amount_usd_micros and the artifacts. When the fork wins, keep it as the Format your backend calls, and leave the integration code alone.

Poll expires_at on a non-terminal receipt as your timeout. Sume force-finalizes a run as failed after 90 minutes from created_at, or sooner if it is older than 25 minutes and silent for 10.

Related posts

More in Formats

All Formats posts

Written by Sume