Agent Completion plans the season; a Format bulk queue renders it

One Agent Completion drafts episode beats as JSON; each beat becomes an item in a Format bulk run. Spend caps at both steps.

5 min readSume
All posts

How do you have an AI agent plan a micro-drama season and then produce it? Call Agent Completions once to turn a premise into a JSON list of episode beats, then pass each beat as one item in a Format bulk run. The completion decides what to make; the Format knows how to make it.

Sume's docs separate three surfaces cleanly: a Format stores how to do a job, a schedule stores what to do on a clock, and an Agent Completion stores nothing and takes the task on each call. Planning is the one-off task, so it belongs in a completion.

Step 1: the outline completion

POST /v1/agent/completions needs instruction or messages (exactly one) and generation_spend_cap_usd, which has no default. Send an output_schema so the result is data, not prose. The call returns 202 with an agent.run receipt; poll status_url until next_action is no longer poll_status. The key needs the agent_completions:write scope, and a service-account key is refused.

curl -sS -X POST https://api.sume.com/v1/agent/completions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: season1-outline-v1" \
  -d '{"instruction":"Outline 8 vertical episodes; each ends on a cliffhanger.",
       "generation_spend_cap_usd":2,
       "output_schema":{"name":"outline","schema":{"type":"object",
         "additionalProperties":false,"required":["beats"],
         "properties":{"beats":{"type":"array","items":{"type":"string"}}}}}}'

Step 2: fan out

Read output.beats from the finished run and build the items array: one item per beat, each with its own instruction and input, concurrency between 1 and 16, up to 100 items. Mint a fresh Idempotency-Key for the batch. The bulk queue has no webhook; put communication.webhook_url on each item if you want one.

Which surface does which job (Sume docs read 2026-10-07)
SurfaceStoresStart with
Agent CompletionNothing; task each callPOST /v1/agent/completions
FormatHow to do one kind of videoPOST /v1/formats/{handle}/{slug}/runs
Bulk runA queue of Format runsPOST .../bulk-runs
ScheduledWhat to do on a clockPOST /v1/actions/{id}/runs

Guard the spend

Every run has a cap. The completion's cap is required; a Format run inherits its Format's cap ($400 if none) unless the item sets generation_spend_cap_usd, up to $500. Set a tight cap on the outline call and a realistic one per episode, and meter against the rates on the pricing page.

An outline you have not read is not a license to spend. Check the beats before you start the queue.

What can go wrong

The outline is model output, so treat it as untrusted input to the next step. Validate beats against your own limits before building items: count them (a bulk run takes 1 to 100 items), trim each to the 8,000-character instruction limit, and drop any that repeat an earlier beat.

If the completion returns a degraded or empty result, do not start the queue. Run the outline again with a clearer instruction and a new Idempotency-Key, since reusing a key replays the first answer.

  • Poll the receipt until it settles; do not assume the first read is final.
  • Keep the outline run id with the season record so you can audit it later.
  • Start the bulk run with a concurrency you are willing to pay for at once.

Checklist before you run the season

Confirm the key has the scopes both calls need, because a service-account key is refused on completions. Set a cap on the outline call and a cap per episode. Use an Idempotency-Key on each create, derived from the season and episode so a retry never double-bills. Finally, decide where the finished episodes land: a webhook on each item, or a poll of the queue.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume