dry_run or generation_admission_preview first for 12 Sume clips?

Use dry_run on a paid call for its estimate, and generation_admission_preview before a burst of 12 clips to check balance and queue room. Neither submits a job.

4 min readSume
All posts

For a burst of twelve clips, call generation_admission_preview first to see whether the balance and the queue can take the whole wave. Then run dry_run=true on the first real call to read the exact estimate for that request. Sume's docs say a normal single create does not need either step, but recommend them before expensive bursts. Neither one submits a job.

What each one does

The tools-and-gates page calls dry_run an optional admission and cost preview that does not submit the job. Playbook B names the same step for a single paid tool: call generation_admission_preview or the paid tool with dry_run=true, then examine the estimate, the balance and the queue behavior, and submit with a new idempotency_key. Generation admission describes the read-only preview as an optional preflight that must not create a job, reserve credits, capture usage, refund usage or call generation providers.

Preflight choices on hosted MCP, from Sume docs read 2026-10-08
StepBest forSpends anything?
dry_run=true on the paid toolThe estimate for this exact requestNo job is submitted
generation_admission_previewBalance and queue room before a burstNo job, no reservation
max_spend_usd on the real callA hard cap, enforced only when you send itRefuses a call above the cap
balance_getCurrent balance snapshotRead only

Twelve clips, worked through

Say the plan is twelve 8-second clips on sume/auto at 720p. At 100 cents per clip, the wave is 12 x 100 = 1,200 cents, or $12.00. Check that against two limits.

  • Balance: if balance_get shows less than $12.00, the wave cannot all be reserved, and a submit returns 402 insufficient_credits before provider work starts.
  • Queue: Pro accepts 24 paid jobs at once (4 processing plus 20 queued), so twelve fit. Free accepts 6, so six would be accepted and the rest would return 429 queue_full.
  • Cap: pass max_spend_usd of about 1.05 on each call (a clip is $1.00), so a mistaken payload cannot spend a lot more than a clip.

Pace the wave from the snapshot

Where Sume returns generation_limits, use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs) as the budget for new in-flight work, capped at queue_capacity_remaining. Do not treat wave_size_hint as concurrency: it is max(1, floor(queue_capacity_remaining * 0.75)), a submission hint only. For an empty Pro workspace, queue_capacity_remaining is 24 and the hint is floor(24 x 0.75) = 18.

The counts are a snapshot and can change right after the response, so refresh the preview before you add the next group.

The order of operations

Put these in the agent's instructions, in this order: read balance_get; call generation_admission_preview; call the paid tool once with dry_run=true; submit with distinct keys and a cap; wait with one batch jobs_wait; read the results in one batch jobs_result. If any step reports less room than needed, stop and ask, rather than submitting a partial wave and hoping.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume