Storyboard to finished video over Sume MCP: the tool order
The order a Sume MCP client should follow: stills, inspect, wordless clips, probe, then a timeline dry run and render. Where each step bills and what to skip.

Over hosted MCP, the order Sume's own server guidance gives for a multi-shot video is: generate and approve stills, generate each wordless clip from its approved still, probe the clips, then assemble with timeline_create, dry-running first. Paid steps are the stills, the clips and the render; probing and frame extraction are unbilled.
This post follows that order and notes where a client should stop and look, because the expensive mistakes in storyboard work happen early, before any clip is generated.
What is the order of tools?
Write calls need an idempotency_key, and under OAuth the mcp:write scope. If a client connected with Write off, the paid tools answer insufficient_scope instead of spending anything.
| Step | Tool | Billed? | Why here |
|---|---|---|---|
| 1. Stills | generate_image | Yes | One approved still per shot is the cheapest place to fix design |
| 2. Look | Inspect the stills yourself | No | Accept or regenerate before any video spend |
| 3. Clips | generate_video | Yes | Wordless beats; the accepted still goes in input_references |
| 4. Probe | video_inspect | Unbilled probe and stills | Real durations before you place slots |
| 5. Frames | video_frames_create | Unbilled | Exact stills for seam and drift checks |
| 6. Dry run | timeline_create with dry_run, or POST /v1/timeline-1.0/plan | Unbilled | Cost, duration and segment count |
| 7. Render | timeline_create, then jobs_wait, then timeline_get | Yes, per output minute | One MP4 on the audio spine |
Why are approved stills the gate?
The server guidance says to inspect and accept each still before video spend. A still is the cheap artifact: if the product label is wrong or the character differs from shot to shot, you fix it with an image regeneration, not a clip regeneration. Once a still is approved, the same image goes into input_references for a wordless clip, which is how a character or product carries across shots.
On the REST side of the same models, frame_images with first_frame (and last_frame where the model supports it) is the image-to-video path. If both frame_images and input_references are sent, frame_images wins and the request is image-to-video. Check each model's supported_frame_images and supported_input_references in the model list before choosing.
What about shots where someone speaks?
The same guidance is blunt: video models do not lip-sync, so a speaking face is not a generate_video job with narration laid under it. A talking shot is made as speech first, then a talking-clip job on the accepted still. Only wordless beats go through generate_video. In a storyboard, mark each panel as wordless or spoken at the start so you route it correctly.
All of it is then assembled with one continuous audio spine and video[] shots, each with start, duration and optional transition. Every source_url and audio.url must be Sume-hosted, so mirror outside media first with media-imports_create.
What should the client check before the render?
Run the dry run and read the result. It shows duration_seconds, segment_count and estimated_cost_usd_micros. It cannot predict warnings about short sources, so probe real durations first with video_inspect, which is read-only and free by default (transcription bills the speech-to-text rate if you request it).
Then pull a few stills at the seams with video_frames_create: the last second of one shot and the first second of the next. at values must be inside the clip or the worker fails with frame_time_out_of_range. When the render finishes, warnings[] are soft: padded or looped short sources, snapped transitions and an ignored still motion do not fail the job and do not need a re-render.
Two last notes. Sume does not accept ffmpeg arguments anywhere in this chain; ffmpeg_fields_rejected is the refusal. And every paid write should carry its own idempotency_key, so a retried call returns the original job instead of billing a second time.
Sources
Related posts
More in Agents
- Sume Action run 403 insufficient_scope: old keys and service accounts
A 403 on POST /v1/actions/{id}/runs means the key lacks actions:write, predates the scope, or is a service-account key. Why scopes cannot be added and the fix.
- sume skills export: review the Sume agent skill before you install it
Run sume skills export to read the packaged Sume skill's source before sume skills install writes it into .agents/skills or .claude/skills. Commands and gates.
- Verify a Sume run webhook in Python: replay window and empty secret
A Python verifier for the sume-v1 signature on a Sume run webhook: raw body, five-minute replay window, constant-time compare, and no empty secrets.
- A weekly scheduled agent run for podcast clips: cron, cap and trigger
Set a Sume Scheduled Action to run every week, with a cron schedule, an IANA time zone and a spend cap that a manual or API run can lower but never raise.
Written by Sume