Get a typed podcast clip plan from an Agent Completion

Send a transcript and an output_schema to POST /v1/agent/completions, and get back start and end times for each clip, ready for video-trim.

5 min readSume
All posts

To get a clip plan out of a podcast transcript, call POST /v1/agent/completions with the transcript in input, an output_schema describing the list of clips, and a generation_spend_cap_usd. The run returns structured output that your code can feed straight to video-trim. The cap is required: the endpoint has no default and refuses a request without it.

Request shape

Send either instruction or messages (not both, and assistant turns are rejected), plus optional input, attachments (images only), output_schema and primary_output_key. The only model value is sume-agent. The call returns 202 with an agent.run object whose id starts with agrun_; poll GET /v1/agent-runs/{id} or pass communication.webhook_url to get one signed agent.run.terminal event.

Facts from docs.sume.com/agents/completions, checked 2026-10-01
FieldRule
generation_spend_cap_usdRequired, no default
modelsume-agent only
Scopesagent_completions:write to create, agent_completions:read to poll
Service-account keysRefused
ThreadEach completion runs in a fresh thread; no streaming

A schema for the plan

Keep the schema flat: one array, each item with a start, an end and a title. output_schema is a binding object with a name and a schema, and the schema must sit inside the strict subset: every object, including the array items, sets additionalProperties: false and lists every property in required. A schema outside the subset is refused with 400 output_schema_invalid before anything runs. Put the transcript in input, not in instruction, because input is treated as caller data rather than as instructions.

{
  "instruction": "Pick up to 5 self-contained segments of 20-50 seconds from input.transcript.",
  "input": {"transcript": [{"start": 0.0, "end": 4.2, "text": "Welcome back."}]},
  "generation_spend_cap_usd": 1,
  "output_schema": {
    "name": "acme/clip-plan/v1",
    "strict": true,
    "schema": {
      "type": "object",
      "additionalProperties": false,
      "required": ["clips"],
      "properties": {"clips": {"type": "array", "items": {
        "type": "object",
        "additionalProperties": false,
        "required": ["start", "end", "title"],
        "properties": {"start": {"type": "number"}, "end": {"type": "number"},
                       "title": {"type": "string"}}}}}
    }
  }
}

Then trim

Loop over output.clips and send each as a video-trim with start and end. Use an Idempotency-Key per clip so a retry does not cut twice. Validate the output yourself: an end before a start, or a window outside the source, should be rejected in your code, because the agent can be wrong. Trim itself clamps an end that runs past the source and says so with trim_clamped_to_source.

Setting a sensible cap

Pick the cap from the task, not the other way round. Picking five windows from a transcript is a small reasoning job, so a cap of a dollar or two is a guard rail rather than a budget. Because the field is required and has no default, a missing cap is a 400, which is easier to notice than a surprise charge. Log the cap you sent with the run id on your side so a later cost question has an answer.

Limits

The transcript has to come from somewhere: video-inspect with transcribe:true is the Sume route. The docs page for completions states no size figure for input, so keep a long transcript compact: sentence segments with start, end and text, not raw word timings. The completion does not watch the video unless you give it frames as image attachments, so its picks are based on words, not on what is on screen.

Related posts

More in Agents

All Agents posts

Written by Sume