Timeline plan first: the unbilled estimate for a 30-second Short
POST /v1/timeline-1.0/plan compiles your cut without a job or a charge and returns duration, segments and estimated cost. A 30-second render bills one minute.

Call POST /v1/timeline-1.0/plan before you render. It runs the schema checks, the Sume-host URL checks and the compiler, and it returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros without creating a job or reserving credits. A render is $0.10 per output minute rounded up, so a 30 second Short bills as one minute.
What the plan does
The behavior is in the Timeline 1.0 docs. The plan does not download media, does not need an Idempotency-Key, and does not predict the soft warnings that appear when a source is shorter than its slot. Those are padding or looping warnings that show up only on the real render.
The numbers for a 30 second cut
The numbers below are from the docs (read 2026-10-05). The cost line is arithmetic on the public rate; the live rate is on GET /v1/catalog.
| Item | Value |
|---|---|
| Output length | 30 s |
| Public rate | $0.10 per ceil(output minute) |
| Billable minutes | ceil(30 / 60) = 1 |
| Cost of the render | $0.10 |
| Provider inference | None, worker ffmpeg only |
| Default output | 1080x1920 MP4 |
Mind the minute boundary
The reserve on a real render is ceil(audio.duration_seconds / 60) minutes. That means a 61 second Short bills two minutes, and a 59 second one bills one. If you are choosing between 58 and 62 seconds for a cut, the plan shows you the difference before you pay it.
A plan request
The plan body is the same as the render body, minus the idempotency requirement. Here is a two-slot cut over a 30 second silent spine, where the first slot is 12 seconds and the second 18:
curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio": { "mode": "silence", "duration_seconds": 30 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/a.mp4",
"start": 0, "duration": 12 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/b.mp4",
"start": 12, "duration": 18,
"transition": { "type": "fade", "duration": 0.25 } }
]
}'Errors the plan catches
The first slot must start at 0, later starts must increase, and coverage may stop at most 0.5 seconds before the end of the spine. A transition may only sit on a slot after the first, with a duration of at most 1 second and at most half of the shorter neighbor. If the plan finds a problem, you get a stable code such as timeline_must_start_at_zero or transition_too_long instead of a rendered file.
Then render
When the plan looks right, send the same body to POST /v1/timeline-1.0/render with an Idempotency-Key. The default mode is async, so poll GET /v1/jobs/:id/status and read the result. You can send mode: "sync" to wait up to 30 seconds for the finished job.
The plan knows nothing about platforms
Platform length limits are separate. YouTube Help (read 2026-10-05) gives 3 minutes for a Short, so a 30 second plan is well inside it; the plan does not know about any platform limit and will happily plan a longer file.
Read the estimate in micros
The estimate is in micros, so 100000 means ten cents. Divide by 1,000,000 before you show it to a person, and guard the division so a missing field never shows up as a broken number. If the field is absent, show nothing rather than a computed value, and treat that as a bug in your request.
One body, two calls
Keep the plan and the render in sync by building the body once and sending it to both endpoints. A diff between the two bodies is the usual way a cost surprise happens. If you change a slot after the plan, plan again.
Longer cuts
For a longer Short, the same call scales. The audio spine can be up to 1800 seconds and the cut up to 200 slots, so even a 3 minute Short, the YouTube Help maximum, is a three-minute render at $0.30 by the same arithmetic. Check the live rate on the catalog before you rely on it.
Fewer slots, simpler render
The plan also helps when you are choosing between cutting styles. Fewer slots mean fewer joins, and the docs say a render with more than 12 segments is chunked by default (render.strategy: auto), while forcing single above 12 slots is refused with render_strategy_unsafe. For a 30 second Short you will usually be well under that number, so leave the strategy on its default.
The takeaway
Run the plan every time, read billable_minutes, and render only when the numbers match what you expected.
Sources
Related posts
More in Developers
- Timeline plan for a 40-second Omni stitch: four segments, one minute
Run Sume's unbilled timeline plan on four 10-second Omni clips to see segment_count, billable_minutes and the cost before you render. Request body included.
- A 30-line Node proxy so a browser can start a Sume video, no key
A node:http server with only POST /render and GET /status/:id. It fixes the model and clip size and keeps SUME_API_KEY on the server, away from the browser.
- A tiny Node proxy so a browser can order a 4K Omni clip safely
Sume keys are server-side only. A short Node proxy exposes POST and GET routes for one 4K Gemini Omni Flash 1.1 clip, with an Idempotency-Key and pinned fields.
- tqdm in Jupyter for a 30-second AI video render, no percentage
Sume gives job statuses, not a percent. Use a tqdm elapsed-time bar with the status as its label, and stop on terminal. Runs in a notebook cell.
Written by Sume