ffmpeg script or Sume Timeline for a weekly vertical series
Compare a self-hosted ffmpeg script with Sume Timeline 1.0 for a weekly 9:16 series: what you maintain, what the API refuses, price per output minute.

The short version
If you cut a weekly vertical series and your edits are a list of clips, a voice track and a music bed, Sume Timeline 1.0 replaces the part of your ffmpeg script that handles slots, fades and ducking, and the bill is $0.10 per whole output minute. If your edits need arbitrary filters, Timeline will not do them: it has a fixed set of fields, and the video filter API allows dimming and cropping but refuses setpts, trim and drawtext. This post compares what you maintain in each case. It does not benchmark speed; we have no measurement to give.
The trigger is practical. The October platform roundup describes Shorts series with seasons and episodes, plus custom thumbnails, rolling out from 23 September. A season means a dozen of the same job, and that is where a hand-edited script starts to cost time.
What each side owns
The right-hand column is what Sume's docs say it does. The left column is generic: it names work any self-hosted script has, with no claim about a particular ffmpeg version.
| Part of the job | A script you host | Sume Timeline 1.0 |
|---|---|---|
| Machine and install | You provide and patch it | Runs on Sume workers |
| Slot timing | Your code | video[] slots with start, duration, source_in; starts must increase and slot one starts at 0 |
| Transitions | Your filter graph | fade, wipeleft, wiperight, slideup, slidedown, dissolve; at most 1 s and 8 chained |
| Music under a voice | Your mix commands | soundtrack with gain_db, loop, fade_out_seconds up to 10, duck_db 0 to 20 |
| Cost estimate | Your own guess | Unbilled plan: duration, segment count, billable minutes |
| Free-form filters | Anything ffmpeg allows | Not available in Timeline |
| Failure handling | Your retries | Jobs with status, result and webhooks |
Where the self-hosted script still wins
It wins when the edit is not a timeline of whole clips. Per-frame effects, text drawn at arbitrary positions, and anything that needs a filter Sume's allowlist refuses are on your side of the line. It also wins when you must keep media off a third-party host: Timeline takes media.sume.com sources, so your clips have to be imported first.
It also wins on cost if you already own the compute and have spare capacity. We will not put a number on that, because the cost of your machine is yours.
Where Timeline wins
A series where every episode is the same shape: a hook, N clips, a voice, a bed, an ending fade. The body is a small JSON document you can generate, and the API refuses the combinations that would fail halfway, such as a transition on the first slot or a soundtrack duck without a real audio spine. Two documented refusals are worth reading before you migrate: duck_requires_audio_spine and render_strategy_unsafe. They show how strict the checks are.
For a 3-minute episode, which YouTube's help page says Shorts can reach, the spine limit is 1800 seconds, so length is not the constraint; the 200-slot cap is. If your editing style needs more than about one cut per second for three minutes, split the episode.
A migration plan that is cheap to try
Keep the script. Pick one episode and write its Timeline body by hand. Post it to the plan call, read the billable minutes, render it once, and compare the result against your script's output by eye. If the cut, the fades and the mix are acceptable, generate the next episode's body from your existing edit list. If a part of the look is missing, you have learned which part of your script was doing the work Timeline cannot.
Whichever way you go, write down the edit as data. A list of clips with start times and durations is easier to diff, review and regenerate than a shell command with seventeen arguments. If you can express your series that way, the move to Timeline is mostly renaming fields, and the move back is just as easy.
One more point on cost: a render is billed per whole output minute, so a 61-second episode is two minutes. A script has no such step, but it has the cost of your own time when it breaks on a source with an odd frame rate. Timeline reports that situation as the output_fps_resamples_sources warning, and a gate on warnings turns the problem into a visible state.
If you are undecided, run the comparison on one real episode and keep the numbers you measure yourself: your own machine time, your own failure count, and the plan's billable minutes. Those three figures will answer the question better than any general advice, including this post.
Sources
Related posts
More in Comparisons
- Groq Orpheus TTS: $22 per million characters and a 200-character cap
Groq lists Orpheus V1 English at $22.00 per million characters with input kept under 200 characters. Request-count and cost math against Sume TTS.
- Ideogram 4.5 is 0.8c to 22c per image: which tiers does Sume sell?
Ideogram lists four quality modes from 0.8 cents to 22 cents at native 2K. Sume's Image API sells low, medium and high at $0.0375, $0.075 and $0.275 an image.
- Lemonfox TTS at $2.50 per million characters vs Sume: the honest gap
Lemonfox's page works out to about $2.50 per million characters. Sume lists $47.50. Cost for 900, 22,500 and 1M characters, and what the gap buys you.
- Lip-sync a video you have, or animate a photo: which API?
Re-syncing a mouth in footage and making a still talk are different jobs. A guide using Sync.so docs and Sume lip-sync, avatar and face-swap routes.
Written by Sume