Halloween costume contest voting reel from phone clips, 46 cents

Trim eight phone clips, join them in one Sume timeline render, then burn Costume 1 to 8 number cards with one caption job. Total $0.46, no music.

4 min readSume
All posts

To build a Halloween costume contest voting reel from phone clips, trim each entry to five seconds with Sume's video trim, join the trims in one Timeline 1.0 render with hard cuts, then burn "Costume 1" to "Costume 8" number cards in a single caption job. For eight entries that is $0.16 for trims, $0.10 for the render and $0.20 for the captions, $0.46 in all.

Why number the entries on the video itself?

Voters need to name the costume they liked without knowing the entrant, and a number on screen gives them that. A card burned into the pixels also survives a screen recording, a repost and a muted autoplay, which a comment typed under the post does not.

Sume's caption endpoint can burn authored text with no speech in the clip. You send cues, each with text, start and end in seconds, and Sume does not run speech-to-text on that path. A clip with no audible speech would fail speech captions with caption_no_speech, so cues are the right input here.

What are the steps?

Import each phone clip first, because the trim, timeline and caption calls only take media that Sume hosts or can fetch at a public HTTPS address. Then run these steps.

  • Trim: one POST /v1/video-trim per entry with start and duration: 5. Each trim is a flat $0.02 and returns a new artifact URL.
  • Plan: POST /v1/timeline-1.0/plan is unbilled and tells you the duration, segment count and estimated cost before you spend anything.
  • Render: one POST /v1/timeline-1.0/render with audio.mode: "silence", a declared length of 40 seconds and eight video[] slots starting at 0, 5, 10 and so on.
  • Caption: one POST /v1/video-captions on the rendered MP4 with eight cues, "Costume 1" at 0 to 5 seconds, "Costume 2" at 5 to 10 seconds, and so on.

How does the render request look?

Two slots are shown below; repeat the slot object for the other entries. Leave out transition on every slot. The timeline compiler honors declared starts, but a fade overlaps its neighbors, and the whole point of this reel is that each number card lines up with exactly one costume. Hard cuts keep the cue times equal to the slot times.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: costume-contest-2026-render" \
  -d '{
    "audio": { "mode": "silence", "duration_seconds": 10 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_a/entry1.mp4",
        "start": 0, "duration": 5, "fit": "cover" },
      { "source_url": "https://media.sume.com/artifacts/artf_b/entry2.mp4",
        "start": 5, "duration": 5, "fit": "cover" }
    ]
  }'

What does it cost for six, eight or twelve entries?

Past 12 slots the render splits into chunks automatically, and the caption price Sume documents applies to videos of up to 60 seconds, so the table stops at twelve entries of five seconds.

Costume contest reel at 5 seconds per entry, Sume rates read 2026-10-11
EntriesTrims at $0.02Render (one minute or less, $0.10)Caption jobTotal
6$0.12$0.10$0.20$0.42
8$0.16$0.10$0.20$0.46
12$0.24$0.10$0.20$0.54

What about music and the sound of the entries?

A silent spine keeps the cost at the figures above. If you want a bed, Sume's Music 1.0 is a fixed $0.125 per accepted generation and goes in the render's soundtrack. Ducking under a voice needs a real audio spine, so it does not apply to a silent reel.

Before you post, pull a still from each card time with video frames and check that every number is the one you meant. The caption price does not change if you restyle the same video later with source_caption_id, so a wrong number is cheaper to fix than to explain.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume