Go past 30 s with Seedance 2.5 on Sume: hand off the last frame
Seedance 2.5 extends clips in rounds on the vendor side. On Sume, pull a late frame with video frames, feed it to the next job as a first frame, then join.

ByteDance Seed says Seedance 2.5 generates up to 30 seconds and can be extended over multiple rounds (read 2026-10-05). Sume documents seedance-2.5 at 4 to 30 seconds per job and lists no extension parameter in the pages I read, so a clip longer than 30 seconds is two jobs. The standard hand-off is to pull a late frame from job one, use it as the first frame of job two, and join the clips with Timeline 1.0.
Step 1: pull the hand-off frame
POST /v1/video-frames takes video_url and either at[] (1 to 24 explicit seconds) or fps. Every at value must satisfy 0 <= t < duration, so for a 30-second clip ask for about 29.8, not 30. Frames come back as durable images at source size; send format: "png" for lossless. A submit always returns 202, so poll GET /v1/video-frames/:id until resource_status is ready.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: handoff-frame-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/part1.mp4",
"at": [29.8],
"format": "png"
}'Step 2: start the next job from that frame
The video docs give two image modes. frame_images with frame_type: first_frame starts the clip from a still. Check that the model's supported_frame_images lists first_frame before relying on it. Restate the subject, location and lighting in the prompt, because the model sees only the still.
- Keep the aspect ratio and resolution identical across both jobs.
- Do not describe a new location unless you want a scene change.
- Ask for motion that continues, for example "the camera keeps pushing in".
Keeping continuity across the join
The still carries the subject and place but not the motion. Describe the motion explicitly in the second prompt, and keep camera direction consistent. If the first clip ends with the subject turning left, the next should start with the subject already turned left or the join will jerk.
Lighting drifts across generations. If you see a color shift, apply the same small adjustment to both clips rather than trying to correct only the second, or use a short dissolve. For a long piece, plan a natural cut at the end of the first clip, such as a hand passing the camera, and the hand-off will not need to be seamless.
- Describe continuing motion, not a new scene.
- Keep resolution and ratio identical.
- Hide the seam with a beat, not a fade.
Cost and count
A 60-second video is two 30-second jobs, one frame extract and one render. The frame extract bills by its own Modal compute, so confirm the figure in the job result and in the catalog. The render for a clip under a minute is $0.10 at the docs' public rate. For three parts, repeat the handoff twice.
Step 3: join and hide the seam
Judge the seam on a few frames from each side with video frames before you ship it. For vendor behavior, read the Seed post.
| Step | Route | Notes |
|---|---|---|
| Generate part 1 | seedance-2.5, 30 s | 4 to 30 s per job |
| Pull frame | POST /v1/video-frames | at: [29.8] |
| Generate part 2 | seedance-2.5 with frame_images | First frame from step 2 |
| Join | POST /v1/timeline-1.0/render | $0.10 per output minute, rounded up |
Sources
Related posts
More in Use cases
- Seedance 2.5 reference inventory: 30 images, 10 clips, 10 tracks
A pre-flight checklist for the 30 image, 10 video and 10 audio reference slots Seedance 2.5 offers, and how to check what Sume accepts.
- Seedream 5.0 Flash style blending: a five-step brand board test plan
A five-step plan to test multi-reference style blending on Seedream 5.0 Flash, with a Sume seedream-5-lite request to run the same test on a catalog model.
- Sermon series title loop: Wan 3.0 first and last frame from one image
Make a looping sermon-series title background: send one image as both first and last frame to Wan 3.0 on Sume; 10 seconds costs $1.25 at 720p.
- Incident update as an avatar video: text first, then a 30s clip
A 30-second Sume avatar clip costs $5.52 on Standard, $7.35 on Plus or $16.50 on Max (no product). Post the text update first; the render is asynchronous.
Written by Sume