Loop a 30-second music bed under a 3-minute Reel: the body
A short bed can cover a 180-second Reel with soundtrack.loop, duck_db and a fade-out. One Timeline 1.0 body, the 3-minute bill, and the limits to respect.

Yes: set loop on the soundtrack
A Reel that runs up to three minutes does not need a three-minute music file. In Timeline 1.0, soundtrack is an optional bed with a url, a gain_db, a loop flag, a fade_out_seconds up to 10 and a duck_db from 0 to 20. With loop: true a 30-second bed repeats under the whole render, ducks while the spine speaks, and fades out at the end. The soundtrack URL must be a media.sume.com artifact of your workspace like every other input.
YouTube's help page states that Shorts creation tools make videos up to 3 minutes long, so 180 seconds is the practical ceiling this body is built around. This post does not cite a vendor page for Instagram Reels length, so confirm the current Reel limit in the app before relying on 180 seconds.
The one caveat about looping: the seam
The docs promise that the bed repeats; they do not promise that your 30 seconds loops without an audible seam. That depends on the file. A bed whose last beat does not land on its first beat will thump every 30 seconds. If you generate the bed with Sume's music model, ask for a loopable instrumental in the prompt and listen to the join once before rendering the whole Reel. Our series theme post covers one way to reuse a single track across a season.
Ducking has a precondition. duck_db needs a real audio spine, so it works with a voiceover and fails with duck_requires_audio_spine under audio.mode: "silence". For a silent b-roll Reel, set the bed's gain_db where you want it and skip ducking.
The render body
The spine is a 180-second voice file, two 90-second video slots with a half-second dissolve on the second, a looped 30-second bed at -16 dB ducked by 10 dB, a 4-second bed fade, and a 1-second fade on the whole output. The values are examples to tune, not recommendations.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: reel-180-bed-001" \
-d '{
"audio": {"url": "https://media.sume.com/artifacts/artf_demo/vo-180.wav",
"duration_seconds": 180},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_demo/a.mp4",
"start": 0, "duration": 90},
{"source_url": "https://media.sume.com/artifacts/artf_demo/b.mp4",
"start": 90, "duration": 90,
"transition": {"type": "dissolve", "duration": 0.5}}
],
"soundtrack": {"url": "https://media.sume.com/artifacts/artf_demo/bed-30s.mp3",
"gain_db": -16, "duck_db": 10, "loop": true,
"fade_out_seconds": 4},
"output": {"width": 1080, "height": 1920, "fps": 30, "fade_out_seconds": 1}
}'What it costs and what it checks
The bill follows the declared spine length, not the bed or the number of loops.
| Spine length | Billable minutes | At $0.10 per minute | Note |
|---|---|---|---|
| 59 s | 1 | $0.10 | Under a minute |
| 120 s | 2 | $0.20 | Exact minutes |
| 180 s | 3 | $0.30 | The 3-minute Reel in this body |
| 181 s | 4 | $0.40 | One second over costs a whole minute |
Limits to keep in the margin
The soundtrack fade cannot exceed the output: a fade longer than the spine is refused with soundtrack_fade_exceeds_output. The edge fades on the whole output are 0 to 5 seconds each, and their sum must fit within the output length (edge_fades_exceed_output). A transition is at most 1 second and at most half of the shorter neighbor. The two video slots must start at 0 and then increase, and coverage can end at most 0.5 seconds before the spine, so the second slot here reaches exactly 180.
Before paying, post the same body to POST /v1/timeline-1.0/plan. It returns billable_minutes and estimated_cost_usd_micros without creating a job. If the plan says four minutes when you expected three, you declared 181 seconds somewhere. Confirm live rates in GET /v1/catalog, since the docs defer to it.
If you need the bed to start later than the first frame, or to change between sections of the Reel, a single soundtrack field is not the tool: it is one bed under the whole render. For two musical sections, join the audio first with timeline audio concat, which splices up to 20 parts in the sample domain with no re-synthesis, and use the merged file as the spine or the bed. That keeps the render body simple and moves the musical decision to where it can be reused.
Finally, keep the finished render's warnings[]. A looped or padded source is reported as a soft warning and is not a failure, and the plan cannot predict it because the plan never downloads media. Reading the warnings is the only way to learn that a clip was shorter than the slot you gave it.
Sources
Related posts
More in Media tools
- MAI-Voice-2.1 is not on Sume: make talking clips with Sume TTS
Microsoft MAI-Voice-2.1 launched 2026-10-01 but Sume does not list it. Here is the path that ships: Sume TTS audio, then H3 Max lip sync on a still.
- MiniMax H3 Max lip sync API: your first clip in Python
Submit a still and a Sume-hosted audio file to POST /v1/minimax/h3-max/lip-sync, poll the job, and read the result. Python stdlib, with the 5 to 14.8 s rule.
- MiniMax H3 lip sync at 2K? Sume offers 480p, 768p and 1080p
MiniMax H3 the video model lists 2K, but Sume's H3 Max lip-sync route stops at 1080p. The three resolutions, their derived per-second prices and how to choose.
- MiniMax H3 Max video or H3 Max lip sync: fixing the wrong_tool error
generate_video refuses minimax/h3-max/lip-sync with wrong_tool. Use avatar-image-to-video_create for lip sync, minimax-h3-max for text or image video.
Written by Sume