Can one avatar UGC ad change location mid-video? One scene only

Sume's avatar video renders one avatar and one shared scene per final video. For a second location, make two jobs and join them with a Timeline 1.0 render.

5 min readSume
All posts

No: a single Sume avatar video does not move the creator between locations. The docs say current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene. Multi-scene video_inputs give you hook, demo and call-to-action beats in one composed video, but they share that one background. To show a second location, render two avatar videos and join them with Timeline 1.0.

This is from Sume's Generate avatar video page, with the join steps from Timeline 1.0 and Audio detach, all read on 2026-10-02.

What does `video_inputs` give me, then?

Ordered scenes with their own voice lines. Each scene has an id, a voice (type: text with exactly one of script or input_text, or type: silence with a required duration) and a background. Total planned duration must land in the 4 to 60 second window. The docs' example uses the same background prompt on all three scenes, which is the supported shape. For first-frame review, Avatar video previews say a multi-scene preview with a shared scene continues later stills from the first frame's pose.

What is the join recipe for two locations?

The scene for each clip is set with scene: { "type": "prompt", "prompt": "..." } or a photo scene with a public HTTPS image_url. Timeline 1.0 needs an audio spine, which is why the voice is detached first. Timeline is priced per output minute and audio detach per job; the docs say to confirm both live in GET /v1/catalog.

Steps to a two-location UGC ad from Sume's docs (read 2026-10-02)
StepCallDocs note
Render clip APOST /v1/avatar-1.0/talking-video with scene ASame avatar_handle for both clips
Render clip BSame call with scene BEach has its own Idempotency-Key
Pull each voicePOST /v1/audio-detach per clipReturns audio_url; wav by default
JoinPOST /v1/timeline-1.0/renderaudio.parts[] gapless slices plus ordered video[] slots

What does the render call look like?

audio.duration_seconds is the output length. Declared part lengths must not sum to less than it (audio_parts_shorter_than_duration). video[0].start must be 0, and later starts must increase. The URLs below are placeholders; use your own artifact URLs.

# 1) Two avatar jobs, same avatar_handle, different scene.prompt, each with its own Idempotency-Key.
# 2) POST /v1/audio-detach for each finished clip -> audio_url (wav).
# 3) Join them:
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ugc-two-locations-001" \
  -d '{
    "audio": {
      "duration_seconds": 18,
      "parts": [
        { "url": "https://media.sume.com/artifacts/artf_a/voice-a.wav", "duration": 8 },
        { "url": "https://media.sume.com/artifacts/artf_b/voice-b.wav", "duration": 10 }
      ]
    },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_a/clip-a.mp4", "start": 0, "duration": 8 },
      { "source_url": "https://media.sume.com/artifacts/artf_b/clip-b.mp4", "start": 8, "duration": 10 }
    ]
  }'

What should I watch for?

Three things to check before you publish the joined clip:

  • Whether the creator looks the same in both clips. The avatar handle is the same, but the docs make no promise about identical framing across separate jobs, so review both before joining.
  • Use a hard cut or a short fade transition; Timeline transitions go on slots after the first.
  • If only one beat needs a change, regenerate that part rather than the whole video.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume