Reference remix take kinds: which model renders each shot
A remake in Sume is planned as takes, and each take kind names its model: talking, performed, voice, still, recast, genjutsu, product swap. The map, and why.

A remake is a list of takes
When Sume's agent remakes a reference video shot for shot with your product, it does not render one long clip. It writes a plan of takes, and every take has a kind. The kind decides which model makes it. The list of kinds lives in the repository's contract package: talking, performed, voice, insert, still, recast, genjutsu and product swap.
The map matters because each model has a narrow job, and the wrong pairing is the usual way a remake drifts into a different video.
The map
| Take kind | Model | Used for |
|---|---|---|
| talking | Fabric, through the avatar image-to-video tool | A person speaking to camera |
| performed or insert | Seedance, through generate_video | Action acted to a voice-over, or a moving insert |
| voice | TTS, through tts_create | The spoken track |
| still | generate_image with the product image | A still the reference shows |
| recast | H3 Max Recast | Swap the people, keep motion, camera, cuts and sound |
| genjutsu | Genjutsu | Keep the motion, swap characters, product or scenery |
| product_swap | Seedance 2.5 reference-to-video | Same shot with your product |
Why the take is measured after it renders
The instructions say the remake's length and every cut come from the new voice, not the reference's seconds. After a take exists, a reference_take call transcribes that take's own audio and aligns it with its script, so each word has the frame it is actually spoken on. A reference_timeline call then lays the takes out so a cut sits on a word or a take edge.
That is why a recast take is special. A Recast clip keeps the source's own camera, so the instructions say that when one of these source-footage models remakes the whole clip, the clip it returns is the program, and only the music bed and any overlay go over it.
What this means for you
You do not pick take kinds by hand. You point at a reference and say what changes. If the change is the people and the footage is one take, expect a recast take. If the reference is a talking head and you have a new script, expect a talking take. Each model's limits still apply: Recast needs 5 to 30 seconds with no shot over 15, and the Video Router docs list the other rows' windows.
Sources
Related posts
More in Agents
- Remake a viral clip with my product: edit it or rebuild it?
Sume has two remake routes: models that keep the source footage, and reference remix that rebuilds shot by shot. How the agent decides and when to override.
- Scheduled run missing from /v1/jobs? Read /v1/action-runs instead
A Sume scheduled agent run is not a generation job. It never appears in /v1/jobs and has its own statuses, so poll /v1/action-runs/{run_id}.
- Scheduled run body: unknown fields are silently dropped, not a 400
A typo in a Sume scheduled run body is dropped without an error, unlike a Format run. Check names, then read the receipt to see what applied.
- Scheduled AI runs in Q4 2026: ceiling by cadence at $1 a run
At the $1.00 default cap, weekly runs top out at $13 for Q4, daily at $90 and hourly at $2,160. How the schedule cap works and what null changes.
Written by Sume