Reference remix take kinds: which model renders each shot

A remake in Sume is planned as takes, and each take kind names its model: talking, performed, voice, still, recast, genjutsu, product swap. The map, and why.

5 min readSume
All posts

A remake is a list of takes

When Sume's agent remakes a reference video shot for shot with your product, it does not render one long clip. It writes a plan of takes, and every take has a kind. The kind decides which model makes it. The list of kinds lives in the repository's contract package: talking, performed, voice, insert, still, recast, genjutsu and product swap.

The map matters because each model has a narrow job, and the wrong pairing is the usual way a remake drifts into a different video.

The map

Take kinds and the model each one names (Sume agent tool descriptions, read 2026-10-03)
Take kindModelUsed for
talkingFabric, through the avatar image-to-video toolA person speaking to camera
performed or insertSeedance, through generate_videoAction acted to a voice-over, or a moving insert
voiceTTS, through tts_createThe spoken track
stillgenerate_image with the product imageA still the reference shows
recastH3 Max RecastSwap the people, keep motion, camera, cuts and sound
genjutsuGenjutsuKeep the motion, swap characters, product or scenery
product_swapSeedance 2.5 reference-to-videoSame shot with your product

Why the take is measured after it renders

The instructions say the remake's length and every cut come from the new voice, not the reference's seconds. After a take exists, a reference_take call transcribes that take's own audio and aligns it with its script, so each word has the frame it is actually spoken on. A reference_timeline call then lays the takes out so a cut sits on a word or a take edge.

That is why a recast take is special. A Recast clip keeps the source's own camera, so the instructions say that when one of these source-footage models remakes the whole clip, the clip it returns is the program, and only the music bed and any overlay go over it.

What this means for you

You do not pick take kinds by hand. You point at a reference and say what changes. If the change is the people and the footage is one take, expect a recast take. If the reference is a talking head and you have a new script, expect a talking take. Each model's limits still apply: Recast needs 5 to 30 seconds with no shot over 15, and the Video Router docs list the other rows' windows.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume