Luma Ray 3.2 edit controls: pose, depth, normals, and Sume's edit

Luma's Ray 3.2 edit takes auto_controls, nine strength presets, or per-signal pose, depth, normals, trajectory and face controls. Sume's edit is prompt-only.

5 min readSume
All posts

Luma's Ray 3.2 video edit can be steered three ways: auto_controls: true, one of nine strength presets, or per-signal controls for pose, depth, normals, trajectory and face. Sume's edit on gemini-omni-flash-1.1 takes a source video_url and a prompt that describes the change, with no control signals to tune.

The Luma details are from its video editing guide and models page, read 2026-10-03. Sume's are from the Video Router docs.

What are Luma's three ways to steer an edit?

auto_controls: true is the default: the model derives conditioning from the source video. It cannot be combined with a manual strength or with controls.

The nine strength presets sit in three bands. The adhere band (adhere_1 to adhere_3) keeps the source motion and composition, the flex band (flex_1 to flex_3) is the middle, and the reimagine band (reimagine_1 to reimagine_3) follows the prompt most. Per-signal controls are mutually exclusive with auto controls.

Luma Ray 3.2 per-signal controls, read 2026-10-03
SignalSetting
poseenabled, strength precise or coarse
depthenabled, blur from 0 to 1
normalsenabled, augmentation from 0 to 1
trajectoryenabled, sparsity from 0 to 1
faceenabled on or off only

What are the source limits?

The guide lists a source of at most 18 seconds and at most 200 MB for a URL or inline upload. The source can be passed as a generation_id, url, data or file_id. The output keeps the source aspect ratio, and resolution is 360p, 540p, 720p or 1080p, with HDR needing 720p or 1080p.

The guide also says keyframes are the strongest control lever: pin a guide image at a chosen source frame and the model honors that look.

How does Sume's edit differ?

On Sume the edit lives on gemini-omni-flash-1.1. You send video_url and a prompt such as "Replace the bottle with an apple. Keep everything else the same." Resolution is optional and defaults to 720p, and aspect_ratio and duration are not sent. The video_url is the edit source, not a reference, so it cannot be combined with image_url, end_image_url or reference_*_urls.

There is no strength, pose or depth control in the documented request. If you need an edit that preserves motion exactly, Luma's adhere presets and pose control are the more direct tool; if you need a quick prompt-described change inside one hosted pipeline, Sume's is simpler. For a different job, swapping the people in a clip, Sume also lists h3-max-recast, which takes a source video and one to four person photos.

Which should you test first?

Take one clip of 10 seconds or less, which is inside both limits, and run it through each with the same change request. On Luma, try adhere_2 and then flex_2. On Sume, vary only the wording. Compare how much of the original motion survives. The result tells you whether you need Luma's controls or whether a prompt is enough. Neither page publishes a quality benchmark, so your own clip is the only evidence.

What do the per-signal controls buy you?

Each signal is a separate piece of the source video the model can be told to respect. Pose holds body position, depth holds the scene layout, normals hold surface shape, trajectory holds the path things move along and face holds identity. Turning on only the ones you care about leaves the model free on the rest, which is how you change a background while keeping a performance.

That is a finer dial than a prompt can provide, but it is also more to tune. If you do not already know which signal matters for a given change, start with auto_controls: true and add manual control only when the first result drifts.

Sources

Related posts

More in Models

All Models posts

Written by Sume