H3 Max Recast prompt: optional, and what Sume does with one
On Sume, H3 Max Recast runs without a prompt: the 1-4 photos say who to swap in. A prompt is allowed up to 2000 characters. Request body and limits below.

No, H3 Max Recast does not need a prompt on Sume. Send one source video in video_url, one to four person photos in reference_image_urls, and the inspected source length in duration; the photos tell the model who to swap in. If you add a prompt, Sume accepts up to 2000 characters after trimming whitespace.
This is unusual in the Sume video catalog, where every other video model requires a non-empty prompt. The rule comes from the Sume API code, and the model's behavior comes from fal's H3 Max Recast page (read 2026-10-02): recast the people in a video using reference photos while preserving the source motion, camera, cuts, and audio. The Video Router docs describe the same row: prompt optional, duration is the source length.
What does a request without a prompt look like?
Recast goes through the Video Router, not through a text-to-video shape. There is no aspect_ratio, no generate_audio and no first or last frame, because the output keeps the source's framing and soundtrack. resolution is 768p (Sume's default) or 1080p.
Submit it async with an Idempotency-Key, then follow the status and result URLs in the envelope, as described in Jobs and results.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recast-no-prompt-001" \
-d '{
"model": "h3-max-recast",
"video_url": "https://example.com/source.mp4",
"reference_image_urls": ["https://example.com/new-person.jpg"],
"duration": 12,
"resolution": "768p",
"mode": "async"
}'How does Sume treat a prompt you do add?
Sume trims the prompt. A prompt that is empty or only spaces counts as no prompt, so it does not trigger the "Required" error that other video models return. A trimmed prompt over 2000 characters is refused before any paid submission, with the message that the prompt must be at most 2000 characters.
The prompt is forwarded to fal along with the video, the photos and the resolution. Sume never forwards duration: fal bills and sizes the output from the source, and Sume uses duration only to price and reserve the job.
| Prompt value | What Sume does |
|---|---|
| Omitted | Accepted; the photos drive the swap |
| Empty or whitespace only | Treated as omitted |
| 1 to 2000 characters after trim | Trimmed and forwarded to fal |
| Over 2000 characters | Refused before reservation or submit |
When is a prompt worth adding?
Add one when a clip has several people and you want to say which one changes. Keep it short and about the swap, because the video already carries the scene, motion, camera and audio. A prompt shaped like "Person 1 becomes the woman in the photo" is the form Sume's own tests use.
One honest gap: neither the Sume docs nor the fal page I read states how photos are matched to people when there are several, or whether person numbering in a prompt is the supported syntax. Test on a short clip at 768p first, and do not assume order from the list.
Where does the duration number come from?
duration is not a length you choose. For Recast it is the inspected length of the source, 5 to 30 seconds, and Sume rounds it up to whole seconds for pricing. Run a video inspect on the clip first: probe and stills are unbilled, and the probe carries the facts you need.
Video inspect reads clips that already live on media.sume.com in your workspace, so a clip from elsewhere needs a media import first. The Recast video_url itself only has to be a public https URL without embedded credentials, and one that is not localhost or a private address.
Why this matters for a prompt-free request: with no prompt, the duration, the photos and the resolution are the only inputs, so a wrong duration is the most likely thing to go wrong. A source of 4 seconds or 31 seconds is refused up front, so inspect the source rather than guessing.
What else does the request refuse?
Recast is a narrow contract. Photos must be public https URLs, one per new person, from one to four. A duration outside 5 to 30 seconds is refused, model_params is refused, and so are fields that belong to other video models. The full list is in the rejected-fields post. Price is per output second, so a 12 second clip at 768p reserves 12 seconds at list times 1.25; the per-clip cost post has the math.
Sources
Related posts
More in Models
- Kling 3.0 4K in Sume's Videos panel: what actually gets submitted
The panel lists 4K for Kling 3.0 but its live submit maps 4K down to 1080p. Where 4K does exist on Sume's API, and how to check the model before you pay.
- Kling motion control keep_original_sound: get a silent clip
keep_original_sound on Sume's Kling 3.0 Motion Control defaults to true, so the driving video's audio rides along. Send false for a silent clip.
- LTX-2.3 LoRAs on LTX-2.5: most run unchanged; Sume has no LoRA field
Lightricks says the large majority of LTX-2.3 LoRAs and IC-LoRAs run on LTX-2.5 unchanged, with a few exceptions to test. Sume has no LoRA field.
- LTX-2.5 distilled or dev transformer: which file to download
Download the distilled bf16 transformer to generate (fixed 8-step schedule, CFG 1) and the dev one to train. int8 files are ComfyUI-only. Sume lists no LTX id.
Written by Sume