Why sume/auto never picks H3 Max Recast, and what to send
A person swap needs a source video and photos. Sume's Auto router never routes to Recast, so you name h3-max-recast yourself. Here is the request.

Auto picks for prompts, Recast needs inputs
If you call Sume with model: "sume/auto", Sume chooses the video model from the request. That works when the request is a prompt, a first frame or references. A person swap is different: it needs one source video and one to four photos, one per new person, and it has no text-only mode.
Sume's design notes for the Video Router say the rows that take a source video_url besides Gemini Omni Flash 1.1 are Genjutsu and H3 Max Recast, that both are explicit picks, and that sume/auto never routes to them. The Auto code in the repository has no Recast branch, which matches the note.
The request that works
Send the catalog id h3-max-recast to the Video Router. The Video Router docs describe it as swapping the people in a source video_url for 1 to 4 reference_image_urls, one photo per person, at 768p or 1080p, with the prompt optional. The duration is the source length, 5 to 30 seconds.
Read the catalog first. GET /v1/video-router/models returns capabilities, so you do not have to trust a blog post for the envelope.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recast-001" \
-d '{
"model": "h3-max-recast",
"video_url": "https://example.com/source.mp4",
"reference_image_urls": ["https://example.com/new-presenter.jpg"],
"resolution": "768p",
"duration": 12,
"mode": "async"
}'What fal says about the model
fal's page, read 2026-10-03, describes Recast as recasting the people in a video with reference photos while preserving the source motion, camera, cuts and audio, at $0.30 per second at 768p and $0.45 at 1080p. fal's own default is 1080p. Sume sends the resolution every time and defaults to 768p, so a request that omits it stays on the cheaper tier.
Checks before you pay
- The model id is
h3-max-recast, not a prompt that says "swap the person". durationis the source length rounded up, between 5 and 30.- No single shot is over 15 seconds, per the Sume catalog constraints.
- Do not send
aspect_ratio,generate_audio,bitrate_modeormodel_params; Recast rejects them. - Keep the Idempotency-Key when you retry a timed-out submit, so you do not pay twice. See Jobs and results.
Sources
Related posts
More in Models
- AI video releases and shutdowns in 2026: one dated timeline
Twelve dated events from Google, Luma, Runway, OpenAI and ElevenLabs, January to October 2026, each taken from a vendor page read on 2026-10-03.
- eleven_v4 vs eleven_v4_turbo: model IDs, endpoints, which to pick
ElevenLabs lists eleven_v4 for expressive speech with cloning in 90+ languages and eleven_v4_turbo at about 100 ms median latency. Which fits a video pipeline.
- ElevenLabs languages: Flash v2.5 has 32, Multilingual v2 29, v4 90+
ElevenLabs lists 32 languages for Flash v2.5, 29 for Multilingual v2 and 90+ for v4 and v4 Turbo. Check your markets against the model, then log it per job.
- Gemini 3.8 Flash TTS tops Hume's VoiceEQ board: run your own test
Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. Why a vendor-run board is only a lead, and how to run a blind A/B on your script.
Written by Sume