H3 Max 3D to Video: previs to photoreal on fal, not on Sume

fal's H3 Max 3D-to-Video turns a blockout render into photoreal video for $0.50 a request plus per second. Sume lists no such endpoint; what it offers instead.

5 min readSume
All posts

H3 Max 3D to Video is a fal endpoint that takes a 3D render, blockout or previs animation and returns photorealistic video that follows the input's layout, camera movement and timing. Sume does not list it: the Sume video catalog has minimax-h3-max for text, frame and reference-to-video, and no row for converting a 3D render.

The facts about the endpoint come from fal's H3 Max 3D-to-Video page (read 2026-10-02). What Sume lists comes from its Video generation and Video Router docs, and a search of the Sume repository for the endpoint name found nothing.

What does the fal endpoint take and charge?

Per fal's page, the video input is required and can be MP4, MOV, WEBM, M4V or GIF. Reference images are optional: if you provide none, up to two are generated for you. A scene intent text is optional, and quality is 480p, 768p or 1080p. Each request has a 5 second minimum, durations round up to whole seconds, and the price has a flat fee on top of the per-second rate.

The page's own worked example is a 5 second 16:9 clip at 768p with your own reference images, at about $0.96. Those are fal's numbers, not Sume's. If Sume added this endpoint it would bill list times 1.25, as it does for every video row, but that is a rule about Sume pricing and not a plan for this endpoint.

fal H3 Max 3D-to-Video pricing, from fal's page (read 2026-10-02)
ComponentPrice
Processing fee$0.50 per request
Video, 480p$0.05 per second
Video, 768p$0.08 per second
Video, 1080p$0.16 per second
Input tokens beyond 4,096 included$0.02 per 1,000
Auto-generated reference image$0.10 per image

What does Sume offer for the same job?

The closest rows are generation with references. minimax-h3-max accepts image, video and audio references, 5 to 15 seconds, at 480p, 768p or 1080p (1080p is a latent refinement from native 768p), with native stereo audio. You can pass your blockout as a reference video and describe the look in the prompt.

That is not the same product. A reference video on a general generation row is documented as a reference, and the Sume docs do not say it follows a blockout's layout and timing the way a dedicated previs conversion does. Test it on a short clip before you plan a pipeline around it. The limits for references (images, videos, audio, total) are in the reference limits post.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: previs-test-001" \
  -d '{
    "model": "minimax-h3-max",
    "prompt": "Photoreal street at dusk, camera path and pacing follow the reference video",
    "duration": 5,
    "resolution": "768p",
    "reference_video_urls": ["https://example.com/blockout.mp4"],
    "mode": "async"
  }'

Should you wait for Sume to add it?

I cannot say. Nothing in the repo or docs I searched announces a 3D-to-video row, so treat it as absent. Check the live catalog with GET /v1/video-router/models for what is listed today; the docs tell you to read capabilities from there rather than assuming.

If previs is central to your work, run fal's endpoint directly for that step and bring the result back as an input. If you only need loose guidance from rough footage, the reference route above is the nearest thing Sume ships.

Sources

Related posts

More in Models

All Models posts

Written by Sume