Genjutsu 30 reference images vs Sume's 1 to 8
Higgsfield's Genjutsu takes up to 30 reference images. Sume's higgsfield-genjutsu takes one source video plus 1 to 8 reference images, at 480p or 720p.

Higgsfield's Genjutsu takes a 4 to 30 second video and up to 30 reference images per generation, with output up to 1080p. Sume's higgsfield-genjutsu is narrower: one source video, 1 to 8 reference images, 480p or 720p, and 4 to 30 seconds. If you plan more than 8 references or a 1080p output, the Sume limits are the ones to design around.
Limits side by side
The Higgsfield column is from the Genjutsu announcement; the Sume column is from the Video Router and video generation docs.
| Item | Higgsfield | Sume |
|---|---|---|
| Source video | 4-30 s | One video_url, 4-30 s |
| Reference images | Up to 30 per generation | 1 to 8 reference_image_urls |
| Output | Up to 1080p | 480p or 720p |
Staying inside the Sume limit
Pick the 8 references that matter most: the ones that carry the subject's identity or the object you want swapped in. Fewer, clearer references are easier to review than a crowd. The model is called through the Video Router route with video_url plus reference_image_urls, and it accepts video references but not audio.
Because the model is listed only when its provider is configured, read GET /v1/video-router/models before you build around it.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: genjutsu-001" \
-d '{
"model": "higgsfield-genjutsu",
"prompt": "Swap the mug for the ceramic teapot in the references",
"video_url": "https://example.com/source.mp4",
"reference_image_urls": [
"https://example.com/teapot-front.png",
"https://example.com/teapot-side.png"
],
"resolution": "720p",
"mode": "async"
}'If you need more than 8
- Split the work into two passes, each with its own set of up to 8 references.
- Merge several views of one object into a single contact-sheet image to carry more views per slot.
- For 1080p, plan an upscale as a separate step, since the model tops out at 720p on Sume.
Sources
Related posts
More in Models
- GPT Image aspect ratios: Firefly's 10 vs Sume's 17
Firefly's September 2026 update gives GPT Image 2 ten aspect ratios instead of three. Sume's image API documents 17 ratios plus auto, and custom pixels.
- Hume Octave 2 covers 11 languages: is yours one?
Hume Octave 2 is reported to cover 11 languages. Before you pick a voice engine, check your language, then verify what Sume's tts_create recorded for the job.
- HunyuanVideo 1.5 on 14 GB of VRAM: run locally or call an API
The HunyuanVideo 1.5 repo lists 8.3B parameters, 480p to 1080p and a 14 GB VRAM minimum with offloading. A local-run versus API checklist.
- HunyuanVideo 1.5 SSTA and FP8: or just call an API
HunyuanVideo-1.5 has 8.3B parameters, SSTA for a 1.87x speedup at 720p, FP8 GEMM and a 14GB VRAM floor. Decide whether to run it or call a hosted video API.
Written by Sume