WorldCrafter camera control: Base weights, adapter and LoRA
TencentARC WorldCrafter-Base ships transformer weights with a matching camera adapter and LoRA. On Sume, steer camera motion in the video prompt instead.

WorldCrafter-Base is a set of weights for TencentARC's inference code, not a hosted service: the README describes base transformer weights with their matching camera adapter and LoRA, tagged image-to-video and camera-control. Sume has no adapter input. To steer the camera on Sume, describe the movement in the video prompt and start from a still with first_frame.
Weights facts are from the README, read 2026-10-01; Sume facts are from Video generation.
What does the WorldCrafter-Base README say?
It says the directory holds "Base transformer weights and their matching camera adapter and LoRA for the WorldCrafter inference code". Base is placed beside a WorldCrafter-Fast directory and reads the shared repencoder/, text_encoder/, tokenizer/, vae/ and scheduler/ folders from it, and the Base-specific transformer/ and adapter/ folders must stay together. The run command is python inference.py --model-type base. The linked paper is titled "WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory". The README does not describe how the camera adapter is driven, so I do not either.
How do I steer the camera on Sume?
Sume's docs advise including "details about motion, camera angles, lighting, and scene composition" in a video prompt. Name one camera move per shot, such as a slow dolly in or a static frame, and keep the subject description short. Examples are in AI video camera movement prompts.
| Approach | WorldCrafter-Base | Sume video request |
|---|---|---|
| Camera input | Camera adapter weights loaded by the inference code | Text in the prompt |
| Start image | Image-to-video pipeline tag | first_frame on models that list it |
| Length | Not stated in the README | Per model; minimax-h3 lists 5-15 s |
| Unsupported field | Depends on the code | Rejected with 400 unsupported_parameter |
What if I need an exact camera path?
A text prompt gives direction, not coordinates. If the README's adapter gives you the control you need, run it on your own hardware. If a described move is enough, check GET /v1/videos/models for the model's supported_frame_images and durations, then iterate on a short clip first. Billing is reserved on submit at provider list x 1.25, so short test clips are cheap to repeat.
Sources
Related posts
More in Models
- YouTube fhd, qhd, uhd thumbnails in the API: a 3840 AI image
YouTube's API notes fhd, qhd and uhd thumbnail keys for some videos. Sume's ChatGPT Image 2.5 image_size accepts a 3840 edge, enough for a 4K image.
- YuE2-3B license is CC-BY-NC-4.0: what that means vs Music Router
YuE2-3B is tagged cc-by-nc-4.0, a non-commercial license. Sume's Music Router routes sume/music-auto, lyria-3.5 and lyria-3-pro instead.
- YuE2 48 kHz stereo on a 24GB GPU vs a hosted Sume music job
YuE2-3B needs a 24GB NVIDIA GPU with BF16 for 48 kHz stereo songs. A hosted Sume music job needs no GPU, only a prompt and an Idempotency-Key for retries.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
Written by Sume