Can I use my own LoRA on a hosted video API? Sume has no LoRA field
Sume's /v1/videos has no LoRA, seed or provider-option fields. MiniMax's H3 license allows LoRAs on the open weights. What that means for a character or style.

Not on Sume. The documented POST /v1/videos request has no LoRA field, no seed, and rejects any non-empty provider.options with a 400, so a custom LoRA cannot ride along with a hosted job. MiniMax's H3 FAQ says LoRAs and fine-tunes of the open weights are allowed and that you own the derivative models, with modified-file notices when you redistribute. So character or style training means running H3 weights yourself, if your region is not excluded. Facts from Sume docs and the H3 FAQ, read 2026-10-08.
What a Sume request can carry
The request table in Sume's video docs is the full list of inputs.
| Need | Field | What the docs say |
|---|---|---|
| Custom weights or LoRA | None | No such field in the request table |
| Fixed random seed | seed | No v1 model accepts it; each model reports seed false and rejects the field |
| Provider options | provider.options | A non-empty value returns 400 unsupported_parameter |
| Identity or style guidance | input_references | Reference images, plus video or audio on models that accept them (Wan 3.0, MiniMax H3, MiniMax H3 Max) |
| Start or end frame | frame_images | first_frame and last_frame entries |
What you can do on the hosted path
Reference inputs are the closest hosted substitute. Sume says Wan 3.0, MiniMax H3 and MiniMax H3 Max accept image, video and audio references, and H3's vendor page says its reference variant takes up to 9 images, 3 video clips and 3 audio clips. A reference set does not teach the model a new identity the way a trained LoRA does, but for a single character in a few shots it can be enough.
If consistency across hundreds of clips is the goal, a LoRA is the stronger tool and the hosted route does not offer it.
What the open-weights path asks for
MiniMax's FAQ says AI-Toolkit supports H3 and that character and style LoRAs, acceleration LoRAs and full fine-tunes already exist on Hugging Face. The weights need hardware (12 GB at 480p with the pruned int8 checkpoint, according to the same page) and a legal territory: the H3 license excludes the EU, UK, US and South Korea, and commercial products must show 'MiniMax H3' in the UI. HunyuanVideo's license also covers Model Derivatives, with its own territory limit.
Choosing
Match the need to the route.
- One-off clips from a prompt and a few references: hosted, $0.375 for five seconds on
minimax-h3at 768p. - A trained character or style at volume: open weights in a permitted region, on hardware you control.
- Both: train on the weights, then compare the hosted output on the same prompts before you commit.
A check you can run today
Call GET /v1/videos/models and read supported_input_references and allowed_passthrough_parameters for the id you want. Sume's docs say the passthrough list is empty for every v1 model. If a future catalog version adds a field, that endpoint is where it appears first. Until then, plan on references, not custom weights.
Sources
Related posts
More in Developers
- Phonon-2 is CC-BY-4.0: what attribution means for shipped transcripts
Phonon-2 allows commercial use under CC-BY-4.0 with attribution. What that asks of an app that ships transcripts, and where hosted Sume STT differs.
- PHP: submit a 30-second Wan 3.0 clip, poll it, save the MP4
Plain PHP 8 with curl and no SDK: send a 30-second wan-3.0 request, poll the job, and write the MP4 to disk. The 720p run is reserved at $3.75.
- Pick the cheapest Video Router model from the live catalog
Read GET /v1/video-router/models, filter by resolution, duration and capability, and rank by list x 1.25 in integer micros. Node code with an offline test.
- Pin Gemini Omni Flash 1.1 or send sume/auto: six checks
sume/auto picks the model for video and defaults to 720p and 8 s. Pin gemini-omni-flash-1.1 when you need edit mode, 4K or fixed references. A decision table.
Written by Sume