HunyuanVideo 1.5 prompt rewrite: the vLLM endpoints to set
HunyuanVideo 1.5 rewrites prompts by default through a vLLM endpoint you set; if none is set it runs without rewriting. The env vars, the flag, and the trade.

HunyuanVideo 1.5 has a prompt rewrite step, and --rewrite defaults to true. The README says rewriting calls a vLLM-compatible endpoint that you deploy and configure, and that if no vLLM endpoint is configured the pipeline runs without remote rewriting. So you do not need a server to run generate.py, but you do need one to get the rewrite.
Which endpoints does the README name?
The README recommends Gemini or models deployed via vLLM, and says the codebase currently supports only models compatible with the vLLM API. Text-to-video and image-to-video use different recommended models and variables (README, read 2026-10-09).
| Mode | Recommended model | Environment variables |
|---|---|---|
| Text-to-video | Qwen3-235B-A22B-Thinking-2507 | T2V_REWRITE_BASE_URL, T2V_REWRITE_MODEL_NAME |
| Image-to-video | Qwen3-VL-235B-A22B-Instruct | I2V_REWRITE_BASE_URL, I2V_REWRITE_MODEL_NAME |
What does the flag do?
The argument table lists --rewrite as optional, default true: "use --rewrite false or --rewrite 0 to disable, may result in lower quality video generation". The example script has REWRITE=true with the comment to ensure the rewrite vLLM server is deployed and configured.
The README also says you may set the model names to any other vLLM-compatible model you have deployed. The two recommended models are 235B-parameter models, so a rewrite server is a separate deployment from the 8.3B video model, which needs at least 14 GB of GPU memory with offloading (model card).
- Endpoint variables set and
--rewrite true: your prompt is rewritten by the model you pointed at. - No endpoint variables set: the README says the pipeline runs without remote rewriting.
--rewrite false: no rewrite, and the README warns of possibly lower quality.
What does this look like on a hosted API?
Sume does not list HunyuanVideo; its catalog code on origin/main (read 2026-10-09) has no HunyuanVideo row. The video docs describe no prompt-rewrite option for the listed ids, so write the detail in yourself: motion, camera, lighting and scene composition, as the video docs advise.
If your reason for HunyuanVideo is cost per clip, compare the full stack: GPU time for the video model, the cost of the rewrite model, and your own hours. The low-VRAM post covers the offloading side.
How do you test rewrite on or off?
Pick ten prompts that match your real work, from a one-line idea to a long scene description. Run each twice with the same seed, once with --rewrite true and once with --rewrite false, and keep both outputs. Judge them in a blind order, and record the flag value with each clip. If long prompts do not improve, you can skip the rewrite deployment for them.
Sources
Related posts
More in Models
- HunyuanVideo 1.5: which checkpoint to download (480p, 720p, SR)
HunyuanVideo 1.5's README names 13 checkpoints, 11 with download links: 480p, 720p, CFG- and step-distilled, sparse and SR. How to choose, with the 14 GB floor.
- Hy Image 3.5 Preview: 17.44 s median latency vs Sume's 30 s wait
OpenRouter shows a 17.44 s median for Hy Image 3.5 Preview. Sume lists no Hy row, but its image calls wait 30 s, then return a 202 job. Handle both.
- Hy Image 3.5 Preview at 91.57% availability: retry math for 200 images
OpenRouter shows 91.57% availability for Hy Image 3.5 Preview. For 200 images that is about 17 failed first tries. The math, and Sume's failed-job billing.
- Hy Image 3.5 Preview footnote watermark: 16 characters, none on Sume
Tencent's Hy Image 3.5 Preview can print a custom footnote of up to 16 characters in the lower right. Sume has no watermark field; how to add one yourself.
Written by Sume