PixVerse API and R2 world model: Sume has no live session

PixVerse R2 is a persistent real-time world. Sume has no live session or action controls: video jobs return a clip you poll for or get by webhook.

4 min readSume
All posts

PixVerse R2 is a real-time world model, and Sume does not ship that kind of product. Sume's video API is request and response: you submit a job to POST /v1/videos, then poll it or receive one terminal webhook, and you get a fixed clip.

R2 facts are from PixVerse's announcement; Sume facts are from Video generation and Webhooks, read 2026-09-30.

What is R2, according to PixVerse?

The September 22, 2026 post says a real-time world model "generates a world that keeps running while a user is in it" instead of returning a clip. R2 takes text, references, audio and action controls into the same running world, and the post says input persists in the world's state. The example shows a player moving with WASD keys while prompts change the scene.

What does Sume's video API do instead?

Sume's docs say video generation is asynchronous: submit, receive a job id and polling URL, poll until completed, then download. You can pass callback_url and Sume POSTs to it. It sends terminal job events only, with no progress or partial deliveries, so nothing arrives while a clip renders.

Which R2 inputs map to a Sume request field?

PixVerse R2 inputs versus the Sume /v1/videos request fields, read 2026-09-30.
R2 input (vendor post)Closest Sume fieldMatch
TextpromptYes, one prompt per job
Referencesinput_references (reference images)Partly: stills for style guidance
Audiogenerate_audioPartly: audio generated with the clip, not steered live
Action controlsNoneNo request field exists
Persistent world stateNoneEach job is independent

How do I see what a model accepts?

The catalog lists each model's capabilities, including supported_input_references, which says which reference types the model accepts. Frame control exists as frame_images (first and last frames for image-to-video). See OpenRouter-compatible video generation for the request shape.

When is a clip job the wrong tool?

If the user steers the output while it plays, as in R2's game example, a request-response clip cannot do that. If you know the shot ahead of time and want a file for editing or publishing, a job fits.

Sources

Related posts

More in Models

All Models posts

Written by Sume