Runway GWM Worlds 2 is real-time 720p; Sume video jobs are async

Runway's GWM Worlds 2 (Sept 3) streams steerable 720p 24 fps video in real time. Sume video is submit-and-poll jobs; here is when each fits.

4 min readSume
All posts

Runway's GWM Worlds 2, listed on its research page with a date of September 3, 2026, is a research preview that generates continuous 720p video at 24 frames per second with 48,000 Hz audio, steered live by text actions. Sume does not offer anything like it. Every Sume video model is an asynchronous job: you submit a request, get a job id, and poll or take a webhook until the finished file is ready. If you need a world you can walk through in real time, Sume is the wrong tool. If you need a finished clip, Sume's rows are the right shape.

Two different products

Runway's page describes GWM Worlds 2 as a way to define a world, with its environment, subjects, visual style, physical rules and ambience, and then steer it with free-form text actions. The page we read gives no API availability for it. Nearby on the same page, Solaris (September 1, 2026) is described as an interface world model that generates interactive UI without traditional code.

Sume's video guide describes the opposite pattern: submit to POST /v1/videos, receive a job id and polling URL at once, poll until completed, then download the file. Video generation there usually takes from 30 seconds to several minutes, as a function of the model and parameters.

Interaction model, read 2026-10-08
QuestionRunway GWM Worlds 2Sume video jobs
OutputContinuous stream, 720p at 24 fps, 48 kHz audioA finished file per job
ControlText actions while it runsOne prompt and inputs per job
API listed on the page readNot stated (research preview)Yes, POST /v1/videos
Typical waitInteractive30 seconds to several minutes
Completion signalNot applicablePolling or signed callback_url

Approximate a walkthrough on Sume

If what you want is an explorable-feeling sequence rather than true interactivity, you can chain jobs. Generate a first clip, take its last frame as the first_frame of the next, and change the prompt for each step. It is not real time, and each step is a paid job, but it gives you a path the viewer can watch.

Steps:

  • Generate the opening shot, for example 5 seconds on Wan 3.0 at 720p, which is $0.63.
  • Extract the last frame of the result.
  • Submit the next job with that frame as first_frame and a new action in the prompt.
  • Repeat for the number of steps you need, then join the clips.
{
  "model": "wan-3.0",
  "prompt": "The camera turns left into the glass atrium, light shifting",
  "resolution": "720p",
  "duration": 5,
  "frame_images": [
    {
      "type": "image_url",
      "image_url": {"url": "https://example.com/previous-last-frame.png"},
      "frame_type": "first_frame"
    }
  ]
}

Matching the product to the need

Ask three questions. Does a person need to steer the output while it plays? Does the result need to be a file you can edit, caption and publish? Can the work wait a few minutes? If the answers are yes, no and no, you want a real-time world model. If they are no, yes and yes, you want a job-based video API.

Most marketing and product video falls in the second group, which is why a file-per-job model with a documented price per second is easy to budget. A research preview, by definition, is not something to put under a launch date.

What Sume does not do

Sume does not stream video frames, does not accept live text actions, and does not simulate physical rules over time. Chained clips are separate generations, so continuity of objects is not guaranteed. For agent workflows, hosted MCP gives you jobs_wait to hold on a job rather than poll by hand, but that is still a job, not a live stream.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume