Runway Dev image_to_video, video_to_video: one Sume route?
Runway Dev names image_to_video, video_to_video and text_to_image paths. Sume's POST /v1/videos infers the mode from your fields. What it changes in ad code.

Runway Dev's docs home lists three endpoints by path: /v1/image_to_video, /v1/video_to_video and /v1/text_to_image. On Sume you call POST /v1/videos for video and POST /v1/images for stills, and the video route works out the generation mode from which fields are present. That means an ad pipeline that makes text-to-video hooks, image-to-video product shots and reference-to-video variants uses one request shape instead of choosing a path per mode.
What the Runway Dev docs list
We read the Runway Dev docs home on 2026-10-07. It named the endpoints above, and it listed Seedance 2.5, Gen 4.5, Aleph 2.0 and GPT Image 2 as available models. The text we received did not describe request fields, so we do not compare field by field here.
| Runway Dev path | Kind of work named on the page | Closest Sume route |
|---|---|---|
| /v1/image_to_video | Image-to-video generation | POST /v1/videos with frame_images |
| /v1/video_to_video | Video transformation | POST /v1/videos with a video_url edit on supported models |
| /v1/text_to_image | Text-to-image creation | POST /v1/images |
How Sume infers the mode
The /v1/videos contract says the API infers the mode and the caller never declares it. If frame_images is present the job is image-to-video, and if both fields are sent frame_images wins. If input_references is present the job is reference-to-video. If neither is present the job is text-to-video.
A pinned opening frame goes in frame_images with frame_type first_frame. A last frame uses frame_type last_frame. A reference image that should condition the whole clip goes in input_references as an image_url, video_url or audio_url object, subject to what the chosen model accepts.
- Text-only hook: model, prompt, duration, aspect_ratio.
- Product still to motion: add frame_images with frame_type first_frame.
- Use a winning clip as a look reference: add input_references with a video_url, on models that list video_url in supported_input_references.
What this changes in ad code
With one route, a variant is just a JSON body. You can generate your hook matrix as an array of bodies, post each with its own Idempotency-Key, and treat the response uniformly: a 202 with id, polling_url, status and model. Status maps to pending, in_progress, completed, failed or cancelled.
curl -sS https://api.sume.com/v1/videos/models \
-H "Authorization: Bearer $SUME_API_KEY"
# each model lists supported_frame_images and
# supported_input_references, so you can check a body
# before you submit itWhere Sume is stricter
Sume rejects fields a model does not advertise with a 400 instead of silently dropping them. size, seed and a non-empty provider.options are 400 unsupported_parameter on /v1/videos. For ad tests this matters: you cannot pin a seed to repeat a clip, so plan variants as separate prompts, not as seed sweeps.
We did not test Runway's endpoints, so we make no claim about their field-level behavior. The point is narrower: if your code today picks a path per mode, the Sume route lets you pick a body per variant.
Checking a body before you spend
Because the catalog controls validation, the cheapest habit is to read GET /v1/videos/models once and check each variant body against it before you submit. Each model entry lists supported_resolutions, supported_aspect_ratios, supported_durations, supported_frame_images and supported_input_references. If a value is not advertised, the API answers 400 instead of dropping it, which is good for correctness but means a bad variant in a batch fails loudly.
For example, the catalog documents that Kling and Grok reject references that Seedance accepts, which is why Sume added supported_input_references to the descriptor. An ad pipeline that sends a video_url reference to every model in a loop will get 400s on the models that do not take one. Filter your model list on the descriptor first, then build bodies.
- Read the catalog once per batch, not once per variant.
- Skip models whose supported_frame_images lacks last_frame if your variants need an end card.
- Keep a table of model id to allowed durations next to your hook sheet.
Polling and fetching the result
The polling shape is the same for every mode. GET /v1/videos/{job_id} returns the job, and when status is completed the unsigned_urls array holds a content URL. GET /v1/videos/{job_id}/content?index=0 redirects to the artifact. If you call content while the job is still running you get a retryable 409 job_not_completed. After a terminal failure you get 409 job_failed with retryable false, so your loop should stop instead of polling forever.
You can also read the same job from GET /v1/jobs/{id}/status and GET /v1/jobs/{id}/result in the normal Sume envelope. The id and generation_id are the same Sume job id.
Sources
Related posts
More in Comparisons
- Seedream 5.0 Lite or Nano Banana 2.1 for product shots on Sume
Seedream 5.0 Lite bills $0.04375 per image on Sume against $0.10 for Nano Banana 2.1 at 1K. Ratios, tiers, references and when the extra cost is worth paying.
- Text-to-video or image-to-video: when the prompt alone is enough
Use text-to-video when the look is open, and image-to-video when a frame, a face or a product must match. The Sume models that take each, and a decision rule.
- Can I use Suno Speech beta for an ad voiceover? What to check first
Suno Speech beta makes one track with voice and music. Before using it for ads, check price, languages and edits, then see how Sume splits voice and music.
- Zoom in on a small product in frame: AI recompose or crop and upscale
Product too small in the photo? A Pillow crop plus Sume Image Upscale ($0.20) keeps real pixels; an Ideogram 4.5 recompose edit costs $0.075 and redraws them.
Written by Sume