Runway Characters is real-time; Sume Avatar video is a 4-60 s job

Runway Characters makes a real-time conversational video agent from one image. Sume Avatar 1.0 renders a scripted talking video of 4 to 60 seconds.

4 min readSume
All posts

Runway Characters, dated July 19, 2026 on Runway's research page, turns a single reference image of any style into a real-time conversational video agent with lip-sync, gaze and head movement at 24 frames per second in HD. Sume's Avatar 1.0 does a different job: it renders a talking video from a ready avatar and a script, between 4 and 60 seconds, as an asynchronous job at 720p. If you need a character that answers people live, Sume does not offer that. If you need a scripted presenter clip to publish, it does.

What each one does

Runway's page says Characters has about 37 ms of model time per frame and a server response time of 1.75 seconds from the end of speech to the character's reply. It does not state API availability in the text we read.

Sume's avatar video guide says the canonical route is POST /v1/avatar-1.0/talking-video. You pass an avatar_handle and either a script or video_inputs, never both. Sume accepts scripts whose estimated duration is 4 to 60 seconds, supports aspect ratios 1:1, 3:4, 9:16, 4:3 and 16:9 with 9:16 as the default, and uses 720p resolution at this time.

Real-time agent vs rendered presenter, read 2026-10-08
ItemRunway CharactersSume Avatar 1.0 talking video
ModeReal-time conversationAsynchronous render
InputOne reference imageA ready avatar plus a script or scenes
LengthOpen-ended session4 to 60 seconds per video
Resolution and rateHD at 24 fps720p
Response1.75 s server time after speechPoll the job until completed
Quality controlNot statedstandard, plus (default) or max

Render a scripted presenter on Sume

Steps:

  • Create or choose a ready avatar and note its handle.
  • Write a script that reads in 4 to 60 seconds; split longer ones into several jobs.
  • Submit with an Idempotency-Key so a retry does not duplicate the job.
  • Poll the job or take a signed webhook, then download the result.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: presenter-demo-001" \
  -d '{"avatar_handle":"your-avatar","script":"Welcome back. Here are this week\u2019s three launches.","aspect_ratio":"9:16","quality":"plus"}'

When to choose which

Choose a real-time agent when a person types or speaks and expects an immediate on-screen reply: a support concierge, a tutor, a game character. Choose a rendered presenter when the words are known in advance and quality per take matters more than latency: product explainers, announcements, social posts.

A rendered clip can be reviewed and re-rolled before anyone sees it. A live agent cannot, which is a risk to weigh for anything customer-facing.

Budgeting the rendered route

The cost of a talking video on Sume depends on its length and the quality tier. Sume publishes the tiers as standard, plus and max, with plus as the default. Estimate by script length first: a spoken minute of English is usually 130 to 160 words, so a 60-second cap holds a script of about that size. Split anything longer into chapters, and keep each chapter's opening line self-contained so viewers who land on one part still follow it.

Because jobs are asynchronous, a batch of presenter clips can run in parallel, and idempotency keys keep retries from creating duplicates.

What Sume does not do

Sume does not run a live conversational character, does not stream avatar video, and does not take a single image as the only input for the talking-video route; it needs a ready avatar. Videos longer than 60 seconds must be split into several jobs.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume