Runway Characters API: avatars, live sessions and videos

Runway Characters builds avatars from one image, a voice and a personality, for live sessions or rendered videos. Endpoints, limits and pricing.

5 min readSume
All posts

Runway Characters is Runway's avatar API. You create an Avatar from a single reference image plus a voice, a personality and optional knowledge documents, then either talk to it live in a realtime session over WebRTC, or render a video of it speaking with POST /v1/avatar_videos from a text script or an audio file.

The Runway facts come from Runway's developer docs, read on 2026-09-28: its AI context page, the Characters docs, the API reference and the pricing page. The Sume facts come from the Create new avatar and Generate avatar video docs.

What is in the Runway Characters API?

Runway describes Characters as real-time conversational characters powered by GWM-1, its General World Model. An Avatar is the reusable configuration (reference image, voice, personality, knowledge); a Session is one live conversation. Characters also support tool calling, so an Avatar can invoke your functions mid-conversation.

Endpoint groups from Runway's AI context page, read 2026-09-28.
GroupPathWhat it does
Avatars/v1/avatars, /v1/avatar_conversations (9 endpoints)Create and manage Avatars, list conversations, read usage
Avatar Videos/v1/avatar_videos (1 endpoint)Render a video from an Avatar instead of holding a live Session
Realtime Sessions/v1/realtime_sessions (3 endpoints)Open, inspect and end a live conversation
Knowledge/v1/documents (5 endpoints)Documents an Avatar can answer from
Voices/v1/voices (6 endpoints)Designed and cloned custom voices, plus previews

How do I create a custom Runway avatar?

Call POST /v1/avatars. The documented fields and limits:

  • name (required, up to 50 characters) and referenceImage (required): an HTTPS URL, a Runway URI or a data URI.
  • personality (required, up to 10,000 characters): the system prompt for how the avatar behaves in conversations. startScript (optional, up to 2,000) is its opening line.
  • voice (required): a runway-live-preset voice by presetId, or a custom voice by id. The Voices API designs a voice from a text prompt or clones one from an audio sample.
  • documentIds (optional, up to 50) attaches knowledge documents.
  • The avatar's status moves through PROCESSING, READY or FAILED.

Can Runway render an avatar video instead of a live call?

Yes. POST /v1/avatar_videos with model: "gwm1_avatars" starts an asynchronous task that renders the avatar speaking. avatar is a runway-preset (nine preset ids, such as influencer and cooking-teacher) or a custom avatar by avatarId. speech is either type: "audio" with an audio file, or type: "text" with a script of up to 5,000 characters; an optional voice overrides the avatar's voice for that text. Poll GET /v1/tasks/:id for the result.

For live calls, POST /v1/realtime_sessions takes the same avatar object, and maxDuration defaults to 300 seconds.

How much does Runway Characters cost?

Runway's pricing page lists gwm1_avatars under Real-time Pricing at "2 credits upfront, then 2 credits per 6 seconds", and credits cost $0.01 each. Runway bills realtime sessions while the Character worker is active. Its docs say every new account includes 600 free credits, "roughly 30 minutes of Character video". The pricing page we read gives no separate rate for avatar_videos renders, so confirm that rate with Runway before you budget rendered clips.

What does the same job look like on Sume?

Sume renders avatar video files; its docs describe no live avatar session (see Real-time AI avatar vs video avatar API). The rendered path works in two calls:

  • Create the avatar once with POST /v1/avatar-1.0/generate, from a prompt, profile traits, or a reference image (image_url must be a public HTTPS image). It costs $0.95 per avatar and returns an avatar handle you reuse.
  • Make a clip with POST /v1/avatar-1.0/talking-video: avatar_handle plus a script that Sume estimates at 4–60 seconds, at 720p, the documented resolution. Rates: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image).
  • Avatar speech is English only in current code. For a face that speaks a separate audio track, such as Sume text-to-speech in another language, the docs pair a still with that audio in VEED Fabric 1.0 (POST /v1/veed/fabric-1.0). Its audio_url must be Sume-hosted and at most 10 MB, per the Sume API reference, at $0.1875 per audio second (720p). More in What languages can an AI avatar speak?.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume