Runway Characters API: avatars, live sessions and videos
Runway Characters builds avatars from one image, a voice and a personality, for live sessions or rendered videos. Endpoints, limits and pricing.
Runway Characters is Runway's avatar API. You create an Avatar from a single reference image plus a voice, a personality and optional knowledge documents, then either talk to it live in a realtime session over WebRTC, or render a video of it speaking with POST /v1/avatar_videos from a text script or an audio file.
The Runway facts come from Runway's developer docs, read on 2026-09-28: its AI context page, the Characters docs, the API reference and the pricing page. The Sume facts come from the Create new avatar and Generate avatar video docs.
What is in the Runway Characters API?
Runway describes Characters as real-time conversational characters powered by GWM-1, its General World Model. An Avatar is the reusable configuration (reference image, voice, personality, knowledge); a Session is one live conversation. Characters also support tool calling, so an Avatar can invoke your functions mid-conversation.
| Group | Path | What it does |
|---|---|---|
| Avatars | /v1/avatars, /v1/avatar_conversations (9 endpoints) | Create and manage Avatars, list conversations, read usage |
| Avatar Videos | /v1/avatar_videos (1 endpoint) | Render a video from an Avatar instead of holding a live Session |
| Realtime Sessions | /v1/realtime_sessions (3 endpoints) | Open, inspect and end a live conversation |
| Knowledge | /v1/documents (5 endpoints) | Documents an Avatar can answer from |
| Voices | /v1/voices (6 endpoints) | Designed and cloned custom voices, plus previews |
How do I create a custom Runway avatar?
Call POST /v1/avatars. The documented fields and limits:
name(required, up to 50 characters) andreferenceImage(required): an HTTPS URL, a Runway URI or a data URI.personality(required, up to 10,000 characters): the system prompt for how the avatar behaves in conversations.startScript(optional, up to 2,000) is its opening line.voice(required): arunway-live-presetvoice bypresetId, or acustomvoice byid. The Voices API designs a voice from a text prompt or clones one from an audio sample.documentIds(optional, up to 50) attaches knowledge documents.- The avatar's
statusmoves throughPROCESSING,READYorFAILED.
Can Runway render an avatar video instead of a live call?
Yes. POST /v1/avatar_videos with model: "gwm1_avatars" starts an asynchronous task that renders the avatar speaking. avatar is a runway-preset (nine preset ids, such as influencer and cooking-teacher) or a custom avatar by avatarId. speech is either type: "audio" with an audio file, or type: "text" with a script of up to 5,000 characters; an optional voice overrides the avatar's voice for that text. Poll GET /v1/tasks/:id for the result.
For live calls, POST /v1/realtime_sessions takes the same avatar object, and maxDuration defaults to 300 seconds.
How much does Runway Characters cost?
Runway's pricing page lists gwm1_avatars under Real-time Pricing at "2 credits upfront, then 2 credits per 6 seconds", and credits cost $0.01 each. Runway bills realtime sessions while the Character worker is active. Its docs say every new account includes 600 free credits, "roughly 30 minutes of Character video". The pricing page we read gives no separate rate for avatar_videos renders, so confirm that rate with Runway before you budget rendered clips.
What does the same job look like on Sume?
Sume renders avatar video files; its docs describe no live avatar session (see Real-time AI avatar vs video avatar API). The rendered path works in two calls:
- Create the avatar once with
POST /v1/avatar-1.0/generate, from a prompt, profile traits, or a reference image (image_urlmust be a public HTTPS image). It costs $0.95 per avatar and returns an avatar handle you reuse. - Make a clip with
POST /v1/avatar-1.0/talking-video:avatar_handleplus ascriptthat Sume estimates at 4–60 seconds, at 720p, the documented resolution. Rates: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). - Avatar speech is English only in current code. For a face that speaks a separate audio track, such as Sume text-to-speech in another language, the docs pair a still with that audio in VEED Fabric 1.0 (
POST /v1/veed/fabric-1.0). Itsaudio_urlmust be Sume-hosted and at most 10 MB, per the Sume API reference, at $0.1875 per audio second (720p). More in What languages can an AI avatar speak?.
Sources
Related posts
More in Sume Avatar 1.0
- Synthesia alternatives with an API: access, pricing, limits
Argil, Creatify, HeyGen and Sume each document an avatar video API. They differ in how API access is sold, the price unit and the length limits.
- Synthesia personal avatar vs studio avatar: what differs
A Synthesia personal avatar is self-serve, from one photo, ready in minutes; a studio avatar is a green-screen shoot sold as a paid add-on.
- Talking head vs B-roll: what's the difference?
A talking head is the shot of someone speaking to camera; B-roll is the footage you cut to while their voice keeps playing. How the two fit together.
- UGC video prompt for AI avatars: what goes in each field
A UGC video prompt for an AI avatar is three inputs, not one: the creator's words, a scene prompt for the setting and light, and a product image.
Written by Sume