Runway Characters voice clone: 10 s to 5 min, 10 MB; the Sume route
Runway clones a Character voice from a 10-second to 5-minute sample up to 10 MB. Sume avatar video documents text scenes, and Fabric takes your own audio_url.

Runway's custom voices page says a clone sample must be 10 seconds to 5 minutes long and at most 10 MB, or you can design a voice from a text description of at least 20 characters. Sume's avatar video request documents voice as a text script or a silence beat, with no sample upload; if you need your own recorded voice, the Fabric route takes an audio_url.
What Runway documents
From the Runway custom voices page, read 2026-10-02:
- Design models named on the page are
eleven_ttv_v3andeleven_multilingual_ttv_v2. - Voices are assigned to an avatar with
type: customand the voice id.
| Item | Documented detail |
|---|---|
| Voice design | Text description, at least 20 characters |
| Voice cloning | Audio sample, 10 seconds to 5 minutes, max 10 MB |
| Sample source | Public HTTPS URL, runway:// upload URI or data URI |
| Status values | PROCESSING, READY, FAILED |
| Voice name | Up to 100 characters |
What the Sume avatar request documents
In the avatar video guide, each video_inputs scene has a voice object: type: "text" with a script and duration, or type: "silence". The documented route has no field for uploading a voice sample, and these pages do not describe voice cloning for avatar video, so this article does not claim it.
If you already have the voice recorded
The Models page lists VEED Fabric 1.0 at POST /v1/veed/fabric-1.0: a talking still plus your audio. The body needs audio_url, a measured duration_seconds and one visual source. That moves the voice question to you: record or generate the audio first, host it at a public HTTPS URL, and Sume animates the face to it.
Mind the file size before you send it. Fabric audio has its own limit, so check the live reference rather than reusing Runway's 10 MB figure.
- Voice must match a person? Record them, then use Fabric.
- Voice just needs to sound professional? Use the avatar script route.
- Do not clone a voice you have no rights to, on any platform.
Summary
Runway makes the voice a managed object attached to a live Character. Sume keeps it simpler: speech comes from the script, or you bring finished audio to a still. Choose by whether you want a reusable cloned voice object, which Runway documents, or a render per clip.
Sources
Related posts
More in Comparisons
- Runway Characters knowledge base: 50,000 tokens per avatar vs Sume
Runway attaches up to 50,000 tokens of text or Markdown knowledge to a Character. Sume avatars have no knowledge base; the script you send is the whole content.
- Runway Characters session cap is 5 minutes; what Sume Avatar 1.0 caps
Runway Characters docs list a 5-minute session, 10,000-character personality and 2,000-character start script. Sume Avatar 1.0 renders 4-60 second clips.
- Segmind API vs Sume: PixelFlow workflows or saved Formats
Segmind turns visual PixelFlow graphs into API endpoints. Sume saves an agent thread as a Format you call over the API. How the two reuse recipes.
- Sonilo segment-level music controls vs Sume section markers
Sonilo's text-to-music lets you set styles and moods per section. On Sume you write section markers like [0:00-0:30] Intro: inside one 5000-character prompt.
Written by Sume