Hedra Character 3 allows 10 minutes; what is the Sume equivalent?

Hedra lists Character 3 at up to 10 minutes. Sume's Avatar 1.0 takes 4-60 seconds per job, so 10 minutes means ten or more jobs joined on a timeline.

5 min readSume
All posts

Hedra's Character 3 page lists a maximum duration of 10 minutes for one generation. Sume's Avatar 1.0 talking-video route accepts a script whose estimated length is 4 to 60 seconds, so the Sume equivalent of a 10-minute talking head is at least ten separate jobs with the same avatar, joined afterwards.

That is a real difference in shape, not a gap you can configure away. Hedra is audio-driven: you bring a start frame and the full audio, and it renders the face. Sume is script-driven: you send text and an avatar_handle, and each job stays inside one minute.

What Hedra lists for Character 3

Vendor facts below come from the Hedra Character 3 page, read 2026-10-02.

  • At the listed 720p rate, a full 600-second clip is 600 x $0.05 = $30.00. That is plain arithmetic on Hedra's number, not a quote.
Hedra Character 3 as listed (read 2026-10-02)
ItemListed value
Max duration10m
Required inputsStart frame and audio
Resolutions540p, 720p, 1080p
Price per second540p 2.5 cents, 720p 5 cents, 1080p 6.25 cents

What one Sume job covers

The Generate avatar video guide says scripts and multi-scene plans are accepted when Sume estimates the duration at 4 to 60 seconds inclusive, and tells you to shorten longer scripts or split them into multiple jobs. Output resolution is currently 720p, and quality is standard, plus (the default) or max.

A single job also resolves to one avatar and one shared scene, so each part you render stands alone: same avatar_handle, same scene prompt, different slice of the script.

A ten-minute plan on Sume

Split the script at natural breaks into parts of 45 to 60 seconds of speech, submit each with its own Idempotency-Key, and poll each job as described in Jobs and results. Then assemble the parts with Timeline 1.0, which takes one audio spine of 1 to 1800 seconds plus ordered video slots.

Review the seams yourself. Each part is rendered independently, so a pose or lighting change between parts is possible and Sume does not promise otherwise. The avatar previews endpoints let you approve a first frame before paying for the full render.

  • Ten 60-second parts cover 10 minutes only if each script fills its window.
  • Costs scale per second on both sides; check the live rates before you budget.
  • Hedra requires you to supply audio up front. Sume writes the speech from your script.

When to pick which

If you already have a finished 10-minute recording and want a face on it in one call, Hedra's listed limit fits that directly. If you start from a script and want per-part control, previews and reuse of one avatar across many short clips, Sume's one-minute window is a reasonable fit, at the cost of stitching.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume