71.6% of small businesses shelved a video: a 60-second avatar fix

HeyGen's September 2026 survey says 71.6% recorded a business video and never posted it. Replace the shelved take with a 4-60 second Sume avatar clip.

5 min readSume
All posts

If you recorded a business video and never posted it, the fastest fix on Sume is to stop re-shooting: write the 30 to 60 spoken seconds as a script, create one reusable avatar, and generate a talking video with POST /v1/avatar-1.0/talking-video. The script has to estimate to 4-60 seconds, and you can look at first-frame stills before paying for the full render.

The trigger is a number from HeyGen's The State of AI Avatars 2026, a September 2026 survey of 1,000+ small business owners and solopreneurs: 71.6% had recorded a business video and never posted it. HeyGen sells avatar video, so treat the survey as a vendor's own data about its market, not a neutral census. The how-to below uses only Sume's Avatar guide and avatar video guide.

Why do people shelve videos they already recorded?

The same survey page gives the reasons, and all three are about the person on screen rather than the message:

  • 53.9% cite not looking professional enough.
  • 48.6% dislike how they look on camera.
  • 35.4% dislike how they sound.

A script-driven avatar video removes the camera from the loop. It does not remove the editing work: you still need a script worth watching, and you still decide whether the result is something you would publish.

What are the steps from shelved take to postable clip?

Each step is one job on the same API. Jobs are polled at /v1/jobs/:id/status and read at /v1/jobs/:id/result once completed, as the Avatar guide describes.

Shelved take to avatar clip on Sume Avatar 1.0 (read 2026-10-03)
StepRouteWhat you send
1. Create the presenterPOST /v1/avatar-1.0/generateavatar_handle plus an input of type prompt, props or photo
2. Review first frames (optional)POST /v1/avatar-video-previewsSame fields as the video: script, quality, aspect_ratio
3. RenderPOST /v1/avatar-1.0/talking-videoavatar_handle and exactly one of script or video_inputs
4. Read the resultGET /v1/jobs/:id/resultJob id from the submit response

What does the request look like?

Create the avatar once. The handle may include a leading @; Sume stores it without it. Then reuse the handle for every later clip.

Below, the second call renders a 9:16 clip at the default plus quality. Use a fresh Idempotency-Key for each new request.

curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: shelved-avatar-001" \
  -d '{"avatar_handle":"shop_owner","input":{"type":"prompt","prompt":"Friendly shop owner, warm, plain background"}}'

# after the avatar job is completed
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: shelved-video-001" \
  -d '{"avatar_handle":"shop_owner","aspect_ratio":"9:16","script":"Three things we changed in the shop this month, in under a minute."}'

What if the shelved video is the actual footage?

Sume does not re-cut or retouch your recording. The nearest documented option is the face swap Beta: it applies a ready avatar face onto a public HTTPS source video, planned for about 4-15 seconds with usable audio, and quality is required. Face swap keeps the source clip as the base, so it only helps for short takes you are already willing to publish.

For anything longer than 60 seconds of speech, split the script into several jobs, as the guide says.

What should you check before posting?

Use an avatar video preview to approve first frames before the full render, and read the finished clip once end to end. If the clip will carry an AI-generated presenter, say so wherever the platform or the law where you sell requires it; Sume does not decide that for you.

Sume produces clips, not live presenters, and it does not promise the clip will look better than your take. Whether it does is a judgement you make on the preview.

How much script fits in 60 seconds?

Sume does not publish a words-per-second rule in the guide; it estimates duration from the script and rejects what falls outside 4-60 seconds. The practical approach is to submit a preview first, and shorten the script if the estimate is rejected. For a longer message, split it into several jobs and join the clips.

Add inline captions when the clip will play with sound off. They are burned in after generation, and a caption failure does not fail the avatar job.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume