Code2Video: 168 briefs and how to score your own

HeyGen's Code2Video benchmark uses 168 briefs from real launch videos. Here is how to build a smaller scoring set and render it from structured input on Sume.

4 min readSume
All posts

HeyGen's September 2026 release describes Code2Video, a benchmark launched on Kaggle with 168 briefs drawn from real launch videos. To score your own motion graphics, write a fixed brief set, render each from structured input, and grade every render on the same axes. Sume's Timeline takes a declarative document, which makes reruns comparable.

What HeyGen published

The facts below come from the HeyGen September 2026 release post, read on 2026-10-03. Sume has not run this benchmark and this post reports no scores.

Code2Video benchmark as described by HeyGen (read 2026-10-03)
ItemDetail
Brief set168 briefs from real launch videos
Scoring axes5 (engagement, prompt-intent, composition, temporal, craft)
Judge accuracy predicting human preference82%, under 1 second per verdict
Top 4 modelsWithin 12 Elo of each other

What Sume Timeline is, and is not

Timeline 1.0 takes one audio spine plus ordered video[] slots and returns one MP4. The server compiles ffmpeg; callers never send filtergraphs. It is an assembly surface for sequencing, transitions and audio, not a code-to-video generator. Timeline compose puts one still and one video on screen at once with operation set to stack or overlay, which covers overlay-style graphics built from assets you already have.

Every input URL must already be a media.sume.com artifact or asset, so import media first with POST /v1/media-imports.

A small scoring harness

You do not need 168 briefs to learn something. Start with 12 to 20 briefs that mirror your real work and keep them fixed.

  • Write each brief as a timeline document so the structure is the input, not a prompt that drifts.
  • Run POST /v1/timeline-1.0/plan first. It is unbilled and returns duration_seconds, segment_count and estimated_cost_usd_micros without creating a job.
  • Render with an Idempotency-Key per brief and read the result from GET /v1/jobs/:id/result.
  • Score every render on the same written axes, and keep a human spot check on a sample of the judge's verdicts.
  • Store warnings[] from the result; padded or looped short sources are soft warnings, not failures.
curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "audio": {"url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 24},
    "video": [
      {"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 8},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "start": 8, "duration": 16}
    ]
  }'

Keep the judge honest

HeyGen reports its judge predicts human preference 82% of the time. Treat any automated judge as a screen, not a verdict. For your own set, measure agreement between the judge and two human raters on a sample before you trust it on the rest.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume