Code2Video: 168 briefs and how to score your own
HeyGen's Code2Video benchmark uses 168 briefs from real launch videos. Here is how to build a smaller scoring set and render it from structured input on Sume.

HeyGen's September 2026 release describes Code2Video, a benchmark launched on Kaggle with 168 briefs drawn from real launch videos. To score your own motion graphics, write a fixed brief set, render each from structured input, and grade every render on the same axes. Sume's Timeline takes a declarative document, which makes reruns comparable.
What HeyGen published
The facts below come from the HeyGen September 2026 release post, read on 2026-10-03. Sume has not run this benchmark and this post reports no scores.
| Item | Detail |
|---|---|
| Brief set | 168 briefs from real launch videos |
| Scoring axes | 5 (engagement, prompt-intent, composition, temporal, craft) |
| Judge accuracy predicting human preference | 82%, under 1 second per verdict |
| Top 4 models | Within 12 Elo of each other |
What Sume Timeline is, and is not
Timeline 1.0 takes one audio spine plus ordered video[] slots and returns one MP4. The server compiles ffmpeg; callers never send filtergraphs. It is an assembly surface for sequencing, transitions and audio, not a code-to-video generator. Timeline compose puts one still and one video on screen at once with operation set to stack or overlay, which covers overlay-style graphics built from assets you already have.
Every input URL must already be a media.sume.com artifact or asset, so import media first with POST /v1/media-imports.
A small scoring harness
You do not need 168 briefs to learn something. Start with 12 to 20 briefs that mirror your real work and keep them fixed.
- Write each brief as a timeline document so the structure is the input, not a prompt that drifts.
- Run
POST /v1/timeline-1.0/planfirst. It is unbilled and returnsduration_seconds,segment_countandestimated_cost_usd_microswithout creating a job. - Render with an
Idempotency-Keyper brief and read the result fromGET /v1/jobs/:id/result. - Score every render on the same written axes, and keep a human spot check on a sample of the judge's verdicts.
- Store
warnings[]from the result; padded or looped short sources are soft warnings, not failures.
curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio": {"url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 24},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 8},
{"source_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "start": 8, "duration": 16}
]
}'Keep the judge honest
HeyGen reports its judge predicts human preference 82% of the time. Treat any automated judge as a screen, not a verdict. For your own set, measure agreement between the judge and two human raters on a sample before you trust it on the rest.
Sources
Related posts
More in Comparisons
- Edits First Draft is iOS only: the Android and web route
Instagram's Edits First Draft is iOS only, per a secondary source. On Android or web, cut clips with Sume video trim at $0.02 per job and join them in Timeline.
- ElevenLabs TTS on fal at $0.10 per 1,000 characters vs direct and Sume
fal lists ElevenLabs TTS at $0.10 per 1,000 characters, ElevenLabs lists $0.08 for v3, Sume is $0.0475. A 2,000-character script: $0.20, $0.16, $0.095.
- fal image model list vs the Sume catalog: how to check by query
fal.ai lists Seedream 5.0, GPT Image 2.5, Flux 2, Nano Banana 2, Ideogram 4 and Krea 2. How to see which ones a Sume key can call, with a short Python script.
- fal retry budgets and 1-hour grace vs Sume job error categories
fal's changelog lists per-condition retry budgets and termination grace of up to one hour. Sume job errors carry a category, a next action and retry-after.
Written by Sume