Creatify Boreal talking clips vs Sume's still-plus-audio route
Creatify says Boreal's gains are smallest on single-person talking clips. Sume makes every speaking shot from an accepted still plus TTS audio via Fabric.

Creatify's own numbers show single-person talking clips are where its Boreal model gains least. Sume sidesteps the question: every on-camera speaking shot is Fabric, built from an accepted still plus TTS audio, and video models are used only for wordless beats.
Boreal claims are from Creatify's 2026-09-15 launch post; Sume's routing is from the Models docs, both read 2026-10-01.
What does Creatify say about talking clips?
In its blind review against its own untouched base model, evaluators preferred Boreal in 81% of decisive comparisons overall: 83% on product ads and 85% on creator and UGC scenes, but 60% on single-person talking clips. These are Creatify's own measurements across 40 production cases, not a general result.
How does Sume route a speaking shot?
The docs say it plainly: "Every on-camera speaking shot is Fabric with an accepted still + TTS", covering short UGC, presenter ads and testimonials. They add that video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is never a video-model clip with narration laid underneath.
Which route fits which shot?
| Shot | Route in the docs |
|---|---|
| Person speaking to camera | Fabric: accepted still + TTS audio |
| Speaking, still plus audio, other model | POST /v1/minimax/h3-max/lip-sync |
| Wordless beat, B-roll, product motion | Auto image, inspect, then Auto video |
What should I do with a talking-head brief?
Generate and inspect the still, produce the audio, then send both to Fabric with the measured audio length. Keep product shots on the video route. See avatar vs lip sync vs motion control for how the endpoints differ.
Sources
Related posts
More in Models
- Deepgram nova-3-pharma vs Sume STT: drug-name transcripts
Deepgram added nova-3-pharma for English drug names. Sume STT has one public model, sume/stt-1.0, so check each drug name against word timings.
- ElevenLabs character limits by model vs Sume TTS 20,000
ElevenLabs lists 5,000 characters for v3, 10,000 for v4 and 40,000 for Flash v2.5. Sume TTS 1.0 takes up to 20,000 characters in one request.
- ElevenLabs speech to speech API: what Sume offers instead
ElevenLabs lists speech-to-speech voice changer models. Sume has no voice-to-voice endpoint: transcribe with STT, edit the text, then run TTS.
- FLUX.2 flex steps and guidance: BFL has them, Sume does not
BFL lists adjustable steps and guidance only for FLUX.2 flex ($0.06/MP). Sume lists flux.2-flex but rejects unlisted parameters with 400.
Written by Sume