Colossyan intrinsicDurationTrackReference vs Sume scene duration

Colossyan can size a scene to the actor's speech. In Sume avatar videos you set voice.duration per scene, silence beats need it, and the total stays 4-60 s.

4 min readSume
All posts

Colossyan lets a scene take the length of its actor's speech by pointing intrinsicDurationTrackReference at a track's referenceId. In Sume's avatar-video docs the multi-scene plan carries its own duration on each voice object, and a silence beat requires one, so you declare scene timing instead of deriving it.

Colossyan facts are from its Timing page; Sume facts from Generate avatar video, both read 2026-10-01.

How does Colossyan time a scene?

Its page lists ways to set scene duration. One is an explicit duration in milliseconds. Another is intrinsicDurationTrackReference, which Colossyan resolves by working out the referenced track's length in isolation first. Whichever you use, it says everything longer than the scene is cut out.

What does a Sume scene declare?

Each item in video_inputs has a voice. A text voice carries script or input_text plus a duration (the docs example uses "duration": 3). A silence voice, voice.type: "silence", is a non-speaking beat: duration is required and script / input_text are not allowed.

Scene timing compared, read 2026-10-01.
QuestionColossyan Timing pageSume avatar-video docs
Explicit lengthduration in ms on the scenevoice.duration on the scene
Length from speechintrinsicDurationTrackReferenceNot documented; declare it
Silent gapNot covered on this pagevoice.type: "silence" with required duration
Whole videoNot stated on this pageEstimated 4-60 seconds inclusive

What should I do about a script that runs long?

Sume accepts scripts and multi-scene plans only when the estimated duration lands in 4-60 seconds, and the docs say to shorten longer scripts or split them into multiple jobs. Because the plan is declared, add up your duration values before you submit rather than relying on the speech length.

To check the first frames of a plan before a full render, see first-frame previews for avatar videos.

Is there a way to size scenes from speech on Sume?

The avatar-video page does not document one. If you need scene lengths that follow audio, measure the audio yourself and write the result into voice.duration.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume