Do [pause] tags and dashes in an avatar script cost extra seconds?
Sume Avatar 1.0 counts every space-separated token as a word, so [pause] tags and lone dashes add planned seconds. See the math and the silence-scene fix.
Yes. In Sume Avatar 1.0 the length estimate for a script counts every run of characters between spaces as one word, so a stray [pause], a lone dash or a bullet character is counted like a spoken word and can add planned seconds. Those seconds are what the tier rate multiplies, so a script padded with markers can cost more than the same script without them.
This post shows the counting rule from the Avatar 1.0 workflow code on origin/main, one worked example priced at the three tiers, and the documented way to put a real beat in a video: a silence scene. The request shapes and the 4 to 60 second window come from the Generate avatar video page.
How does Sume count words in an avatar script?
The docs say Sume accepts scripts when it estimates the target duration at 4 to 60 seconds inclusive. The estimate in the workflow code splits the script into sentences at ., ! and ?, splits each sentence on whitespace, and treats each piece as a word. There is no filtering of punctuation-only tokens.
Each planned clip then gets ceil(words / 2.8) seconds, clamped to 4 to 12 seconds, and the estimator packs sentences into clips of at most 33 words. The public price is the sum of those clip seconds times the per-second rate for your quality tier. That is also why a script that crosses 33 words can jump in cost by more than the extra words suggest, a case covered in 33 words versus 34 words.
Hellois 1 word,Hello,is 1 word,Hello - welcomeis 3 tokens.[pause]is 1 token. So is...when it has a space on each side, and so is a lone—.$14.70andhttps://sume.comare each 1 token, however long they sound when read aloud.- A script of only whitespace has no words and is refused before any spend.
What does padding with markers cost? A worked example
Take a five-sentence product script with 31 real words. The estimator puts all of it into one clip of ceil(31 / 2.8) = 12 seconds. Now add three lone dashes and two [pause] tags between sentences. That is 5 extra tokens, 36 in total, which no longer fits one 33-word clip. The estimator splits it into a 24-token clip (9 seconds) and a 12-token clip (5 seconds), for 14 seconds.
Two extra seconds from five characters of formatting. The table prices both versions with the no-product rates in the Sume catalog constants (standard $0.184, plus $0.245 and max $0.55 a second), read from the repository on 2026-10-11.
| Quality | Clean (12 s) | With markers (14 s) | Difference |
|---|---|---|---|
| standard | $2.21 | $2.58 | $0.37 |
| plus | $2.94 | $3.43 | $0.49 |
| max | $6.60 | $7.70 | $1.10 |
Do the markers change what the avatar says?
Sume's docs do not define any pause markup, SSML or bracket syntax for a script, so there is no documented effect from [pause]. Do not assume it becomes a pause. It may be read aloud, ignored or interpreted unpredictably, and either way the estimator already counted it.
That makes markers a poor trade: a documented cost, an undocumented effect. If you want rhythm, use sentence breaks. Commas and full stops cost nothing extra because they attach to the neighbouring word, and a short sentence is a natural beat.
How do I put a real beat in the video?
Use ordered video_inputs with a silence scene. The docs describe voice.type: "silence" as a beat with no speech: duration is required and script or input_text are not allowed. You pay for the planned seconds either way, but the silence is an exact length you chose, not a side effect of token counting.
Spoken scenes in the same request use type: "text" with exactly one of script or input_text. The whole plan must still land in the 4 to 60 second window, and the current execution supports one avatar and one shared scene per final video. A hook, silent demo, call-to-action layout is shown in the flash-sale example.
{
"avatar_handle": "product_host",
"quality": "standard",
"video_inputs": [
{ "id": "hook", "voice": { "type": "text", "script": "Meet the new travel mug.", "duration": 4 } },
{ "id": "beat", "voice": { "type": "silence", "duration": 2 } },
{ "id": "cta", "voice": { "type": "text", "script": "Order today and get free shipping.", "duration": 5 } }
]
}A pre-flight habit that catches this
Before a batch, strip tokens that are not words and count what is left. A one-line check is enough: split on whitespace and drop any token with no letter or digit. Then compare the count with the 33-word clip boundary.
To see the composition before you pay for the full render, create an Avatar video preview. It generates first-frame stills only, and you can still change the final render tier afterwards. Changing the script is a structural change and needs a new preview, so fix the wording first. The price ladder shows what each length costs at every tier.
Sources
Related posts
More in Sume Avatar 1.0
- Is my Sume avatar ready? Check resource_status and preview_video_url
Read a Sume Avatar 1.0 resource before your first talking video: resource_status, preview_image_url, nullable preview_video_url, voice.status, handle filter.
- UGC avatar scene prompt: why dim, glary lighting words are kept
Sume's first-frame step copies lighting words from your scene prompt as written, so ordinary light like TV glow or a side window stays. How to write it.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
Written by Sume