Perfume unboxing captions with the scent names spelled right
Fix mis-heard product names in auto-captions by sending script_text: Sume keeps the speech timings and burns your spelling, $0.20 per clip up to 60 seconds.

To make auto-captions spell a perfume's scent names correctly, send the clip to POST /v1/video-captions with script_text set to the exact words you want on screen. Sume keeps the speech-to-text word timings as the time source and aligns your text to them, so the timing stays natural and the spelling is yours. A caption job is $0.20 for clips up to 60 seconds.
Fragrance names are the worst case for speech-to-text: invented words, borrowed words and house spellings. A caption that says "Oud Noir" as "odd no air" is worse than no caption. script_text is the cheap fix.
The price of the fix is nothing extra: a caption job costs the same $0.20 whether it uses script_text or not, so the spelling control is free. The cost is your time writing the script, which is why it is worth keeping the scent list in one place.
Which field to send
The alignment can fail in two typed ways, script_alignment_mismatch and script_alignment_failed, and the recommended next action is simplify_script_text_or_omit. A mismatch means your text and the audio are too far apart to line up, which usually means the script is not what the person said. Fix the script to match the speech, drop it, or switch to words or cues.
You can send only one of script_text, words, cues and segments. words is word-level with text, start and end in seconds and it skips speech-to-text entirely; use it when you already have a transcript with times. script_text is the lightest touch when the speech is audible and only the spelling is off.
| Field | Runs speech-to-text | Best when |
|---|---|---|
| none | Yes | Plain speech with no tricky names |
| script_text | Yes, for timing | The voice is clear and only names are mis-heard |
| words | No | You have word times from another tool |
| cues or segments | No | The clip has no speech, or you author overlay text |
Silent clips, language and restyles
Speech-to-captions work only when the clip has audible speech. A silent clip fails with caption_no_speech and next_action: use_overlay_captions; for that case send cues with text, start and end. language is only a hint to speech-to-text, such as en or fr, and it never selects a style or a font. If you do not send a style, Latin text defaults to slam.
If you need to try the same clip in another look, send source_caption_id instead of video_url: Sume reuses the source video and the word timings it already holds, so speech-to-text does not run again, and the price is still a render, so $0.20 again.
One caution on names: type each scent name exactly as it is printed on the box, with its capitals and accents. Sume aligns the burned-in text to your script, so a typo in the script is a typo in the video.
A good habit is to run one clip first with script_text and read the result at phone size before you run the rest; a mismatch shows up on the first clip, not on the tenth.
Send the script
import os
import uuid
import requests
script = ("Today I am opening the winter set. First, Oud Noir, "
"then Amber Hearth, and last Snow Fig.")
body = {
"video_url": os.environ["UNBOX_URL"], # a public HTTPS clip
"script_text": script,
"style": "slam",
"language": "en",
}
r = requests.post("https://api.sume.com/v1/video-captions",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": f"unbox-{uuid.uuid4()}"},
json=body, timeout=60)
print(r.status_code, r.json().get("request_id"))
Batch and checks
Standalone video captions take a public HTTPS video URL, and the job returns a captioned video. Poll the job, then read the result for the burned clip. If a job fails with a script alignment error, fix the script and submit again with a new Idempotency-Key, because the body changed.
For a campaign of ten scents, write the ten scripts in your own repository, keep the scent names in one list and build each script from it. That way a rename in the catalogue is a one-line change, and a re-run is ten jobs at $0.20, or $2.00. Make no claims in the script about how long the fragrance lasts, because nothing in the clip can prove it. For cue-based price text over silent clips see the 12 price variants post.
Sources
Related posts
More in Media tools
- Pet adoption video music: a 30-second hopeful shelter bed
Brief a 30-second hopeful acoustic bed for a shelter's pet adoption appeal: one generation and a one-minute render, $0.225 on Sume.
- Pick the cleanest last frame: sample 24 stills before chaining
Before chaining AI clips, sample up to 24 stills with Sume video frames (fps up to 2) and choose the best one as the next first_frame, not just the last.
- Pinterest video ads take H.264 or H.265: do you need H.265?
Pinterest video ads accept H.264 or H.265 in MP4, MOV or M4V, so an H.264 clip is fine. Sume's exact trim uses libx264; confirm any other output with a probe.
- Podcast quote clips end abruptly: STT boundary_lead_ms tail padding
Sume STT segmentation boundary_lead_ms (0 to 500 ms, default 70) sets how long a sentence's tail runs before the cut. Tune it, then cut with timeline audio.
Written by Sume