Symphony Agent takes prompts, images or example videos: Sume's inputs
Symphony Agent is reported to start from a prompt, images or an example video. Which Sume Video fields carry each of those inputs, and what the docs cap.

Each of the three inputs reported for Symphony Agent has a matching field on Sume's video endpoint, with caps. A search-result snippet from Stackmatix (read 2026-10-02; snippet only) says Symphony Agent analyses trends and produces promotions from prompts, images or example videos. Sume's Video route takes a prompt, an image_url, and reference images and videos. What Sume does not do is analyse trends inside the same call.
Which Sume field matches which input?
The field names come from the Video 1.0 page, which is a compatibility alias for Video Router Auto. New integrations use POST /v1/videos with model: "sume/auto", which the same page points to.
| Symphony input (reported) | Sume field | Limit in docs |
|---|---|---|
| Prompt | prompt | Required, non-empty |
| Images | image_url, reference_image_urls | 1 to 9 reference images |
| Example videos | reference_video_urls | 1 to 3 reference videos |
| Example audio | reference_audio_urls | 1 to 3; needs an image or video reference |
What happens to a reference video?
It guides generation; it is not copied. The docs describe reference_video_urls as reference-guided input, and URLs must be public HTTPS. Use footage you have the right to use as a reference. If you want to read a reference clip before planning a remix, video inspect returns probe facts and stills at no charge for the probe.
Where does trend analysis happen on Sume?
Not inside the video call. Sume has a separate trending-videos search for TikTok research, covered in the stored posts below, and you decide which trend to follow. A bundled agent that picks the trend for you removes that step; with the API it is your step.
Pick a trend, write the prompt, pass a reference if you have the rights, and set aspect_ratio to 9:16 for a vertical clip. Then label the result if it shows a synthetic face or a photorealistic scene.
What does a minimal request look like?
A prompt alone is a valid request. Add image_url when you want a first frame, end_image_url for an end frame, and reference arrays when you want the look of existing assets. Choose duration from the integers the catalog lists for the model, since supported lengths differ by model and GET /v1/videos style catalog reads show them.
Run one test clip per input type before a batch. A prompt-only clip, an image-led clip and a reference-led clip give you three samples of how each input behaves, and you can then pick the one that suits your product.
Sources
Related posts
More in Comparisons
- TikTok product avatars skip shoes, hats, sunglasses and bracelets
TikTok's Symphony help page lists products its avatars cannot show. What to try for those SKUs, and what Sume's avatar product_image does and does not claim.
- Symphony Voiceover Avatars in 30+ languages vs Sume Avatar 1.0
Symphony is reported to offer licensed-actor Voiceover Avatars in 30+ languages. What Sume's avatar docs state about languages, and what to test first.
- Together AI dynamic rate limits (429, 503) vs Sume rate_limited
Together AI publishes no fixed tiers: limits track live capacity and your recent traffic. How its 429 and 503 map to Sume's rate_limited and queue_full.
- Together AI video API alternative: Sume /v1/videos compared
Together AI's video API is create-then-retrieve with five statuses. How that maps to Sume /v1/videos, and what each side does better.
Written by Sume