pplx-embed-v2-late for storyboard frames: Sume has no embeddings route

Perplexity's pplx-embed-v2-late embeds images and pages. Sume lists no embeddings route; use video_frames_create for stills and index them on your side.

4 min readSume
All posts

You can use pplx-embed-v2-late to search frames from Sume videos, but Sume does not run it. I found no embeddings route in Sume's public docs or API contract docs. The workable path is to pull stills with video_frames_create, embed them with Perplexity's open-weight model on your own machine, and keep the index yourself.

What is pplx-embed-v2-late?

Per its Hugging Face card, it is a late-interaction (ColBERT-style) retriever for text, images and visual documents, built on Qwen3.5, MIT licensed, with a blog post dated October 7, 2026. It stores a 128-dimension vector per token and scores with MaxSim. The 9B card says it shares an embedding space with the 0.6B variant, and that mixed text-plus-image inputs are not supported. It requires sentence-transformers 6.0.0 or newer.

That last limit matters for a storyboard: a frame with burned-in caption text counts as an image, not as a text-and-image pair.

Does Sume expose embeddings anywhere?

No. The one Sume doc that mentions the word, Sume's repository note on single-clip video understanding, lists "no indexes, collections, hosted search, embedding, batch, generation, or new transcription surfaces" for that feature, and that feature is a development-only opt-in. For the general case, the same answer applies as for other embedding models: there is no Sume endpoint to call.

Who does what in a frame-search setup, read 2026-10-11
StepWhere it happensTool
Cut stills from a finished clipSumevideo_frames_create
Wait for and fetch the stillsSumejobs_wait, jobs_result
Embed the stillsYour machinepplx-embed-v2-late (0.6B or 9B)
Store and query vectorsYour machineYour own index
Regenerate a matching shotSumegenerate_video with an idempotency_key

Which size should I pick for frames?

Because the two models share an embedding space, an index built with the 9B model can be queried with the 0.6B model's query vectors. That is stated on the 9B card. For a few hundred storyboard frames, the 0.6B model is the lower-cost place to start; judge both on your own frames, since the card's retrieval scores are for documents, not video stills.

  • Pull one still per shot rather than per second to keep the index small.
  • Keep the Sume job id next to each vector so a hit leads back to the clip.
  • Cap any regeneration with max_spend_usd and test it with dry_run first.

What can go wrong?

Late-interaction indexes are bigger than single-vector ones, because each frame yields many 128-dimension vectors rather than one. That is a property of how the model scores, as the card describes it, and it means a long project's stills can add up on disk. Decide how many frames you really need before you embed them.

Signed result URLs from a finished job expire, so download the stills when you pull them instead of storing the link. Keep the job id as your stable reference.

Last, the card's headline numbers are for document retrieval. Whether a model tuned for pages ranks cartoon frames or cut-down product shots well is something only your own queries will show.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume