ChatGPT Sora clip as reference video: Omni takes 3 s each, trim

Sora 2 left the API but a ChatGPT clip can still seed a new render. Gemini Omni Flash 1.1 on Sume takes up to 3 reference videos, each 3 s at most.

5 min readSume
All posts

If you still have Sora clips, you can use them as reference_video_urls on Sume's gemini-omni-flash-1.1, but each reference clip must be 3 seconds or shorter and you can send at most 3 of them, plus up to 10 reference images. Cut the useful 3 seconds out first with Sume's video trim, then pass the trimmed file.

The Pondero report says the Sora 2 video and audio capability remains available inside ChatGPT behind its paywall, with no programmatic route in the OpenAI API after 2026-09-24. That leaves a gap: a human can still make a Sora-style clip, but a pipeline cannot call it. Using that clip as a reference closes part of it.

The Omni reference rules

Gemini Omni Flash 1.1 reference inputs on Sume (read 2026-10-07)
InputLimitPrompt tag
reference_image_urlsUp to 10<IMAGE_REF_0> (0-based, list order)
reference_video_urlsUp to 3, each 3 s or less<VIDEO_REF_0>
reference_audio_urlsNot acceptedNone
video_url (edit mode)Cannot combine with image or reference fieldsNone

Prepare the clip in three steps

First, make the clip reachable. Trim reads only a media.sume.com video, and the docs say it does not fetch from the open internet. Import your file with POST /v1/media-imports. Second, cut a 3-second range with POST /v1/video-trim, sending video_url, start and duration: 3. The result is a new MP4 artifact, and trim is priced at $0.02 per job in the docs. Third, pass that artifact URL in reference_video_urls.

Choose precision: exact (the default) for a frame-accurate cut. keyframe can start a GOP early, so the 3 seconds you get may not be the 3 you chose.

Write the prompt against the tag

Name the reference in the prompt: for instance, keep the camera move from <VIDEO_REF_0> and put the product from <IMAGE_REF_0> in frame. Indexes are zero-based and follow the order of your lists. A prompt that mentions no tag leaves the model to guess what the reference is for.

Native audio on Omni is always on and the API rejects generate_audio: false. Output is 3 to 10 seconds at 16:9 or 9:16, at 360p, 720p, 1080p or 4K.

What this does not do

A reference is guidance, not a continuation. The docs do not describe an extend or remix call for Sora clips, so do not expect the new render to pick up where the old one ended. For changing an existing clip, Omni's video edit mode takes a video_url and a prompt, as covered in the remix post.

Check what rights apply to clips you made in ChatGPT before using them in a client deliverable; that is a terms question this page cannot answer.

A worked cost line

A 6-second 720p Omni render is priced per output second at the provider list times 1.25. The code lists $0.10 per second at 720p, so Sume's price is $0.125 per second and a 6-second clip is $0.75, plus $0.02 for each trim job. Three trimmed reference clips add $0.06 for the trims. Reference clips are inputs; the docs bill Omni by output second at the chosen resolution.

If a reference clip comes out wrong, redoing the trim costs two cents. Redoing the generation costs the full clip price, so check the trimmed file (video inspect shows its duration) before you submit the paid job.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume