AI car video generator: from a car photo or a prompt

An AI car video generator animates a photo of your car or invents one from a prompt. How to get a driving shot, what drifts, and what it costs.

5 min readSume
All posts

An AI car video generator makes a short clip of a car driving, turning or parked in a scene, either from a text prompt or from a photo of the car. To show a real car, one you own or sell, start from its photo as the first frame and ask for one camera move; a prompt alone gives you a car the model imagines, not yours.

Below is how to do it with Sume's video API. Facts come from the Video generation, Video frames, Video captions and Timeline 1.0 docs, read on 2026-09-29; limits marked as current behavior are read from Sume's code.

How do I make a video of my car from a photo?

  • Use a sharp photo with the whole car in frame, at a public HTTPS URL.
  • Send it to POST /v1/videos as the first_frame in frame_images. That frame is your photo; every frame after it is generated.
  • Ask for one move in the prompt: a slow orbit, a push in on the front, or the car pulling away down the road. There is no camera setting in the request; the move is prompt text. Camera movement prompts lists the terms.
  • Set duration, and aspect_ratio from the model's supported_aspect_ratios: 16:9 for a site or YouTube, 9:16 for Reels and Shorts.
  • Set generate_audio to false if you will add your own music or voice; left out, it follows the model's audio capability.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: car-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "The silver hatchback pulls away slowly down a wet coastal road at dusk. Camera stays low and still.",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/car-front.jpg" }, "frame_type": "first_frame" }
    ],
    "duration": 5,
    "resolution": "720p",
    "aspect_ratio": "16:9"
  }'

Can AI make a car driving video from text only?

Yes. Send only a prompt: body style, color, road, weather, light and one camera move. The model invents the car, so use text-only clips for mood shots, intros and B-roll, not to show a specific vehicle. Name a body style rather than a make and model, and don't present an invented car as a real one.

What goes wrong in AI car videos?

Fine detail the model must redraw in every frame: the badge lettering, the characters on the number plate, the spoke pattern of the wheels, the trim line along the door. They can shift as the car turns. Pull stills with POST /v1/video-frames at the times you choose in at[] and compare them with the photo; the extract is unbilled. A clip of a car you are selling must show the same trim and color the buyer will get. If a detail changes, generate again with a smaller move, or cut the clip before the drift.

Can a dealership make one video per car?

Yes: one clip per car photo, joined with POST /v1/timeline-1.0/render. In current code each clip's own sound is dropped, so the video's sound comes from the render's audio spine, a voice-over or music, and an optional soundtrack. Add the price or offer as timed caption cues on POST /v1/video-captions; in current code the caption job refuses a video over 60 seconds or one with no audio stream, so add the sound first. The ad side, formats and placements, is in car dealership video ads with AI.

How much does an AI car video cost?

From Video generation, Video captions, Timeline 1.0 and API pricing, read 2026-09-29. Each price is plus a 5.5% agent fee by default.
StepPrice
Clip, POST /v1/videosProvider list × 1.25, by model: $1.89 for 5 s of seedance-2 at 720p 16:9; $1.05 for 5 s of kling-3 with audio
Check stills, POST /v1/video-framesUnbilled
Price text, POST /v1/video-captions$0.20 per job, videos up to 60 s
Join, POST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes

What are the limits?

  • Most catalog models make at most 15 seconds per clip; seedance-2.5 and wan-3.0 go to 30.
  • Aspect ratios differ by model. In current catalog code kling-3 takes 16:9, 9:16 and 1:1.
  • Photos must be at public HTTPS URLs; the render and the frame extract take this workspace's media.sume.com files, such as the clips Sume returned.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume